Papers by Jiwen Zhang

6 papers
VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models (2025.naacl-long)

Copied to clipboard

Challenge: Despite the impressive capabilities of large multi-modal models, their effectiveness in handling complex tasks has been limited by the prevailing singlestep reasoning paradigm.
Approach: They propose a visuallygrounded object-centric Chain-of-Thought reasoning framework for LMMs that is based on a multi-modal interleaved and aligned representation of object concepts.
Outcome: The proposed model outperforms SOTA models in CLEVR and EmbSpatial benchmarks.
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs (2026.acl-long)

Copied to clipboard

Challenge: Evaluating multimodal large language models (MLLMs) is becoming increasingly expensive as benchmarks grow in scale and cross-modality complexity.
Approach: They propose an adaptive evaluation framework for efficient benchmarking that treats evaluation as an interview-like process by keeping a hypothesized ability structure of the evaluated model and actively selecting the informative questions.
Outcome: Experiments on four representative multimodal benchmarks show that **A2-Judger significantly improves sample efficiency while maintaining reliable evaluation results.
MAGNET: Towards Adaptive GUI Agents with Memory-Driven Knowledge Evolution (2026.acl-long)

Copied to clipboard

Challenge: Mobile GUI agents powered by large foundation models can perform tasks autonomously, but frequent updates that alter UI appearance and reorganize workflows cause agents trained on historical data to fail.
Approach: They propose a memory-driven adaptive agent framework with stationary memory that links visual features to stable functional semantics and procedural memory that captures stable task intents across varying workflows.
Outcome: The proposed framework improves performance over memory-augmented baselines and offline benchmarks on AndroidWorld.
UI-Hawk: Unleashing the Screen Stream Understanding for Mobile GUI Agents (2025.emnlp-main)

Copied to clipboard

Challenge: Existing GUI agents depend on current visual observations and plain-text action history, ignoring the significance of history screens.
Approach: They propose a multi-modal GUI agent specifically designed to process screen streams . they propose UI-Hawk incorporates a history-aware visual encoder to handle the sequences .
Outcome: The proposed GUI agent can process screen streams encountered during GUI navigation.
Android in the Zoo: Chain-of-Action-Thought for GUI Agents (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on large language models (LLMs) focus on the semantics of smartphone operations.
Approach: They propose a large language model (LLM) which predicts a sequence of actions of API by analyzing past actions and visual observations.
Outcome: The proposed model improves the prediction of actions on a zero-shot Android-In-The-Zoo dataset compared to previous models .
DELAN: Dual-Level Alignment for Vision-and-Language Navigation by Cross-Modal Contrastive Learning (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies focus on cross-modal attention at the fusion stage, but modality features generated by disparate uni-encoders reside in their own spaces, leading to a decline in the quality of cross-modulation and decision-making.
Approach: They propose a framework to align navigation-related modalities before fusion by cross-modal contrastive learning.
Outcome: The proposed framework integrates with the majority of existing models, resulting in improved navigation performance on various VLN benchmarks, including R2R, R4R, and CVDN.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations