Papers by Jiwen Zhang
VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models (2025.naacl-long)
Copied to clipboard
| Challenge: | Despite the impressive capabilities of large multi-modal models, their effectiveness in handling complex tasks has been limited by the prevailing singlestep reasoning paradigm. |
| Approach: | They propose a visuallygrounded object-centric Chain-of-Thought reasoning framework for LMMs that is based on a multi-modal interleaved and aligned representation of object concepts. |
| Outcome: | The proposed model outperforms SOTA models in CLEVR and EmbSpatial benchmarks. |
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs (2026.acl-long)
Copied to clipboard
| Challenge: | Evaluating multimodal large language models (MLLMs) is becoming increasingly expensive as benchmarks grow in scale and cross-modality complexity. |
| Approach: | They propose an adaptive evaluation framework for efficient benchmarking that treats evaluation as an interview-like process by keeping a hypothesized ability structure of the evaluated model and actively selecting the informative questions. |
| Outcome: | Experiments on four representative multimodal benchmarks show that **A2-Judger significantly improves sample efficiency while maintaining reliable evaluation results. |
MAGNET: Towards Adaptive GUI Agents with Memory-Driven Knowledge Evolution (2026.acl-long)
Copied to clipboard
| Challenge: | Mobile GUI agents powered by large foundation models can perform tasks autonomously, but frequent updates that alter UI appearance and reorganize workflows cause agents trained on historical data to fail. |
| Approach: | They propose a memory-driven adaptive agent framework with stationary memory that links visual features to stable functional semantics and procedural memory that captures stable task intents across varying workflows. |
| Outcome: | The proposed framework improves performance over memory-augmented baselines and offline benchmarks on AndroidWorld. |
UI-Hawk: Unleashing the Screen Stream Understanding for Mobile GUI Agents (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing GUI agents depend on current visual observations and plain-text action history, ignoring the significance of history screens. |
| Approach: | They propose a multi-modal GUI agent specifically designed to process screen streams . they propose UI-Hawk incorporates a history-aware visual encoder to handle the sequences . |
| Outcome: | The proposed GUI agent can process screen streams encountered during GUI navigation. |
Android in the Zoo: Chain-of-Action-Thought for GUI Agents (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies on large language models (LLMs) focus on the semantics of smartphone operations. |
| Approach: | They propose a large language model (LLM) which predicts a sequence of actions of API by analyzing past actions and visual observations. |
| Outcome: | The proposed model improves the prediction of actions on a zero-shot Android-In-The-Zoo dataset compared to previous models . |
DELAN: Dual-Level Alignment for Vision-and-Language Navigation by Cross-Modal Contrastive Learning (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing studies focus on cross-modal attention at the fusion stage, but modality features generated by disparate uni-encoders reside in their own spaces, leading to a decline in the quality of cross-modulation and decision-making. |
| Approach: | They propose a framework to align navigation-related modalities before fusion by cross-modal contrastive learning. |
| Outcome: | The proposed framework integrates with the majority of existing models, resulting in improved navigation performance on various VLN benchmarks, including R2R, R4R, and CVDN. |