Papers by Wentao Zhou
FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs’ Responsiveness to Human Feedback (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing research focuses on benchmarking LLMs in single-turn dialogues, neglecting the nuanced nature of human feedback within real-world usage scenarios. |
| Approach: | They propose a fine-grained, multi-task benchmark designed to evaluate LLMs’ responsiveness to human feedback under real-world usage scenarios in Chinese. |
| Outcome: | The proposed benchmarks show that human feedback can significantly impact LLMs’ responsiveness in real-world usage scenarios. |
MathMixup: Boosting LLM Mathematical Reasoning with Difficulty-Controllable Data Synthesis and Curriculum Learning (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing data synthesis methods suffer from limited diversity and lack precise control over problem difficulty, making them insufficient for efficient training paradigms such as curriculum learning. |
| Approach: | They propose a data synthesis paradigm that generates high-quality, difficulty-controllable mathematical reasoning problems through hybrid and decomposed strategies. |
| Outcome: | The proposed paradigm outperforms existing methods and improves mathematical reasoning abilities. |
Reinforcement Learning with Semantic Rewards Enables Low-Resource Language Expansion without Alignment Tax (2026.findings-acl)
Copied to clipboard
Zeli Su, Ziyin Zhang, Zhou Liu, Xuexian Song, Zhankai Xu, Longfei Zheng, Xiaolu Zhang, Rong Fu, Guixian Xu, Wentao Zhang
| Challenge: | Extending large language models to low-resource languages often incurs an "alignment tax" token-level fine-tuning enforces token-level surface imitation on narrow and biased data distributions. |
| Approach: | They propose a semantic-space alignment paradigm powered by group-level semantic rewards instead of likelihood maximization. |
| Outcome: | The proposed model acquires low-resource capa- bilities while mitigating alignment tax on Tibetan–Chinese machine translation and Ti- betan headline generation. |
Inductive Relation Inference of Knowledge Graph Enhanced by Ontology Information (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to inference knowledge graphs lack ontology information, which is often too sparse. |
| Approach: | They propose a knowledge graph inductive inference method that fuses ontology information to learn the semantic information of entities. |
| Outcome: | The proposed method outperforms large language models like ChatGPT on two benchmark datasets and improves the MRR metrics by 15.4% and 44.1%, respectively. |
MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing (2026.acl-industry)
Copied to clipboard
Junbo Niu, Zheng Liu, Zhuangcheng Gu, Bin Wang, Linke Ouyang, Zhiyuan Zhao, Tao Chu, Tianyao He, Fan Wu, Qintong Zhang, Zhenjiang Jin, Guang Liang, Rui Zhang, Wenzheng Zhang, Yuan Qu, Zhifei Ren, Yuefeng Sun, Zirui Tang, Boyu Niu, Yuanhong Zheng, Dongsheng Ma, Ziyang Miao, Hejun Dong, Siyi Qian, Junyuan Zhang, Fangdong Wang, Jingzhou Chen, Xiaomeng Zhao, Liqun Wei, Wei Li, Shasha Wang, RuiLiang Xu, Yuanyuan Cao, Lu Chen, Qianqian Wu, Huaiyu Gu, Lindong Lu, Dechen Lin, null Shenguanlin, Xuanhe Zhou, Linfeng Zhang, Yuhang Zang, Xiaoyi Dong, Jiaqi Wang, Bo Zhang, Lei Bai, Pei Chu, Weijia Li, Jiang Wu, Lijun Wu, Zhenxiang Li, Guangyu Wang, Zhongying Tu, Chao Xu, Kai Chen, Bowen Zhou, Dahua Lin, Wentao Zhang, Conghui He
| Challenge: | Document images are characterized by higher resolutions, denser content, and more complex structural layouts. |
| Approach: | They propose a 1.2B-parameter document parsing vision-language model that decouples layout analysis from local content recognition. |
| Outcome: | The proposed model surpasses general-purpose and domain-specific models on multiple benchmarks while maintaining significantly lower computational overhead. |
Towards General Agentic Intelligence via Environment Scaling (2026.findings-acl)
Copied to clipboard
Runnan Fang, Shihao Cai, Baixuan Li, Jialong Wu, Guangyu Li, Wenbiao Yin, Xinyu Wang, Xiaobin Wang, Liangcai Su, Zhen Zhang, Shibin Wu, Zhengwei Tao, Yong Jiang, Pengjun Xie, Ningyu Zhang, Fei Huang, Wentao Zhang, Jingren Zhou
| Challenge: | Diverse real-world APIs require precise, robust function-calling intelligence, which needs agents to develop these capabilities through interaction in varied environments. |
| Approach: | They propose a framework that scales up environments to enable agentic intelligence . they use a two-phase agent fine-tuning strategy to first endow agents with basic agentic capabilities, then specializing them for domain-specific contexts. |
| Outcome: | Experiments on -bench, -Bench, and ACEBench show that the model significantly enhances the models’ function-calling capability. |
Dynamic Model-Bank Test-Time Adaptation for Automatic Speech Recognition (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing ASR TTA methods struggle with instability under continual and long-term distribution shifts. |
| Approach: | They propose a continuous adaptive model-bank framework that adapts to domain shifts in ASR test-time scenarios. |
| Outcome: | Experiments on diverse, continuously shifting ASR benchmarks show that DMSUTA outperforms existing continual TTA baselines. |
Memorize Step by Step: Efficient Long-Context Prefilling with Incremental Memory and Decremental Chunk (2024.emnlp-main)
Copied to clipboard
Zhiyuan Zeng, Qipeng Guo, Xiaoran Liu, Zhangyue Yin, Wentao Shu, Mianqiu Huang, Bo Wang, Yunhua Zhou, Linlin Li, Qun Liu, Xipeng Qiu
| Challenge: | Existing methods to optimize LLM for long sequences for long documents are slow and consume memory. |
| Approach: | They propose a method that starts with a small memory size and gradually increases it . they propose Decremental Chunk based on Incremental Memory (IMDC) which reduces chunk size while increasing memory size . |
| Outcome: | The proposed method is faster (1.45x) and reduces GPU memory consumption by 23.3% compared to fixed-size memory. |
Nested Browser-Use Learning for Agentic Information Seeking (2026.acl-long)
Copied to clipboard
Baixuan Li, Jialong Wu, Wenbiao Yin, Kuan Li, Zhongwang Zhang, Huifeng Yin, Zhengwei Tao, Liwen Zhang, Pengjun Xie, Jingren Zhou, Yong Jiang, Wentao Zhang, Zhiqiang Gao
| Challenge: | Existing information-seeking (IS) agents rely on the web for their information acquisition. |
| Approach: | They propose a browser-action framework that decouples interaction control from page exploration through a nested structure. |
| Outcome: | Empirical results show that NestBrowse offers clear benefits in practice. |
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding. (2026.findings-acl)
Copied to clipboard
| Challenge: | LongInsightBench is the first benchmark designed to assess models’ ability to understand long videos, with a focus on human language, viewpoints, actions, and other contextual elements. |
| Approach: | They propose a benchmark to assess models’ ability to understand long videos with a focus on human language, viewpoints, actions, and other contextual elements. |
| Outcome: | The proposed model excels in three key areas: a) long-duration, human-centric videos; b) diversifying and challenging task scenarios; c) quality assurance pipeline; and d) reliability. |
Multimodal Chemical Structure-Text Coreference in Intellectual Property via Rule-guided Reinforcement Learning (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing tools for identifying chemical structures and textual referents are inadequate for this multimodal task. |
| Approach: | They propose a RULE-guided multimodal Reinforcement learning framework for chemical structure-text coreference . RULER is a rule-driven reinforcement learning framework that uses rule-based reward functions to obtain the correct domain knowledge. |
| Outcome: | The proposed framework improves on the baseline framework and shows superior efficacy. |
Training Verifier to Assessing Complex Real-World Tool-Use Trajectories (2026.findings-acl)
Copied to clipboard
Linzhuang Sun, Mingyang Chen, Hao Liang, Tianpeng Li, Zhou Yijie, Chenzheng Zhu, Tianyu Guo, Huanyao Zhang, Jingxuan Wei, Bihui Yu, Fan Yang, Wentao Zhang
| Challenge: | Existing methods for training effective AI agents often resort to synthetic data generation. |
| Approach: | They propose a plug-and-play framework for data quality control in tool-use scenarios . they construct a tool-verify dataset and release a benchmark to assess its performance . |
| Outcome: | The proposed framework surpasses Qwen2.5-72B-Instruct on Tool-V-Bench and the previous APIGen-MT dataset. |
CFBench: A Comprehensive Constraints-Following Benchmark for LLMs (2025.acl-long)
Copied to clipboard
Tao Zhang, ChengLIn Zhu, Yanjun Shen, Wenjing Luo, Yan Zhang, Hao Liang, Tao Zhang, Fan Yang, Mingan Lin, Yujing Qiao, Weipeng Chen, Bin Cui, Wentao Zhang, Zenan Zhou
| Challenge: | Existing evaluations of Large Language Models (LLMs) focus on fragmented constraints or narrow scenarios, but they overlook the comprehensiveness and authenticity of constraints from the user’s perspective. |
| Approach: | They propose a Chinese Comprehensive Constraints Following Benchmark for LLMs that compiles constraints from real-world instructions and constructs a systematic framework for constraint types. |
| Outcome: | The proposed framework integrates multi-dimensional assessment criteria with requirement prioritization, covering various perspectives of constraints, instructions, and requirement fulfillment. |
CoMave: Contrastive Pre-training with Multi-scale Masking for Attribute Value Extraction (2023.findings-acl)
Copied to clipboard
Xinnan Guo, Wentao Deng, Yongrui Chen, Yang Li, Mengdi Zhou, Guilin Qi, Tianxing Wu, Dong Yang, Liubin Wang, Yong Pan
| Challenge: | Existing methods to extract product features from unstructured text still suffer from problems . e-commerce platforms are focusing on multi-scale values, which can be confusing . |
| Approach: | They propose a pre-training technique to automatically obtain attribute value pairs from product descriptions to aid e-commerce. |
| Outcome: | The proposed method improves on the existing token-level masking strategy and achieves state-of-the-art on four benchmarks. |
UniDataBench: Evaluating Data Analytics Agents Across Structured and Unstructured Data (2026.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks do not assess agents’ capabilities across data types . Existing tools only evaluate agents' ability to extract reasonable insights across data formats. |
| Approach: | They propose a multi-source benchmark to evaluate the performance of data analytics agents in handling diverse data sources. |
| Outcome: | The proposed agent performs end-to-end analysis over diverse data sources by automatically discovering cross-source linkages, decomposing goals, and generating robust, self-correcting code to extract actionable insights. |
The Dominance of Text Space: Unveiling the Asymmetric Nature of Cross-Modal Alignment in Large Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods for cross-modal alignment assume a symmetric interaction between visual and textual modalities, implying that both spaces adapt to each other. |
| Approach: | They propose a method that regularizes the projector to maintain the geometric structure of the text embedding space via spectral filtering. |
| Outcome: | The proposed method preserves the LLM’s inherent linguistic capabilities and reduces object hallucination significantly better than standard fine-tuning methods. |
MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought Verification (2025.acl-long)
Copied to clipboard
| Challenge: | MM-Verifier and MM Reasoner are a powerful multimodal reasoning model . large language models (LLMs) have demonstrated exceptional performance across tasks spanning myriad domains. |
| Approach: | They propose a method which combines tree search and verification to generate high-quality chain-of-thought data. |
| Outcome: | The proposed method outperforms all larger models on the MathCheck, MathVista, and MathVerse benchmarks. |
SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse Parsing (2026.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks for scientific diagram generation rely on image-centric metrics or evaluation of intermediate symbolic representations rather than final rendered images. |
| Approach: | They propose a structure-first benchmark for evaluating scientific diagram generation from pixel-level outputs. |
| Outcome: | The proposed benchmark evaluates scientific diagram generation directly from pixel-level outputs. |