Papers by Xiaocheng Zhang
Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models (2026.findings-acl)
Copied to clipboard
Hoang Phan, Xianjun Yang, Yuanshun Yao, Jingyu Zhang, Shengjie Bi, Xiaocheng Tang, Madian Khabsa, Lijuan Liu, Deren Lei
| Challenge: | Reinforcement learning with verifiable rewards (RLVR) has delivered impressive gains in mathematical and multimodal reasoning . however, the recipe introduces a significant risk of capability regression, where models forget foundational skills after prolonged training without employing regularization strategies. |
| Approach: | They propose a replay strategy with dynamic objective reweighting for general knowledge preservation using short-horizon signals of convergence and instability. |
| Outcome: | The proposed method preserves general capabilities and improves reasoning . it can be applied to existing RLVR pipelines without training additional models or tuning . |
CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention (2025.acl-long)
Copied to clipboard
Zekai Ye, Qiming Li, Xiaocheng Feng, Libo Qin, Yichong Huang, Baohang Li, Kui Jiang, Yang Xiang, Zhirui Zhang, Yunfei Lu, Duyu Tang, Dandan Tu, Bing Qin
| Challenge: | Large Vision-Language Models (LVLMs) have impressive multimodal abilities but remain prone to multilingual object hallucination. |
| Approach: | They propose a cross-lingual attention intervention method to mitigate multilingual object hallucination in LVLMs by aligning attention patterns. |
| Outcome: | The proposed method improves 13.56% (up to 30%) on the POPE and 21.75% on the hallucination subsets across languages. |
EscapeBench: Towards Advancing Creative Intelligence of Language Model Agents (2025.acl-long)
Copied to clipboard
Cheng Qian, Peixuan Han, Qinyu Luo, Bingxiang He, Xiusi Chen, Yuji Zhang, Hongyi Du, Jiarui Yao, Xiaocheng Yang, Denghui Zhang, Yunzhu Li, Heng Ji
| Challenge: | Existing language model agents excel in planning and reasoning, but lack creativity in unfamiliar environments. |
| Approach: | They propose a benchmark suite of room escape game environments to challenge agents with creative reasoning, unconventional tool use and iterative problem-solving to uncover implicit goals. |
| Outcome: | The proposed framework can perform with 40% fewer steps and hints and performs robustly across difficulty levels. |
LI4: Label-Infused Iterative Information Interacting Based Fact Verification in Question-answering Dialogue (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing studies on fact verification have failed to fully exploit question structures and ignoring relevant label information during the verification process. |
| Approach: | They propose a new approach for question-answering dialogue based fact verification using label-infused iterative information interacting. |
| Outcome: | The proposed approach achieves remarkable performance on HEALTHVER, FAVIQ, and COLLOQUIAL. |
One for All: Update Parameterized Knowledge Across Multiple Models with Once Edit (2025.acl-long)
Copied to clipboard
Weitao Ma, Xiyuan Du, Xiaocheng Feng, Lei Huang, Yichong Huang, Huiyi Zhang, Xiaoliang Yang, Baohang Li, Xiachong Feng, Ting Liu, Bing Qin
| Challenge: | Existing methods for modifying large language models focus on individual models, resulting in errors and hallucinations. |
| Approach: | They propose an ensemble-based approach that employs a plug-in model as the editing module and a dynamic weight mechanism to enhance its effectiveness. |
| Outcome: | The proposed approach outperforms existing methods while achieving superior editing efficiency. |
Controllable Text Generation via Probability Density Estimation in the Latent Space (2023.acl-long)
Copied to clipboard
| Challenge: | Existing control approaches cannot effectively model complex space with diverse attributes, high dimensionality, and asymmetric structure, leaving subsequent controls unsatisfactory. |
| Approach: | They propose a control framework using probability density estimation in the latent space and an invertible transformation function that maps the complex distributions to simple Gaussian distributions in the prior space. |
| Outcome: | The proposed method outperforms baselines on attribute relevance and text quality, achieving a new SOTA. |
WebAnchor: Anchoring Agent Planning to Stabilize Long-Horizon Web Reasoning (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for reinforcement learning (RL)-based agents struggle with long-horizon planning and strategy coherence. |
| Approach: | They propose a reinforcement learning framework that decouples planning and execution. |
| Outcome: | The proposed framework outperforms baseline and first-step RL frameworks on four benchmarks. |
Learning Which Features Matter: RoBERTa Acquires a Preference for Linguistic Generalizations (Eventually) (2020.emnlp-main)
Copied to clipboard
| Challenge: | Pretraining on self-supervised linguistic tasks is effective for learning features helpful for language understanding, but it requires more data to learn to prefer linguistic generalizations over surface ones. |
| Approach: | They propose a set of 20 ambiguous binary classification tasks to test whether a pretrained model prefers linguistic or surface generalizations. |
| Outcome: | The proposed model can learn to represent linguistic features with little pretraining data, but requires far more data to learn to prefer linguistic generalizations over surface ones. |
A Distributional Lens for Multi-Aspect Controllable Text Generation (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for multi-aspect control suffer from attribute degeneration due to mutual interference of these controllers. |
| Approach: | They propose to use attribute fusion to find the intersections of multiple attributes as their combination for generation. |
| Outcome: | The proposed method outperforms baselines on attribute relevance and text quality and achieves the SOTA. |
When Do You Need Billions of Words of Pretraining Data? (2021.acl-long)
Copied to clipboard
| Challenge: | Pretrained language models (LMs) are dominated by models that can encode billions of words. |
| Approach: | They use classifier probing, information-theoretic probing and unsupervised relative acceptability judgments to evaluate model ability. |
| Outcome: | The proposed models require only about 10M to 100M words to learn to encode most syntactic and semantic features. |
CC-Tuning: A Cross-Lingual Connection Mechanism for Improving Joint Multilingual Supervised Fine-Tuning (2025.acl-long)
Copied to clipboard
Yangfan Ye, Xiaocheng Feng, Zekun Yuan, Xiachong Feng, Libo Qin, Lei Huang, Weitao Ma, Yichong Huang, Zhirui Zhang, Yunfei Lu, Xiaohui Yan, Duyu Tang, Dandan Tu, Bing Qin
| Challenge: | Existing fine-tuning approaches that focus on English-centric training corpora often introduce implicit cross-lingual alignment, overlooking the potential for more profound, latent-level cross-linguistic interactions. |
| Approach: | They propose a multilingual fine-tuning paradigm that explicitly establishes a cross-lingual connection mechanism at the latent level. |
| Outcome: | The proposed model outperforms vanilla SFT and offers a strong latent-level alternative to data-level augmentation methods. |
TrendFact: A Benchmark Towards Hotspot Perception in Automatic Fact-Checking (2026.acl-long)
Copied to clipboard
Xiaocheng Zhang, Xi Wang, Yifei Lu, Jianing Wang, Zhuangzhuang Ye, Mengjiao Bao, Peng Yan, Xiaohong Su
| Challenge: | Existing benchmarks lack social metadata and evaluation framework to meet this urgent evaluation needs. |
| Approach: | They propose a benchmark capable of evaluating HPA and three fact-checking tasks. |
| Outcome: | The proposed framework improves HPA and computational efficiency for RLM-driven systems. |