Papers by Ben He
All Languages Matter: Understanding and Mitigating Language Bias in Multilingual RAG (2026.acl-long)
Copied to clipboard
Dan Wang, Guozhao Mo, Yafei Shi, Cheng Zhang, Bo Zheng, Boxi Cao, Xuanang Chen, Yaojie Lu, Hongyu Lin, Ben He, Xianpei Han, Le Sun
| Challenge: | Existing mRAG systems suffer from a language bias during reranking, systematically favoring English and the query’s native language. |
| Approach: | They propose a language-agnostic utility-driven reranker alignment technique to mitigate language bias during re-ranking. |
| Outcome: | The proposed approach mitigates language bias and consistently improves mRAG performance across languages. |
Navigating the Shadows: Unveiling Effective Disturbances for Modern AI Content Detectors (2024.acl-long)
Copied to clipboard
| Challenge: | Recent research indicates that AI-text detection systems lack robustness and struggle to effectively differentiate perturbed texts. |
| Approach: | They propose to evaluate the robustness of current detection systems by using black-box text perturbation methods and adversarial learning experiments. |
| Outcome: | The proposed methods assess the robustness of current detection models across perturbation granularities and the impact of perturbation data augmentation on the robustity of AI-text detectors. |
DiffLM: Controllable Synthetic Data Generation via Diffusion Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have significantly enhanced their knowledge and generative capabilities, leading to a surge of interest in leveraging LLMs for high-quality data synthesis. |
| Approach: | They propose a controllable data synthesis framework based on variational autoencoder which leverages diffusion models to reserve more information of original distribution and format structure in the learned latent distribution. |
| Outcome: | The proposed framework generates high-quality data with performance exceeding that of real data by 2%–7% on seven real-world datasets. |
Navigating the Infinite Dynamic Web Space: Effective In-Context Exploration via Cognitive Multi-Agent Collaboration (2026.eacl-long)
Copied to clipboard
Guozhao Mo, Yanjiang Liu, Yafei Shi, Jiawei Chen, Yang Li, Yaojie Lu, Hongyu Lin, Ben He, Le Sun, Bo Zheng, Xianpei Han
| Challenge: | Existing methods for dynamic web navigation rely on greedy strategies or value estimation, struggle to achieve effective backtracking and are heavily dependent on proprietary models. |
| Approach: | They propose a cognitive multi-agent collaboration framework that enhances cyberspace exploration capability through In-Context Exploration. |
| Outcome: | The proposed framework surpasses the proprietary model Claude-3.5 Sonnet on the WebArena benchmark. |
Memorizing is Not Enough: Deep Knowledge Injection Through Reasoning (2025.acl-long)
Copied to clipboard
Ruoxi Xu, Yunjie Ji, Boxi Cao, Yaojie Lu, Hongyu Lin, Xianpei Han, Ben He, Yingfei Sun, Xiangang Li, Le Sun
| Challenge: | Existing knowledge injection frameworks focus on knowledge memorization and retrieval, but static nature of large language models leads to outdated information as the real world evolves or when adapting to domain-specific knowledge. |
| Approach: | They propose a four-tier knowledge injection framework that defines the levels of knowledge injection: memorization, retrieval, reasoning, and association. |
| Outcome: | The proposed framework defines the levels of knowledge injection: memorization, retrieval, reasoning, and association. |
ConsistentChat: Building Skeleton-Guided Consistent Multi-Turn Dialogues for Large Language Models from Scratch (2025.emnlp-main)
Copied to clipboard
Jiawei Chen, Xinyan Guan, Qianhao Yuan, Mo Guozhao, Weixiang Zhou, Yaojie Lu, Hongyu Lin, Ben He, Le Sun, Xianpei Han
| Challenge: | Existing instruction data synthesis methods focus on single-turn instructions and neglect cross-turn coherence, resulting in context drift and reduced task completion rates. |
| Approach: | They propose a framework that constrains multi-turn instruction synthesis by explicitly modeling human conversational intent. |
| Outcome: | The proposed framework outperforms existing models trained on single-turn and multi-turn instruction datasets. |
ChatGPT Is a Knowledgeable but Inexperienced Solver: An Investigation of Commonsense Problem in Large Language Models (2024.lrec-main)
Copied to clipboard
| Challenge: | acquiring and representing commonsense in machines has posed a long-standing challenge (Li et al., 2021; Zhang e t al, 2022; Zhou e al. 2023) . |
| Approach: | They use a commonsense-based LLM to evaluate ChatGPT's commonsensing abilities by analyzing 11 datasets and generating knowledge descriptions. |
| Outcome: | The proposed model can achieve good QA accuracies while still struggling with certain domains of datasets. |
PRP-Graph: Pairwise Ranking Prompting to LLMs with Graph Aggregation for Effective Text Re-ranking (2024.acl-long)
Copied to clipboard
| Challenge: | Existing methods for pairing ranking prompting only output the same label for comparison results of different confidence intervals without considering the uncertainty of pairwise comparison. |
| Approach: | They propose a pairwise ranking prompting approach that exploits the output probabilities of target labels to capture the degree of certainty of comparison results. |
| Outcome: | The proposed method shows strong robustness and acceptable efficiency on the BEIR benchmark. |
NPRF: A Neural Pseudo Relevance Feedback Framework for Ad-hoc Information Retrieval (D18-1)
Copied to clipboard
| Challenge: | Existing neural IR models do not have a mechanism for treating expansion terms differently from the original query terms, making it difficult to combine them with existing PRF approaches. |
| Approach: | They propose an end-to-end neural PRF framework that can be used with existing neural IR models by embedding different neural models as building blocks. |
| Outcome: | Extensive experiments on two standard test collections confirm the effectiveness of the proposed framework in improving the performance of two state-of-the-art neural IR models. |
On-Policy Self-Alignment with Fine-grained Knowledge Feedback for Hallucination Mitigation (2025.findings-acl)
Copied to clipboard
Xueru Wen, Jie Lou, Xinyu Lu, Yuqiu Ji, Xinyan Guan, Yaojie Lu, Hongyu Lin, Ben He, Xianpei Han, Debing Zhang, Le Sun
| Challenge: | Large language models exhibit behavior that deviates from the boundaries of their knowledge during response generation. |
| Approach: | They propose a framework that allows large language models to explore their knowledge boundaries and self-correct generation behavior through fine-grained feedback signals. |
| Outcome: | The proposed framework enables LLMs to explore their knowledge boundaries and self-correct generation behavior through fine-grained feedback signals. |
Not All Terms Matter: Recall-Oriented Adaptive Learning for PLM-aided Query Expansion in Open-Domain Question Answering (2025.acl-long)
Copied to clipboard
| Challenge: | Open-domain question answering (ODQA) systems typically adopt a retriever-reader architecture, where the retriever finds relevant documents, and the reader extracts or synthesizes answers. |
| Approach: | They propose a method that iteratively adjusts the importance weights of QE terms based on their relevance, refining term distinction and enhancing the separation of relevant terms. |
| Outcome: | The proposed method improves retrieval accuracy and overall performance on four ODQA datasets and five QE methods. |
Contextual Interaction for Argument Post Quality Assessment (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for assessing the quality of natural language arguments are limited . existing methods focus on evaluating individual argument posts, but they often fail to distinguish between arguments with a narrow quality gap. |
| Approach: | They propose to use supervised contrastive learning to model arguments' quality . large language models with in-context examples harness the power of LLMs . |
| Outcome: | The proposed approach outperforms state-of-the-art models on a publicly available dataset . it shows that the LLMs with in-context examples are more effective than baseline models . |
BERT-QE: Contextualized Query Expansion for Document Re-ranking (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to expand query use pseudo relevance feedback (PRF) but they are under-equipped to evaluate the relevance of information pieces used for expansion. |
| Approach: | They propose a query expansion model that leverages the BERT model to select relevant document chunks for expansion. |
| Outcome: | The proposed model significantly outperforms existing models on the TREC Robust04 and GOV2 test collections. |
CoCoNUTS: Concentrating on Content while Neglecting Uninformative Textual Styles for AI-Generated Peer Review Detection (2026.acl-long)
Copied to clipboard
| Challenge: | Existing detectors rely on stylistic cues to distinguish between surface-level language refinement and genuine content generation. |
| Approach: | They propose a content-based detection paradigm to detect substantive AI-generation . they propose 'CoCoDet' detector that can detect surface-level language refinement . |
| Outcome: | The proposed detector achieves a macro F1 score of 98.24% on permissible machine-polished reviews and maintains 3.89% false positive rate on real-world reviews. |
XMC-Agent : Dynamic Navigation over Scalable Hierarchical Index for Incremental Extreme Multi-label Classification (2024.findings-acl)
Copied to clipboard
Yanjiang Liu, Tianyun Zhong, Yaojie Lu, Hongyu Lin, Ben He, Shuheng Zhou, Huijia Zhu, Weiqiang Wang, Zhongyi Liu, Xianpei Han, Le Sun
| Challenge: | Existing methods for XMC struggle with the growing set of labels due to their static label assumptions, and embedding-based methods struggle with complex mapping relationships due to late interaction paradigm. |
| Approach: | They propose a large language model (LLM) powered agent framework for extreme multi-label classification, XMC-Agent, which can effectively learn, manage and predict the extremely large and dynamically increasing set of labels. |
| Outcome: | The proposed framework can learn, manage and predict the extremely large and dynamically growing set of labels and achieves state-of-the-art performance on three standard datasets. |
Towards Imperceptible Document Manipulations against Neural Ranking Models (2023.findings-acl)
Copied to clipboard
| Challenge: | Current approaches to detect vulnerabilities in neural ranking models often introduce noticeable errors and require a well-imitated surrogate NRM to guarantee the attack effect. |
| Approach: | They propose a framework called Imperceptible DocumEnt Manipulation to produce adversarial documents that are less noticeable to both algorithms and humans. |
| Outcome: | The proposed framework outperforms strong baselines while maintaining fluency and correctness of the target documents. |
OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis (2025.acl-long)
Copied to clipboard
Qiushi Sun, Kanzhi Cheng, Zichen Ding, Chuanyang Jin, Yian Wang, Fangzhi Xu, Zhenyu Wu, Chengyou Jia, Liheng Chen, Zhoumianze Liu, Ben Kao, Guohao Li, Junxian He, Yu Qiao, Zhiyong Wu
| Challenge: | Graphical User Interface (GUI) agents powered by Vision-Language Models (VLMs) have demonstrated human-like computer control capability. |
| Approach: | They propose a GUI data synthesis pipeline that reverse engineers GUI trajectory construction process by executing pre-defined tasks. |
| Outcome: | The proposed GUI data synthesis pipeline overcomes the bottlenecks of previous methods that rely on pre-defined tasks and limited data diversity. |
Self-Steering Optimization: Autonomous Preference Optimization for Large Language Models (2025.findings-acl)
Copied to clipboard
Hao Xiang, Bowen Yu, Hongyu Lin, Keming Lu, Yaojie Lu, Xianpei Han, Ben He, Le Sun, Jingren Zhou, Junyang Lin
| Challenge: | Prior research focused on developing data generation methods, while insufficient attention has been paid to quality control mechanisms and often produces inaccurate and unhelpful data. |
| Approach: | They propose an algorithm that automatically generates high-quality preference data, eliminating manual annotation requirements. |
| Outcome: | The proposed algorithm outperforms baselines in human preference alignment and reward optimization. |
CodeRise: Bootstrapping LLMs for Ultra Low-Resource Programming Languages via Progressive Self-Refinement Curriculum (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for training data generation for low-resource languages suffer from a cold-start problem and lack diversity. |
| Approach: | They propose a two-stage framework that generates a high-quality, diverse, and progressively complex curriculum for Ultra Low-Resource Programming Languages (ULRPLs) they leverage the full formal syntax of the target language as structural guidance and apply a biased sampling strategy over library modules. |
| Outcome: | The proposed framework outperforms training-free and training-based baselines on two ULRPLs, Tengo and Janet. |
Rule or Story, Which is a Better Commonsense Expression for Talking with Large Language Models? (2024.acl-long)
Copied to clipboard
| Challenge: | Experimental results show that stories outperform rules as the expression for retrieving commonsense from LLMs, exhibiting higher generation confidence and commonsensense accuracy. |
| Approach: | They investigate the commonsense ability of large language models expressed through stories and rules to retrieve commonsensing knowledge from LLMs. |
| Outcome: | The stories outperform rules as commonsense expressions on 28 commonsensense QA datasets, exhibiting higher generation confidence and commonsence accuracy. |
Code-SPA: Style Preference Alignment to Large Language Models for Effective and Robust Code Debugging (2025.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated impressive capabilities in coding tasks like code generation and debugging. |
| Approach: | They propose a method which aligns noisy code with the well-structured style familiar to LLMs, mitigating the impact of stylistic inconsistencies. |
| Outcome: | The proposed method improves debugging performance on poorly styled code across the HumanEval, MBPP and EvalPlus datasets. |
Expanding the Boundaries of Vision Prior Knowledge in Multi-modal Large Language Models (2026.eacl-long)
Copied to clipboard
Qiao Liang, Yanjiang Liu, Weixiang Zhou, Ben He, Yaojie Lu, Hongyu Lin, Jia Zheng, Xianpei Han, Le Sun, Yingfei Sun
| Challenge: | Existing research treats MLLMs as unified systems optimized through end-to-end training, but the impact of vision encoder’s prior knowledge is seldom investigated. |
| Approach: | They propose a metric to quantify the effect of prior knowledge on MLLM performance by integrating prior knowledge at the vision encoder level into a training framework. |
| Outcome: | The proposed training framework incorporates prior knowledge at the vision encoder level, and significantly boosts visual understanding capabilities of MLLMs. |
Hidding the Ghostwriters: An Adversarial Evaluation of AI-Generated Student Essay Detection (2023.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) have exhibited remarkable capabilities in text generation tasks, but their utilization carries inherent risks, including plagiarism and the dissemination of fake news. |
| Approach: | They propose to use a dataset to construct an AI-generated student essay that employs a range of text perturbation methods to evade detection. |
| Outcome: | The proposed methods evade detection and maintain quality of the generated essays while avoiding plagiarism and fake news. |
Humanizing Machine-Generated Content: Evading AI-Text Detection through Adversarial Attack (2024.lrec-main)
Copied to clipboard
| Challenge: | Despite the development of large language models, there are still significant challenges in detecting whether text is generated by a machine. |
| Approach: | They propose a framework for a broader class of adversarial attacks to perform minor perturbations in machine-generated content to evade detection. |
| Outcome: | The proposed framework can be compromised in as little as 10 seconds, and improves over iterative adversarial learning. |
SURE or Not? Investigating Semantic Understanding in Dense Retrieval Models (2026.acl-long)
Copied to clipboard
| Challenge: | Dense retrieval models have been successful in a number of applications but it is unclear whether they truly understand semantics. |
| Approach: | They propose a benchmark for semantic understanding in dense retrieval that characterizes semantic precision, semantic abstraction and semantic equivalence along three dimensions. |
| Outcome: | The proposed model characterizes semantic understanding in dense retrieval along three dimensions: semantic precision, semantic abstraction, and semantic equivalence. |
Learning to Bootstrap for Entity Set Expansion (D19-1)
Copied to clipboard
| Challenge: | Existing bootstrapping methods for Entity Set Expansion suffer from two problems: 1) delayed feedback and sparse supervision. |
| Approach: | They propose a method that estimates delayed feedback and adaptively scores entities given sparse supervision signals. |
| Outcome: | The proposed method can estimate delayed feedback for pattern evaluation and adaptively score entities given sparse supervision signals. |
Analyze, Generate and Refine: Query Expansion with LLMs for Zero-Shot Open-Domain QA (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing methods like GAR and EAR rely heavily on supervised training and struggle to maintain effectiveness across domains and datasets. |
| Approach: | They propose a QE approach based on a three-step prompting strategy to enhance query expansion by broadening the scope of queries with additional relevant texts. |
| Outcome: | The proposed approach outperforms state-of-the-art methods in out-domain zero-shot scenarios and outperformed existing methods in end-to-end evaluations. |
Cheems: A Practical Guidance for Building and Evaluating Chinese Reward Models from Scratch (2025.acl-long)
Copied to clipboard
Xueru Wen, Jie Lou, Zichao Li, Yaojie Lu, XingYu XingYu, Yuqiu Ji, Guohai Xu, Hongyu Lin, Ben He, Xianpei Han, Le Sun, Debing Zhang
| Challenge: | Existing Chinese resources are small in scale and limited to specific domains, making them insufficient for LLM post-training. |
| Approach: | They propose a Chinese-annotated reward model and a preference dataset to address this gap . they evaluate Chinese RMs on CheemsBench and construct an RM that captures human preferences . |
| Outcome: | The proposed RM achieves state-of-the-art performance on CheemsBench and CheeMePreference. |
Contrastive Distant Supervision for Debiased and Denoised Machine Reading Comprehension (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Distant supervision (DS) is a promising learning approach for machine reading comprehension (MRC) however, the annotated dataset will inevitably lead to mislabeled instances, resulting in answer bias and context noise problems. |
| Approach: | They propose an algorithm that can learn to distinguish confusing and noisy instances via confidence-aware contrastive learning. |
| Outcome: | The proposed algorithm can learn to distinguish confusing and noisy instances via confidence-aware contrastive learning. |
Global Bootstrapping Neural Network for Entity Set Expansion (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Recent studies have shown that end-to-end bootstrapping methods only leverage local semantics rather than global semantics. |
| Approach: | They propose a global-sighted encoder to capture and encode local and global semantics into entity embedding and an attention-guided decoder to sequentially expand new entities based on these embeddables. |
| Outcome: | The proposed network achieves state-of-the-art on two bootstrapping datasets. |
TDNN: A Two-stage Deep Neural Network for Prompt-independent Automated Essay Scoring (P18-1)
Copied to clipboard
| Challenge: | Existing automated essay scoring (AES) models rely on rated essays for the target prompt as training data. |
| Approach: | They propose a shallow deep neural network to learn a prompt-dependent rating model using rated essays for non-target prompts as training data. |
| Outcome: | The proposed model improves on the standard ASAP dataset. |
Understanding Differential Search Index for Text Retrieval (2023.findings-acl)
Copied to clipboard
| Challenge: | Differentiable Search Index (DSI) is a new information retrieval framework . however, due to the black-box nature of the end-to-end neural architecture, it remains unclear to what extent it possesses basic indexing and retrieval abilities. |
| Approach: | They propose a multi-task distillation approach to enhance the retrieval quality without altering the structure of the model. |
| Outcome: | The proposed method outperforms baselines on various datasets. |
CRUXEVAL-X: A Benchmark for Multilingual Code Reasoning, Understanding and Execution (2025.acl-long)
Copied to clipboard
Ruiyang Xu, Jialun Cao, Yaojie Lu, Ming Wen, Hongyu Lin, Xianpei Han, Ben He, Shing-Chi Cheung, Le Sun
| Challenge: | Existing code benchmarks focus on code generation, while those for code reasoning are insufficient. |
| Approach: | They propose a multi-lingual code reasoning benchmark that contains 19 programming languages and at least 600 subjects for each language. |
| Outcome: | The proposed model trains on Python and achieves 34.4% Pass@1 in other languages, revealing the cross-language generalization of LLMs. |
Can LLMs Clarify? Investigation and Enhancement of Large Language Models on Argument Claim Optimization (2025.coling-main)
Copied to clipboard
| Challenge: | While Large Language Models (LLMs) have demonstrated proficiency in text rewriting tasks such as style transfer and query rewrite, their application to claim optimization remains unexplored. |
| Approach: | They propose to use a sliding window mechanism to evaluate the performance of large language models in claim clarification tasks under different settings. |
| Outcome: | The proposed model improves the performance of three LLMs on the claim clarification task under zero-shot, few-shot and supervised fine-tuning settings. |