Papers by Zhendong Mao
LAFaCT: Attribution-based Localization and Focused Sequential Analysis of Fact-Critical Tokens for Hallucination Detection (2026.acl-long)
Copied to clipboard
| Challenge: | Large Language Models suffer from hallucinations, severely undermining their reliability. |
| Approach: | They propose a framework that localizes fact-critical tokens and performs sequential analysis on their hidden states. |
| Outcome: | The proposed framework localizes fact-critical tokens using Factual Criticality . it then performs a focused sequential analysis on their hidden states . |
Improving Image Captioning via Predicting Structured Concepts (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies on image captioning ignore the relationship between concepts . current methods for image caption generation ignore this relationship . |
| Approach: | They propose a structured concept predictor to predict concepts and their structures . they integrate these predictions into captioning to enhance visual signals . |
| Outcome: | The proposed approach improves image captioning performance by using semantic concepts as a bridge between images and texts. |
Feature-Adaptive and Data-Scalable In-Context Learning (2024.acl-long)
Copied to clipboard
| Challenge: | In-context learning (ICL) is a popular way to stimulate LLM capabilities for downstream tasks due to context length constraints. |
| Approach: | They propose a feature-adaptive and data-scalable in-context learning framework which leverages task-adaptives to promote inference on the downstream task. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on 10 datasets under different data settings and LLM scale. |
E-CORE: Emotion Correlation Enhanced Empathetic Dialogue Generation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Empathy is a desirable human trait that improves the emotional perceptivity in emotion-bonding social activities. |
| Approach: | They propose a framework that integrates emotion correlation learning, utilization, and supervising. |
| Outcome: | The proposed framework improves empathetic perception and expression on a humanized dialogue dataset. |
IDEATE: Detecting AI-Generated Text Using Internal and External Factual Structures (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods to detect AI-generated text rely on internal evidences, but external evidences are not considered. |
| Approach: | They propose a hierarchical graph network that utilizes internal and external factual structures to detect AI-generated text. |
| Outcome: | The proposed network outperforms current state-of-the-art methods on four datasets. |
Zero-Shot Detection of LLM-Generated Text using Temperature Sensitivity (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods for detecting LLM-generated text rely on statistical features that are insufficient for reliable detection. |
| Approach: | They propose a temperature-sensitive detector that modulates decoding temperature and monitors how probability distributions respond to temperature. |
| Outcome: | The proposed method is based on a temperature sensitivity feature and a simple zero-shot detector built upon normalized temperature sensitivity. |
Text Style Transfer with Contrastive Transfer Pattern Mining (2023.acl-long)
Copied to clipboard
| Challenge: | Existing methods for text style transfer only focus on the transformation between styles, yet they do not take into account that this transformation can be achieved via different hidden transfer patterns. |
| Approach: | They propose a novel approach which automatically mines hidden transfer patterns to improve TST . they use a clustering module to automatically discover hidden transfer pattern from the data . |
| Outcome: | The proposed method can be applied in a plug-and-play manner to enhance other methods to further improve their performance. |
LIRE: listwise reward enhancement for preference alignment (2024.findings-acl)
Copied to clipboard
| Challenge: | prevailing approaches to preference alignment focus on pairwise comparisons, with limited exploration into multi-response scenarios. |
| Approach: | They propose a listwise reward enhancement approach that integrates offline rewards of multiple responses into a streamlined listwise framework. |
| Outcome: | The proposed approach outperforms existing methods on dialogue and summarization tasks with good transferability to out-of-distribution data. |
Automated Creativity Evaluation for Large Language Models: A Reference-Based Approach (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for evaluating creativity of machine-generated texts rely on costly manual annotations or fail to align closely with human assessments. |
| Approach: | They propose an automated method based on the Torrance Test of Creative Writing (TTCW) . |
| Outcome: | The proposed method improves the alignment between LLM evaluations and human assessments. |
Curriculum Learning for Natural Language Understanding (2020.acl-main)
Copied to clipboard
| Challenge: | Pre-trained language models can be fine tuned to perform NLU tasks in a straightforward manner. |
| Approach: | They propose a pretrain-finetune paradigm for natural language understanding (NLU) they propose 'a cross-trainset' approach that allows users to distinguish easy from difficult examples . |
| Outcome: | The proposed approach achieves significant performance improvements on a wide range of NLU tasks. |
S2ynRE: Two-stage Self-training with Synthetic data for Low-resource Relation Extraction (2023.acl-long)
Copied to clipboard
| Challenge: | Existing methods for relation extraction suffer from the inadequacy of large-scale annotated data. |
| Approach: | They propose a framework for two-stage self-training with synthetic data for relation extraction . |
| Outcome: | The proposed framework is based on two-stage self-training with synthetic data . it is able to synthesize large quantities of training data and iteratively and alternately learn from synthetic and golden data together. |
Knowledge Context Modeling with Pre-trained Language Models for Contrastive Knowledge Graph Completion (2024.findings-acl)
Copied to clipboard
| Challenge: | Text-based knowledge graph completion methods neglect knowledge contexts in inferring process. |
| Approach: | They propose a framework which models the knowledge context as additional prompts with pre-trained language models for knowledge graph completion. |
| Outcome: | The proposed framework achieves state-of-the-art on FB15k-237, WN18RR and Wikidata5M datasets. |
Fine-grained Knowledge Enhancement for Retrieval-Augmented Generation (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing studies rely on semantic similarity to retrieve knowledge but ignore fine-grained information within documents. |
| Approach: | They propose a fine-grained knowledge enhancement method to fill knowledge gaps with retrieved external information by a Chain-of-Thought prompting procedure and a decoding enhancement strategy to constrain the document-based decoding process. |
| Outcome: | The proposed method can be applied in a plug-and-play manner to enhance its performance with no additional modules or training process. |
Random Entity Quantization for Parameter-Efficient Compositional Knowledge Graph Representation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to learning on Knowledge Graphs (KGs) are not critical for learning on KGs. |
| Approach: | They propose an alternative approach to represent entities by composing entity-corresponding codewords matched from predefined small-scale codebooks. |
| Outcome: | The proposed approach achieves similar results to existing methods. |
Visual-Linguistic Dependency Encoding for Image-Text Retrieval (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing approaches to image-text retrieval ignore semantic discrepancies caused by syntactic structure in natural language expressions and relationships among visual entities. |
| Approach: | They propose a visual-linguistic dependency encoder framework which explicitly models the dependency information among textual words and interaction patterns between image regions. |
| Outcome: | The proposed framework outperforms existing methods on a vision-linguistic compositional structure reasoning dataset. |
UniRel: Unified Representation and Interaction for Joint Relational Triple Extraction (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to extract rich correlations between entities and relations are not fully exploited by existing methods. |
| Approach: | They propose to unify entities and relations by jointly encoding them within a concatenated natural language sequence and unify the modeling of interactions with a proposed Interaction Map. |
| Outcome: | The proposed method is more efficient and efficient than existing methods and can be scaled up to 2021. |
Align Documents to Questions: Question-Oriented Document Rewriting for Retrieval-Augmented Generation (2026.findings-acl)
Copied to clipboard
| Challenge: | Retrieval-Augmented Generation (RAG) enhances the factuality of Large Language Models (LLMs) however, LLMs exhibit a stylistic bias when presented with mixed contexts, revealing a bottleneck in their utility. |
| Approach: | They propose a style-controlled rewriter that aligns retrieved documents with a question-oriented style while preserving facts. |
| Outcome: | The proposed model improves RAG pipelines by 8% with negligible latency overhead. |
WildGraphBench: Benchmarking GraphRAG with Wild-Source Corpora (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks for Graph-based Retrieval-Augmented Generation (GraphRAG) rely on short, curated passages as external knowledge, failing to adequately evaluate systems in realistic settings involving long contexts and large-scale heterogeneous documents. |
| Approach: | They propose a benchmark to assess GraphRAG performance in the wild using Wikipedia's unique structure where cohesive narratives are grounded in long and heterogeneous external reference documents. |
| Outcome: | Experiments with articles across 12 top-level topics show that GraphRAG performs better in the wild than existing methods. |
From Real to Synthetic: Synthesizing Millions of Diversified and Complicated User Instructions with Attributed Grounding (2025.acl-long)
Copied to clipboard
| Challenge: | a pursuit of diverse, complex, and large-scale instruction data is crucial for automatically aligning large language models . authors: methods that generate synthetic instructions at scale suffer from limited grounding sources . attributed grounding is a technique that can be used to align language models with human . |
| Approach: | They synthesize 1 million instructions using attributed grounding and a bottom-up synthesis process that leverages web documents to generate a situation, then a meaningful instruction. |
| Outcome: | The proposed framework achieves leading performance on benchmarks and scales with more web corpora. |
FS-Researcher: Test-Time Scaling for Long-Horizon Research Tasks with File-System-Based Agents (2026.acl-long)
Copied to clipboard
| Challenge: | Long trajectories in deep research often exceed model context limits, compressing token budgets for both evidence collection and report writing. |
| Approach: | They propose a file-system-based framework that scales deep research beyond context window . a Context Builder agent acts as a librarian and a Report Writer agent composes the final report . |
| Outcome: | Experiments on two open-ended benchmarks show that FS-Researcher achieves state-of-the-art report quality across different backbone models. |
EmRel: Joint Representation of Entities and Embedded Relations for Multi-triple Extraction (2022.naacl-main)
Copied to clipboard
| Challenge: | Existing studies only explore entity representations, but propose a novel triple perspective for relation extraction. |
| Approach: | They propose to explicitly introduce relation representation and jointly represent it with entities to identify valid triples. |
| Outcome: | The proposed method is based on ablations and document-level relation extraction and joint entity and relation extraction. |
RESEMO: A Benchmark Chinese Dataset for Studying Responsive Emotion from Social Media Content (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing studies on social media text processing do not focus on responsive emotion analysis. |
| Approach: | They propose a Chinese dataset named ResEmo for responsive emotion analysis, including 3813 posts with 68,781 comments collected from Weibo, the largest social media platform in China. |
| Outcome: | The proposed dataset includes 3813 posts with 68,781 comments collected from weibo, the largest social media platform in China. |
M-RangeDetector: Enhancing Generalization in Machine-Generated Text Detection through Multi-Range Attention Masks (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing supervised methods for text detection are overfitting within their training domains. |
| Approach: | They propose a method that integrates four distinct attention masking strategies into a Multi-Range Attention module to learn various writing strategies for machine-generated text detection. |
| Outcome: | The proposed method improves the generalization capability of existing detectors on three datasets. |
Mitigating Biases in Language Models via Bias Unlearning (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent debiasing approaches target different demographic groups, harming fairness and discrimination. |
| Approach: | They propose a model debiasing framework which targets stereotypes by unlearning stereotype forgetting and anti-stereotype retention. |
| Outcome: | The proposed framework outperforms existing methods in mitigating bias while retaining language modeling capabilities. |
CodeRipple: Wavelet-Based Detection of LLM-Generated Code (2026.acl-long)
Copied to clipboard
| Challenge: | Existing training-free detectors rely on global statistics of the Token Perplexity Sequence (TPS) and struggle with code. |
| Approach: | They propose a training-free detection framework that characterizes TPS morphology across scales. |
| Outcome: | The proposed framework outperforms existing training-free detectors on three challenging benchmarks spanning programming languages, multiple generating LLMs, and various evasion strategies. |
KNN-Instruct: Automatic Instruction Construction with K Nearest Neighbor Deduction (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for generating synthetic instructions for large language models suffer from stale distribution and scalability. |
| Approach: | They propose a method which incorporates KNN deduction to produce meaningful new instructions by summarizing and learning from existing ones. |
| Outcome: | The proposed method outperforms all 7B models on the LMSYS leaderboard. |
Rationales Are Not Silver Bullets: Measuring the Impact of Rationales on Model Performance and Reliability (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing studies have shown that training language models with rationales augmentation is beneficial, but this view does not hold consistently. |
| Approach: | They conduct comprehensive investigations to thoroughly inspect the impact of rationales on model performance and a novel perspective of model reliability. |
| Outcome: | The proposed method outperforms untrained models in several areas and provides informative regulations on the broad utilization of rationales. |
Benchmarking and Improving Compositional Generalization of Multi-aspect Controllable Text Generation (2024.acl-long)
Copied to clipboard
| Challenge: | Existing MCTG methods face a noticeable performance drop in compositional testing. |
| Approach: | They propose a benchmark to evaluate compositional generalization of MCTG methods by combining multi-aspect labeled datasets and a crafted three-dimensional evaluation protocol. |
| Outcome: | The proposed framework improves compositional generalization performance by 3.64% and 94.4% in compositional testing. |
On the Calibration of Large Language Models and Alignment (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models are becoming more popular and are proving to be reliable . however, their reliability is often understudied due to their uncertainty and complex structure . |
| Approach: | They conduct a systematic examination of the calibration of aligned language models throughout the entire construction process including pretraining and alignment training. |
| Outcome: | The results shed light on whether popular large language models are well-calibrated and how the training process influences model calibration. |
Grammatical Error Correction via Mixed-Grained Weighted Training (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Empirical evaluation shows that MainGEC achieves consistent and significant performance improvements on two benchmark datasets. |
| Approach: | They propose to use mixed-grained weighted training to improve the training effect for GEC by analyzing the inherent discrepancies in annotated training data. |
| Outcome: | Empirical results show that the proposed method achieves significant performance improvements on two benchmark datasets. |
Alleviating Hallucinations in Large Language Models via Truthfulness-driven Rank-adaptive LoRA (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to improve truthfulness are training-free without modifying the LLM itself. |
| Approach: | They propose a rank-adaptive LoRA method to improve LLM truthfulness that allocates ranks according to truthfulness correlations of LLM modules. |
| Outcome: | The proposed method outperforms state-of-the-art methods on the LLM family and makes the performance of 7B LLMs exceed GPT-4. |
Improving Chinese Spelling Check by Character Pronunciation Prediction: The Effects of Adaptivity and Granularity (2022.emnlp-main)
Copied to clipboard
| Challenge: | Chinese spelling check (CSC) is a fundamental NLP task that detects and corrects spelling errors in Chinese texts. |
| Approach: | They propose an auxiliary task of Chinese pronunciation prediction to improve CSC . they propose adaptive weighting schemes and a delicate correction strategy . |
| Outcome: | The proposed auxiliary task improves Chinese pronunciation prediction on three benchmarks. |
FlipGuard: Defending Preference Alignment against Update Regression with Constrained Optimization (2024.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in preference alignment have significantly improved Large Language Models' ability to generate texts that align with human preferences and values. |
| Approach: | They propose a constrained optimization approach to detect and mitigate update regression with focal attention. |
| Outcome: | The proposed approach detects and mitigates update regression with focal attention while maintaining excellent overall performance. |
Improve Safety Training of Large Language Models with Safety-Critical Singular Vectors Localization (2025.acl-long)
Copied to clipboard
| Challenge: | Recent work on safety training with modules such as low-rank adaptation (LoRA) to resist jailbreaks shows promise, but these approaches can inadvertently degrade a model’s general utility. |
| Approach: | They propose a plug-and-play method that locates safety-critical singular vectors within the model's parameter space and a dynamic rank number determination strategy to reduce parameter overhead. |
| Outcome: | The proposed method mitigates the impact of safety training on model utility by explicitly locating and leveraging safety-critical singular vectors within the model’s parameter space. |
Air-Decoding: Attribute Distribution Reconstruction for Decoding-Time Controllable Text Generation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Controllable text generation (CTG) aims to generate text with desired attributes, but current methods lack high levels of controllability. |
| Approach: | They propose a lightweight decoding framework that reconstructs attribute distributions to balance the weights between attribute words and non-attribute words to generate more fluent text. |
| Outcome: | The proposed framework achieves state-of-the-art control performance on multiple CTG tasks. |
Chain-of-Question: A Progressive Question Decomposition Approach for Complex Knowledge Base Question Answering (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to answer complex questions rely on decomposition of complex questions into sub-questions . Existing approaches to decompose complex questions are limited by the original question . |
| Approach: | They propose a question decomposition approach to decompose semantically clear questions . they use the decomposed sub-questions to select relevant patterns as auxiliary information . |
| Outcome: | The proposed method achieves state-of-the-art performance on multiple datasets. |
IAEval: A Comprehensive Evaluation of Instance Attribution on Natural Language Understanding (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Instance attribution (IA) aims to identify the training instances leading to the prediction of a test example. |
| Approach: | They propose a systematic and comprehensive evaluation scheme covering four significant requirements: sufficiency, completeness, stability and plausibility. |
| Outcome: | The proposed evaluation scheme covers four significant requirements: sufficiency, completeness, stability and plausibility. |