Papers by Hui Gao

23 papers
Safety Sidecar: Reflection-Driven Runtime Control for Safer Agents (2026.findings-acl)

Copied to clipboard

Challenge: Existing safety controls fail to provide runtime intervention or cross-architecture portability for autonomous LLM agents.
Approach: They propose a model-agnostic, plug-and-play module to provide arbitrary agent safety control and auditability.
Outcome: The proposed module improves the secure-solution rate by 2.9–11.2 percentage points . it adds only 3.2s to end-to-end latency and a negligible average cost of 5.37 10-4 per scenario .
From Adversarial Arms Race to Model-centric Evaluation: Motivating a Unified Automatic Robustness Evaluation Framework (2023.findings-acl)

Copied to clipboard

Challenge: Existing models of robustness evaluation are incomprehensive, impractical, and invalid .
Approach: They propose a framework for automatic robustness evaluation that shifts towards model-centric evaluation to further exploit the advantages of adversarial attacks.
Outcome: The proposed framework is based on a model-centric evaluation protocol and a robustness evaluation protocol.
Knowledge-Selective Pretraining for Attribute Value Extraction (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for AVE are limited on rare attributes due to poor generalization ability.
Approach: They propose to leverage pretraining and transfer learning to address weaknesses in existing methods.
Outcome: The proposed method achieves new state-of-the-art performance without pretraining on rare attributes with limited training resources.
LA-UCL: LLM-Augmented Unsupervised Contrastive Learning Framework for Few-Shot Text Classification (2024.lrec-main)

Copied to clipboard

Challenge: Experimental results show that our model exceeds the baseline models due to the lack of cognitive ability.
Approach: They propose a LLM-Augmented Unsupervised Contrastive Learning Framework which introduces a cognition-enabled Large Language Model (LLM) for efficient data augmentation and presents corresponding contrastive learning strategies.
Outcome: The proposed model exceeds baseline models on six datasets.
Error Comparison Optimization for Large Language Models on Aspect-Based Sentiment Analysis (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for aspect-based sentiment analysis (ABSA) only compare current predictions and labels on each sample, yet fail to perceive and understand its error outputs from different degrees.
Approach: They propose a framework that can perceive and understand the degree of errors by learning from comparative error pairs.
Outcome: The proposed framework exceeds baselines and achieves the desired performance.
ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection (2026.findings-acl)

Copied to clipboard

Challenge: Current mathematical benchmarks focus on evaluating MLLMs’ problem-solving ability, yet there is a crucial gap in addressing more complex scenarios such as error detection.
Approach: They propose to evaluate multimodal error detection by evaluating two sub-tasks error step identification and error categorization.
Outcome: The proposed task evaluates MLLMs' ability to handle multimodal questions compared to text-only models.
Rumor Detection on Social Media with Crowd Intelligence and ChatGPT-Assisted Networks (2023.emnlp-main)

Copied to clipboard

Challenge: Existing research on rumor detection challenges the expressive power of text encoding sequences, and insufficient mining of semantic structural information.
Approach: They propose a Crowd Intelligence-based semantic feature learning module to capture textual content’s sequential and hierarchical features and a knowledge-based structural mining module that leverages ChatGPT for knowledge enhancement.
Outcome: The proposed system achieves performance improvement in rumor detection tasks validating the effectiveness and rationality of using large language models as auxiliary tools.
Anti-Length Shift: Dynamic Outlier Truncation for Training Efficient Reasoning Models (2026.acl-long)

Copied to clipboard

Challenge: Existing efficient reasoning methods rely on explicit length penalties for excessive verbosity on simple queries.
Approach: They propose a training-time intervention that selectively suppresses redundant tokens . they find length shift occurs when models generate unnecessary reasoning on trivial inputs - a phenomenon that is often unexplored .
Outcome: The proposed method reduces inference token usage by 78% while increasing accuracy compared to the initial policy and surpasses state-of-the-art efficient reasoning methods.
EverMemOS: A Self-Organizing Memory Operating System for Structured Long-Horizon Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Existing memory systems for LLMs store isolated records and retrieve fragments . Existing systems store isolated data and fragments, limiting their ability to consolidate evolving experience and resolve conflicts.
Approach: They propose an engram-inspired memory operating system that implements an 'engram'-inspired lifecycle for computational memory.
Outcome: Experiments on LoCoMo, LongMemEval, and PersonaMeM-v2 show that EverMemeOS outperforms state-of-the-art methods on memory-augmented reasoning tasks.
Toward Safe and Human-Aligned Game Conversational Recommendation via Multi-Agent Decomposition (2026.findings-eacl)

Copied to clipboard

Challenge: Existing systems for conversational recommender systems (CRS) have strong results in movies, but games present distinct challenges . MATCHA framework provides specialized agents for intent parsing, tool-augmented retrieval, multi-LLM ranking, and stronger safety.
Approach: They propose a framework for conversational recommender systems that assigns specialized agents for intent parsing, tool-augmented retrieval, multi-LLM ranking and risk control.
Outcome: MATCHA outperforms baselines on real user request dataset, improves Hit@5 by 20%, reduces popularity bias by 24%, and achieves 97.9% adversarial defense.
Live-Aid: A Large-Scale Dialogue Dataset and Benchmark for Interleaved Multi-party Interactions in Live Streaming (2026.findings-acl)

Copied to clipboard

Challenge: Existing Multimodal Large Language Models struggle with dynamic interactions due to the scarcity of high-quality interleaved data.
Approach: They propose a large-scale interleaved live interaction Chinese dataset with human-annotated video responses.
Outcome: The proposed model can be used to evaluate live interactions in Chinese over 1,100 hours and 80,037 dialogue turns.
Iterative Constrained Back-Translation for Unsupervised Domain Adaptation of Machine Translation (2022.coling-1)

Copied to clipboard

Challenge: Existing back-translation methods focus on in-domain lexical knowledge, which may lead to poor translation of unseen in- domain words.
Approach: They propose an iterative constrained back-translation method to incorporate in-domain lexical knowledge into synthetic parallel data from BT.
Outcome: The proposed method improves the BLEU score by up to 3.08 on four domains.
SVD-GCL: A Noise-Augmented Hybrid Graph Contrastive Learning Framework for Recommendation (2025.coling-main)

Copied to clipboard

Challenge: Recent advances in graph neural networks have made it difficult to capture user preferences.
Approach: They propose a graph contrastive learning recommendation model based on noise augmentation that integrates truncated singular value decomposition in the feature engineering stage.
Outcome: The proposed model reduces dimensionality and denoises the original data.
Rethink Rumor Detection in the Era of LLMs: A Review (2025.findings-emnlp)

Copied to clipboard

Challenge: rumor detection has been reshaped by large language models (LLMs) this paper proposes a Cognition-Interaction-Behavior (CIB) framework for rumour detection based on collective intelligence .
Approach: They propose a Cognition-Interaction-Behavior framework for rumor detection based on collective intelligence and explore synergistic relationship between LLMs and collective intelligence in rumour governance.
Outcome: The proposed framework unifies existing methods and reveals synergistic relationship between LLMs and collective intelligence in rumor governance.
The Right Time Matters: Data Arrangement Affects Zero-Shot Generalization in Instruction Tuning (2025.findings-acl)

Copied to clipboard

Challenge: Existing work on instruction tuning has focused on task level, without considering that tasks are artificially defined and, to LLMs, merely consist of tokens and representations.
Approach: They propose a training data arrangement framework that allows for continual learning and loss reduction.
Outcome: The proposed framework promotes continual learning and loss reduction on unseen tasks.
A LSTM Approach with Sub-Word Embeddings for Mongolian Phrase Break Prediction (C18-1)

Copied to clipboard

Challenge: Existing word embedding methods for Mongolian PB prediction are expensive and time-consuming.
Approach: They propose to use Mongolian word embedding to build a robust Mongolian PB prediction model . they encode sub-word units and feed it to LSTM to decode the best corresponding PB label .
Outcome: The proposed model outperforms traditional model using manual features and achieves 7.49% gain.
Persona-E²: A Human-Grounded Dataset for Personality-Shaped Emotional Responses to Textual Events (2026.acl-long)

Copied to clipboard

Challenge: A critical bottleneck is the lack of ground-truth human data to link personality traits to emotional shifts.
Approach: They propose a large-scale dataset to capture reader-based emotional variations across news, social media, and life narratives.
Outcome: The proposed model captures reader-based emotional variations across news, social media, and life narratives.
Deciphering Rumors: A Multi-Task Learning Approach with Intent-aware Hierarchical Contrastive Learning (2024.emnlp-main)

Copied to clipboard

Challenge: Social networks are rife with noise and misleading information, presenting multifaceted challenges for rumor detection.
Approach: They propose a new multi-task learning framework that mines latent intentions and rumor semantic features . they propose to use event-level and intent-level strategies to establish cognitive anchors .
Outcome: The proposed framework improves the effectiveness of rumor detection and addresses the challenges present in the field.
LLM-SLM Collaborative Framework of Idiomatic Expression Generation (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for idiomatic expression generation lack parallel data and manual annotations.
Approach: They propose an iterative LLM-SLM collaborative framework that replaces human supervision for idiomatic expression data generation.
Outcome: The proposed framework outperforms DeepSeek-R1 in Chinese Idiom Polishing with a 25.2% improvement in accuracy.
EcomScriptBench: A Multi-task Benchmark for E-commerce Script Planning via Step-wise Intention-Driven Product Association (2025.acl-long)

Copied to clipboard

Challenge: Goal-oriented script planning is used by humans to plan for typical activities . however, this capability remains underexplored due to several challenges .
Approach: They propose a framework that enables product-enriched scripts by associating products with each step based on the semantic similarity between the actions and their purchase intentions.
Outcome: The proposed framework can generate product-enriched scripts from 2.4 million scripts . human annotations are conducted to provide gold labels for a sampled subset .
Consistency Rating of Semantic Transparency: an Evaluation Method for Metaphor Competence in Idiom Understanding Tasks (2025.coling-main)

Copied to clipboard

Challenge: Idioms condense complex semantics into fixed phrases, making idiom comprehension a test of metaphor competence.
Approach: They propose a method to evaluate the metaphor competence of LLMs for the idiom understanding task: the Consistency Rating of Semantic Transparency (CR-ST).
Outcome: The proposed method assesses the difficulty of understanding idioms through two dimensions: overall semantic transparency and constituent semantic transparency, aiming to gauge LLMs’ mastery of metaphor competence.
Distilling Causal Effect of Data in Continual Few-shot Relation Learning (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for learning relational patterns from data are prone to catastrophic forgetting issues due to limited number of samples and continual training mode.
Approach: They propose a unified causal framework for CFRL to restore causal effects from old data . they establish two additional causal paths from old to predictions by colliding with old data separately in the old feature space.
Outcome: The proposed method is superior to existing state-of-the-art methods in CFRL task settings.
ContrastKV: Robust KV Cache Eviction via Contrastive Signal Fusion for Multi-Query Generalization (2026.acl-long)

Copied to clipboard

Challenge: Existing query-agnostic approaches rely on a single proxy query, leading to fragile eviction decisions under high evict ratios.
Approach: They propose a query-agnostic KV cache eviction algorithm that exploits complementary semantic and non-semantic signals.
Outcome: Experiments show that the proposed algorithm outperforms state-of-the-art methods while retaining up to 92% accuracy with only 20% of the KV cache budget.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations