Papers by Guan Huang
LOT: A Story-Centric Benchmark for Evaluating Chinese Long Text Understanding and Generation (2022.tacl-1)
Copied to clipboard
| Challenge: | Existing benchmarks for natural language processing focus on understanding or generating short texts . lack of standardized benchmarks makes it difficult to assess and compare models . |
| Approach: | They propose a story-centric benchmark for Chinese long text modeling that aggregates two understanding tasks and two generation tasks. |
| Outcome: | The proposed model outperforms similar-sized models on understanding and generation tasks. |
Automatic Comment Generation for Chinese Student Narrative Essays (2022.emnlp-demos)
Copied to clipboard
| Challenge: | Existing studies focus on giving discrete scores for holistic quality or distinct traits, but real-world teachers usually provide detailed comments in natural language, which are more informative than single scores. |
| Approach: | They propose a model which generates comments for specified segments from given student narrative essays using a human-written Chinese dataset. |
| Outcome: | The proposed model outperforms baselines and has 91% success rate. |
AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models (2024.findings-emnlp)
Copied to clipboard
Xiyang Wu, Tianrui Guan, Dianqi Li, Shuaiyi Huang, Xiaoyu Liu, Xijun Wang, Ruiqi Xian, Abhinav Shrivastava, Furong Huang, Jordan Boyd-Graber, Tianyi Zhou, Dinesh Manocha
| Challenge: | Large vision-language models are prone to hallucinations, where contextual cues in an image can trigger the language module to produce overconfident and incorrect reasoning about abnormal or hypothetical objects. |
| Approach: | They propose to automate the generation of hallucination-related questions using images . they propose to use three image manipulation strategies to induce hallucinosity . |
| Outcome: | The proposed approach reduces human bias in crafting such examples and improves accuracy. |
Generating Informative Responses with Controlled Sentence Function (P18-1)
Copied to clipboard
| Challenge: | Sentence function is a significant factor to achieve the purpose of the speaker, but has not been touched in large-scale conversation generation. |
| Approach: | They propose a model to generate informative responses with controlled sentence function using a latent variable and a type controller to deal with compatibility. |
| Outcome: | The proposed model outperforms state-of-the-art models and generates responses with controlled sentence function and informative content. |
Mixture-of-Modules: Reinventing Transformers as Dynamic Assemblies of Modules (2024.emnlp-main)
Copied to clipboard
| Challenge: | Empirical results show that MoMs consistently outperform vanilla transformers . |
| Approach: | They propose an architecture that allows for a mixture-of-modules computation that uses a finite set of modules defined by multi-head attention and feed-forward networks. |
| Outcome: | The proposed architecture outperforms vanilla Transformers and their variants in multiple ways. |
MAGRET: Machine-generated Text Detection with Rewritten Texts (2025.coling-main)
Copied to clipboard
| Challenge: | Existing studies focus on detecting machine-generated text in open-source models, but their performance on closed-source large models is limited. |
| Approach: | They propose a method to detect rewritten text from large language models using a BERT encoder and propose to refine it to achieve semantic alignment. |
| Outcome: | The proposed method outperforms baseline methods on three text-generated datasets. |
EOP-LLM: Energy Oriented Pruning for Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Inference energy consumption has grown rapidly in large language models (LLMs) but existing methods focus on reducing FLOPs or latency rather than modeling or enforcing end-to-end inference energy constraints. |
| Approach: | They propose an energy-oriented dynamic pruning framework that enables LLM inference under explicit per-sequence energy budgets. |
| Outcome: | EOP-LLM outperforms state-of-the-art dynamic pruning baselines while adhering to per-sequence energy constraints. |
UNION: An Unreferenced Metric for Evaluating Open-ended Story Generation (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing metrics for open-ended text generation are poorly correlated with human judgments . despite the success of existing metrics, there are few plausible outputs for the same input . |
| Approach: | They propose a UNreferenced measure for evaluating open-ended story generation . it is built on top of BERT and is trained to distinguish human-written stories from negative samples . |
| Outcome: | The proposed measure is more generalizable than state-of-the-art metrics and correlates better with human judgments. |
Multi-Stage LLM Fine-Tuning with a Continual Learning Setting (2025.findings-naacl)
Copied to clipboard
Changhao Guan, Chao Huang, Hongliang Li, You Li, Ning Cheng, Zihe Liu, Yufeng Chen, Jinan Xu, Jian Liu
| Challenge: | Large language models (LLMs) have made significant progress in knowledge-intensive applications, but they may face a multi-stage continuous learning scenario. |
| Approach: | They propose a multi-stage continuous learning paradigm that includes a preference-based learning bias to identify potential knowledge conflicts and a self-distillation-based data augmentation strategy to expand and enrich the training corpus. |
| Outcome: | The proposed learning paradigm achieves a significant improvement in accuracy after 7 stages of fine-tuning compared to previous methods while preserving general knowledge. |
E2EDev: Benchmarking Large Language Models in End-to-End Software Development Task (2026.acl-long)
Copied to clipboard
| Challenge: | Existing E2ESD benchmarks are limited by coarse-grained requirement specifications and unreliable evaluation protocols. |
| Approach: | They propose a benchmark to assess whether generated software meets user needs . they use a fine-grained set of user requirements and a fully automated testing pipeline . |
| Outcome: | E2EDev is a benchmark to assess whether generated software meets user needs through mimicking real user interactions. |
Large Language Models and Causal Inference in Collaboration: A Comprehensive Survey (2025.findings-naacl)
Copied to clipboard
Xiaoyu Liu, Paiheng Xu, Junda Wu, Jiaxin Yuan, Yifan Yang, Yuhang Zhou, Fuxiao Liu, Tianrui Guan, Haoliang Wang, Tong Yu, Julian McAuley, Wei Ai, Furong Huang
| Challenge: | Large Language Models (LLMs) have shown great potential to enhance Natural Language Processing (NLP) models in areas such as predictive accuracy, fairness, robustness, and explainability. |
| Approach: | They evaluate or improve generative Large Language Models from a causal perspective in areas such as reasoning capacity, fairness and safety issues, explainability, and handling multimodality. |
| Outcome: | The proposed models can be used to perform causal relationship discovery and causal effect estimation tasks. |
Evidence-Augmented Policy Optimization with Reward Co-Evolution for Long-Context Reasoning (2026.acl-long)
Copied to clipboard
| Challenge: | Evidence-Augmented Policy Optimization (EAPO) improves long-context reasoning performance . Xu et al., 2025): large language models are a critical part of NLP . |
| Approach: | They propose an Evidence-Augmented Reasoning paradigm that uses a group-relative reward to improve evidence quality. |
| Outcome: | EAPO significantly improves long-context reasoning performance compared to baselines. |
Self-Supervised Sentence Polishing by Adding Engaging Modifiers (2023.acl-demo)
Copied to clipboard
| Challenge: | a typical way to polish sentences is to add engaging modifiers, which enhance the meaning of the sentence. |
| Approach: | They propose a task that requires polishing sentences while maintaining fluency . they remove engaging modifiers from public resources and fine-tune LongLM to reconstruct original sentences from corrupted ones. |
| Outcome: | The proposed model generates more engaging sentences with suitable modifiers than strong baselines while keeping fluency. |
ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent controllable zero-shot text-to-speech systems can synthesize speech for unseen speakers from a short reference audio clip, but they also inherit the speaking style present in the reference. |
| Approach: | They propose a framework that enables continuous and reference-relative style control in zero-shot text-to-speech systems by combining style-specific LoRAs with Orthogonal LoRA Fusion. |
| Outcome: | The proposed framework reduces the model's dependence on reference style while preserving text fidelity while maintaining intelligibility and speaker timbre. |
Language Models Hallucinate, but May Excel at Fact Verification (2024.naacl-long)
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have produced non-factual outputs . however, current LLMs suffer from the hallucination issue . |
| Approach: | They propose to use instruction-tuned LLMs to generate factual outputs . they find that FLAN-T5-11B performs best as a fact verifier . |
| Outcome: | The proposed method outperforms more capable LLMs like GPT3.5 and ChatGPT in the human evaluation. |
Re3Dial: Retrieve, Reorganize and Rescale Conversations for Long-Turn Open-Domain Dialogue Pre-training (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing pre-training models lack long-turn dialogue sessions due to the scarcity of long-term sessions. |
| Approach: | They propose a framework that can automatically construct billion-scale long-turn dialogues by reorganizing existing short-turn ones. |
| Outcome: | The proposed framework can automatically construct billion-scale long-turn dialogues by reorganizing existing short-turn ones. |
Cross-lingual Multimodal Sentiment Analysis for Low-Resource Languages via Language Family Disentanglement and Rethinking Transfer (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing multimodal sentiment analysis methods are limited to textual data and cannot handle multimodal scenarios. |
| Approach: | They propose a transfer learning framework that allows cross-lingual and cross-modal alignments and a language family disentanglement module that enhances the sharing of language universals within families. |
| Outcome: | The proposed method is superior to existing methods and can handle low-resource languages. |
SAT: Balancing Reasoning Accuracy and Efficiency with Stepwise Adaptive Thinking (2026.acl-long)
Copied to clipboard
| Challenge: | Large Reasoning Models (LRMs) produce excessively long Chains of Thought (COT) Existing solutions that improve token efficiency but sacrifice fine-grained control can disrupt the logical integrity of the reasoning process. |
| Approach: | They propose a framework that performs step-level, difficulty-aware pruning while preserving the core reasoning structure. |
| Outcome: | Experiments show that SAT reduces reasoning tokens by 40% while maintaining or improving accuracy. |
LIST: Linearly Incremental SQL Translator for Single-Hop Reasoning, Generation and Verification (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing schema linking methods are not able to handle complex SQL queries. |
| Approach: | They propose a new algorithm that transforms SQL queries into grammatically verifiable sub-queries which are arranged sequentially to reflect single-hop reasoning steps. |
| Outcome: | The proposed algorithm achieves significant performance gains on the BIRD dataset and surpasses schema linking methods at comparable or better cost. |
Biomedical Question Answering via Multi-Level Summarization on a Local Knowledge Graph (2026.acl-long)
Copied to clipboard
| Challenge: | Existing work on how to effectively capture multi-document relationships remains an open question . Existing techniques to mitigate this problem include hierarchical summarization of semantically related chunks or integrating Knowledge Graphs (KGs). |
| Approach: | They propose a method which constructs a local knowledge graph from retrieved documents . they use propositional claims to construct a knowledge graph and contextualize a small language model . |
| Outcome: | The proposed method outperforms RAG on biomedical benchmarks and is generalizable and effective. |
Boosting Data Utilization for Multilingual Dense Retrieval (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies focus on fine-tuning multilingual dense retrieval models, but data scarcity for low-resource languages makes it difficult to align representations in a shared vector space. |
| Approach: | They propose to obtain high-quality hard negative samples and effective mini-batch data to boost data utilization for multilingual dense retrieval by obtaining high-quality negative samples. |
| Outcome: | The proposed method outperforms existing baselines on a multilingual retrieval benchmark, MIRACL, with 16 languages. |
A Knowledge-Enhanced Pretraining Model for Commonsense Story Generation (2020.tacl-1)
Copied to clipboard
| Challenge: | Existing models for story generation suffer from repetition, logic conflicts, and lack of long-range coherence . |
| Approach: | They propose to utilize commonsense knowledge from external knowledge bases to generate reasonable stories by multi-task learning. |
| Outcome: | The proposed model can generate more reasonable stories than state-of-the-art models, compared with existing models, showing that it can capture useful semantic and syntactic features. |
A Corpus for Understanding and Generating Moral Stories (2022.naacl-main)
Copied to clipboard
| Challenge: | Existing tasks for evaluating story understanding and generation focus on reasoning plots from context, but they focus on bridging plots with implied morals. |
| Approach: | They propose two understanding tasks and two generation tasks to assess machines' ability to bridge story plots and implied morals. |
| Outcome: | The proposed tasks are based on a dataset of Chinese and English moral stories . they show that the proposed models can perform better than existing models . |
Stylized Story Generation with Style-Guided Planning (2021.findings-acl)
Copied to clipboard
| Challenge: | Current storytelling systems focus more on generating stories with coherent plots regardless of the narration style. |
| Approach: | They propose a novel task, stylized story generation, that first plans stylized keywords and then generates the whole story with the guidance of the keywords. |
| Outcome: | The proposed model can generate emotion-driven or event-driven stories based on the ROCStories dataset . |
Multilingual Collaborative Defense for Large Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing safeguards for Large Language Models are vulnerable to "jailbreaking" harmful queries. |
| Approach: | They propose a learning method that optimizes a continuous soft safety prompt automatically to facilitate multilingual safeguarding of LLMs. |
| Outcome: | The proposed method outperforms previous approaches in multilingual jailbreak defense while exhibiting strong cross-lingual generalization. |
Improving Role-Oriented Dialogue Summarization with Interaction-Aware Contrastive Learning (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods for encoding dialogues do not capture interaction information between roles, thus ignore interaction-related key information. |
| Approach: | They propose a contrastive learning based interaction-aware model for the role-oriented dialogue summarization namely CIAM and use it to train the decoder to learn role-level interaction. |
| Outcome: | The proposed model captures interaction information between different roles and produces informative summaries on two public datasets. |
KVFKT: A New Horizon in Knowledge Tracing with Attention-Based Embedding and Forgetting Curve Integration (2025.coling-main)
Copied to clipboard
| Challenge: | Existing knowledge tracing models do not incorporate forgetting features to improve the learning and answering processes. |
| Approach: | They propose a new approach in knowledge tracing with attention-based embedding and forgetting curve integration using four real-world datasets to test the model. |
| Outcome: | The proposed model outperforms the existing knowledge tracing models and eliminates the need for artificial engineering features. |
DoRA: Enhancing Parameter-Efficient Fine-Tuning with Dynamic Rank Distribution (2024.acl-long)
Copied to clipboard
| Challenge: | Existing parameter-efficient fine-tuning methods such as Low-Rank Adaptation ignore the differential parameter budget requirements across weight matrices, which may lead to suboptimal fine-uning outcomes. |
| Approach: | They propose a parameter-efficient low-rank Adaptation method that decomposes high-rank LoRA layers into structured single-rank components and allows dynamic pruning of parameter budget . |
| Outcome: | The proposed method outperforms LoRA and LoRA with the same parameter budget and performance. |
Mitigating the Learning Bias towards Repetition by Self-Contrastive Training for Open-Ended Generation (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing language models generate repetitive texts with greedy decoding or beam search. |
| Approach: | They propose a self-contrastive training technique to penalize the output of a premature checkpoint of the same model when it incorrectly predicts repetition. |
| Outcome: | The proposed training mitigates repetition while maintaining fluency while minimizing the overestimation of token-level repetition probabilities. |
Persona-Guided Planning for Controlling the Protagonist’s Persona in Story Generation (2022.naacl-main)
Copied to clipboard
| Challenge: | Existing methods to control the protagonist's persona in story generation are implicitly and sparsely embodied in stories, so we propose a planning-based generation model called ConPer to explicitly model the relationship between personas and events. |
| Approach: | They propose a model to control the protagonist's persona in story generation by predicting one target sentence and planning the plot as a sequence of keywords with the guidance of the predicted persona-related events and commonsense knowledge. |
| Outcome: | The proposed model outperforms state-of-the-art models for generating more coherent and persona-controllable stories. |
When and Who? Conversation Transition Based on Bot-Agent Symbiosis Learning Network (2020.coling-main)
Copied to clipboard
| Challenge: | a bot-agent symbiosis is a method for transparent conversation transition in online customer service applications. |
| Approach: | They propose a bot-agent symbiosis approach to solve conversation transition problems . they provide user feedback and develop deep neural networks to predict the NPS . |
| Outcome: | The proposed approach outperforms state-of-the-art methods on real-time data generated from an online service support platform. |
OAgents: An Empirical Study of Building Effective Agents (2025.findings-emnlp)
Copied to clipboard
He Zhu, Tianrui Qin, King Zhu, Heyuan Huang, Yeyi Guan, Jinxiang Xia, Hanhao Li, Yi Yao, Ningning Wang, Pai Liu, Tianhao Peng, Xin Gui, Li Xiaowan, Yuhui Liu, Xiangru Tang, Jian Yang, Ge Zhang, Xitong Gao, Yuchen Eleanor Jiang, Changwang Zhang, Jun Wang, Jiaheng Liu, Wangchunshu Zhou
| Challenge: | a recent study shows that agent research practices are far from standard, rigorous . lack of a standard evaluation protocol makes previous works not reproducible, authors say . |
| Approach: | They conduct an empirical study on the GAIA benchmark to investigate agent design choices . they find that lack of a standard evaluation protocol makes previous works not reproducible . |
| Outcome: | The proposed framework achieves state-of-the-art performance among open-source projects. |
OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation Metrics (2021.acl-long)
Copied to clipboard
Jian Guan, Zhexin Zhang, Zhuoer Feng, Zitao Liu, Wenbiao Ding, Xiaoxi Mao, Changjie Fan, Minlie Huang
| Challenge: | Existing automatic metrics are observed to correlate poorly with human evaluation. |
| Approach: | They propose to use OpenMEVA to evaluate open-ended story generation metrics. |
| Outcome: | The proposed test suite assesses the capabilities of open-ended story generation metrics on annotated stories and auto-constructed test examples. |
StoryTrans: Non-Parallel Story Author-Style Transfer with Discourse Representations and Content Enhancing (2023.acl-long)
Copied to clipboard
| Challenge: | Existing studies on text style transfer neglect long style transfer at the discourse level. |
| Approach: | They propose a model that transfers text style into target styles with learnable style embeddings . they use a mask-and-fill framework to explicitly fuse style-specific keywords into generation . |
| Outcome: | The proposed model outperforms baselines in style transfer and content preservation. |
Long Text Generation by Modeling Sentence-Level and Discourse-Level Coherence (2021.acl-long)
Copied to clipboard
| Challenge: | Existing generation models struggle to maintain a coherent event sequence throughout the generated text. |
| Approach: | They propose a long text generation model which can represent prefix sentences at sentence level and discourse level in the decoding process. |
| Outcome: | The proposed model can generate more coherent texts than state-of-the-art models. |