Papers by Cheng Niu

22 papers
UHGEval: Benchmarking the Hallucination of Chinese Large Language Models via Unconstrained Generation (2024.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) produce hallucinated text, compromising their practical utility in professional contexts.
Approach: They have developed an unconstrained hallucination generation evaluation benchmark that contains hallucines generated by large language models with minimal restrictions.
Outcome: The proposed benchmarks are based on a Chinese-language dataset that is lacking in the field.
Rhetorically Controlled Encoder-Decoder for Modern Chinese Poetry Generation (P19-1)

Copied to clipboard

Challenge: Rhetoric is a vital element in modern Chinese poetry, and plays an essential role in improving its aesthetics. however, to date, it has not been considered in research on automatic poetry generation.
Approach: They propose a rhetorically controlled encoder-decoder for modern Chinese poetry generation . their model captures various rhetorical patterns in an encoder and incorporates mixtures .
Outcome: The proposed model outperforms state-of-the-art methods in terms of fluency, coherence, meaningfulness, and rhetorical aesthetics.
Data Efficient RLVR via Off-Policy Influence Guidance (2026.acl-long)

Copied to clipboard

Challenge: Existing data selection methods for RLVR are heuristic-based, lacking theoretical guarantees and generalizability.
Approach: They propose an off-policy influence estimation method that approximates data influence using offline trajectories.
Outcome: The proposed method reduces the computational cost of policy rollouts and improves storage and computation efficiency.
Improving Multi-turn Dialogue Modelling with Utterance ReWriter (P19-1)

Copied to clipboard

Challenge: Recent research has achieved impressive results in single-turn dialogue modelling, but multi-turn models still remain challenging.
Approach: They propose to rewrite human utterances as a pre-process to help multi-turn dialgoue modelling.
Outcome: The proposed architecture achieves remarkably good performance on the utterance rewriting task.
Contrastive Zero-Shot Learning for Cross-Domain Slot Filling with Adversarial Attack (2020.coling-main)

Copied to clipboard

Challenge: Existing approaches to zero-shot slot filling ignore constraints in the latent space and lack robustness.
Approach: They propose a Contrastive Zero-Shot Learning with Adversarial Attack method for slot filling . they propose to map slot value contextual representations to slot description representations .
Outcome: The proposed method outperforms state-of-the-art models under zero-shot and few-shot settings.
Incremental Transformer with Deliberation Decoder for Document Grounded Conversations (P19-1)

Copied to clipboard

Challenge: Existing dialogue systems do not exploit document knowledge effectively enough.
Approach: They propose a Transformer-based architecture for document grounded conversations that incorporates document knowledge into a two-pass decoder to improve context coherence and knowledge correctness.
Outcome: The proposed model outperforms baselines on context coherence and knowledge relevance on a real-world document grounded dataset.
Neural Data-to-Text Generation via Jointly Learning the Segmentation and Correspondence (2020.acl-main)

Copied to clipboard

Challenge: Recent neural attention models conflate all steps into a single end-to-end system and simplify training process.
Approach: They propose to explicitly segment target text into fragment units and align them with their data correspondences.
Outcome: The proposed model outperforms neural attention models on E2E and WebNLG benchmarks.
OpenGenAlign: A Preference Dataset and Benchmark for Trustworthy Reward Modeling in Open-Ended, Long-Context Generation (2026.findings-acl)

Copied to clipboard

Challenge: Existing reward models perform suboptimal on held-out benchmarks, resulting in poor quality outputs.
Approach: They propose a framework and a high-quality dataset to evaluate reward models . they define four key metrics to assess generation quality and develop a pipeline to evaluate outputs .
Outcome: The proposed framework and dataset improves hallucination-free, comprehensive, reliable, and efficient open-ended long-context generation.
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models (2024.acl-long)

Copied to clipboard

Challenge: Retrieval-augmented generation (RAG) is a main technique for alleviating hallucinations in large language models.
Approach: They propose to integrate RAG into large language models to analyze word-level hallucinations using a corpus of 18,000 naturally generated responses from diverse LLMs.
Outcome: The proposed model can fine tune a relatively small LLM and achieve a competitive hallucination detection performance when compared to the existing prompt-based approaches.
LiveFact: A Dynamic, Time-Aware Benchmark for LLM-Driven Fake News Detection (2026.acl-long)

Copied to clipboard

Challenge: Current evaluation frameworks are static and vulnerable to benchmark data contamination . current models are ineffective at assessing reasoning under temporal uncertainty .
Approach: They propose a live-based benchmark that simulates the real-world "fog of war" they propose evaluating models on their ability to reason with evolving, incomplete information .
Outcome: The proposed model outperforms proprietary state-of-the-art models in classification and evidence mode . it also provides a component to monitor BDC explicitly .
Tackling Distractor Documents in Multi-Hop QA with Reinforcement and Curriculum Learning (2026.findings-eacl)

Copied to clipboard

Challenge: Existing work on retrieval-augmented generation systems has shown that retrievers exhibit imperfect recall and precision, limiting downstream performance.
Approach: They propose a retrieval-augmented generation model that generates answers from larger sets of retrieved contexts.
Outcome: The proposed model generates answers and cites relevant information from larger sets of retrieved contexts.
MovieChats: Chat like Humans in a Closed Domain (2020.emnlp-main)

Copied to clipboard

Challenge: Currently, open-domain chatbots are far from satisfactory.
Approach: They propose a unified, readily scalable neural approach which reconciles all subtasks like intent prediction and knowledge retrieval.
Outcome: The proposed approach outperforms commercial systems replying on complex rules on static and interactive tests and shows that the results are remarkably good.
Enhancing Dialogue State Tracking Models through LLM-backed User-Agents Simulation (2024.acl-long)

Copied to clipboard

Challenge: Experimental results show that the model can be used to generate dialogues in new domains quickly.
Approach: They propose to use LLMs to generate dialogue data to reduce dialogue collection and annotation costs.
Outcome: The proposed model performs better than the baseline model trained on real data.
StableToolBench-MirrorAPI: Modeling Tool Environments as Mirrors of 7,000+ Real-World APIs (2025.findings-acl)

Copied to clipboard

Challenge: Existing tool environments face challenges in balancing stability, scale, and realism, especially for benchmarking purposes.
Approach: They propose a framework that trains specialized LLMs to accurately simulate real API responses by supervised fine-tuning and chain-of-thought reasoning.
Outcome: The proposed framework achieves superior accuracy and stability compared to state-of-the-art methods on the newly constructed MirrorAPI-Bench and its integration into StableToolBench.
Diversifying Dialogue Generation with Non-Conversational Text (2020.acl-main)

Copied to clipboard

Challenge: Neural network-based sequence-to-sequence models suffer from low diversity in open-domain dialogue generation.
Approach: They propose a way to diversify dialogue generation by leveraging non-conversational text . they collect large-scale corpus from forum comments, idioms and book snippets .
Outcome: The proposed model produces significantly more diverse responses without sacrificing relevance with context.
Contextual Relevance and Adaptive Sampling for LLM-Based Document Reranking (2026.acl-long)

Copied to clipboard

Challenge: identifying relevant documents for Reasoning-intensive queries remains a challenge . large language models have shown strong performance in zero-shot document reranking .
Approach: They propose a reranking algorithm that estimates contextual relevance by aggregating LLMs' relevance judgments across batches.
Outcome: The proposed algorithm improves nDCG@10 over retrieval and reranking baselines by 15% and 6–21% respectively.
RAG-HAT: A Hallucination-Aware Tuning Pipeline for LLM in Retrieval-Augmented Generation (2024.emnlp-industry)

Copied to clipboard

Challenge: Retrieval-augmented generation (RAG) has emerged as a significant advancement in the field of large language models (LLMs).
Approach: They propose a method that uses hallucination detection labels to correct hallucines by integrating up-to-date information into their initial training.
Outcome: The proposed method is based on the Retrieval Augmented Generation (RAG) method, which has shown to be effective in mitigating hallucinations and improving answer quality.
Machine Unlearning of Pre-trained Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Using curated datasets, we establish a robust benchmark for unlearning performance, demonstrating that these methods are over 105 times more computationally efficient than retraining.
Approach: They propose a framework for machine unlearning in pre-trained LLMs and integrate gradient ascent with gradient descent on in-distribution data to achieve robustness.
Outcome: The proposed framework is over 105 times more efficient than retraining on in-distribution data and provides detailed guidelines for efficient hyperparameter tuning in the unlearning process.
VeraCT Scan: Retrieval-Augmented Fake News Detection with Justifiable Reasoning (2024.acl-demos)

Copied to clipboard

Challenge: generative artificial intelligence has exacerbated the challenge of distinguishing genuine news from fabricated stories.
Approach: They propose a retrieval-augmented system that extracts the core facts from a given piece of news and conducts an internet-wide search to identify corroborating or conflicting reports.
Outcome: The proposed system has demonstrated state-of-the-art accuracy in the realm of fake news detection.
A Contextual Hierarchical Attention Network with Adaptive Objective for Dialogue State Tracking (2020.acl-main)

Copied to clipboard

Challenge: Existing methods for dialogue state tracking ignore the slot imbalance problem and treat all slots indiscriminately, which limits the learning of hard slots.
Approach: They propose to employ a contextual hierarchical attention network to enhance the DST by learning contextual representations.
Outcome: The proposed approach achieves 52.68% and 58.55% joint accuracy on multiWOZ 2.0 and MultiWOZ 2.1 datasets and significantly improves performance (+1.24% and +5.98%)
Evaluating the Expressive Appropriateness of Speech in Rich Contexts (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for evaluating expressive speech focus on word accuracy, naturalness, signal quality, or emotional intensity at the utterance level.
Approach: They propose a framework for Evaluating Expressive Appropriateness in speech that assesses whether a speech sample aligns with the underlying communicative intent implied by its discourse-level narrative context.
Outcome: The proposed framework outperforms existing speech evaluation and analysis systems on a human-annotated test set.
Answer-Supervised Question Reformulation for Enhancing Conversational Machine Comprehension (D19-58)

Copied to clipboard

Challenge: Existing question reformulation models are based on supervised question labels without considering feedback information from answers.
Approach: They propose a question reformulation model that integrates conversational history information with reinforcement learning.
Outcome: The proposed model is more effective in conversational machine comprehension with reinforcement learning.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations