Papers with hallucinations

8 papers
ALOHa: A New Measure for Hallucination in Captioning Models (2024.naacl-short)

Copied to clipboard

Challenge: Existing metric for object hallucination, CHAIR, is limited to MS COCO objects and synonyms.
Approach: They propose a new open-vocabulary metric, ALOHa, which leverages large language models to measure object hallucinations.
Outcome: The proposed metric correctly identifies 13.6% more hallucinated objects than CHAIR on HAT and 30.8% more on nocaps.
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have a tendency to hallucinate false or misleading information, limiting their reliability.
Approach: They examine how architecture-based inductive biases affect the propensity to hallucinate . they find that the models are more reliable and more reliable than traditional models .
Outcome: The proposed models can be used to train and train large language models that are factual or able to explain themselves through their knowledge.
Before Generation, Align it! A Novel and Effective Strategy for Mitigating Hallucinations in Text-to-SQL Generation (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) driven by In-Context Learning (ICL) have improved performance of text-to-SQL.
Approach: They propose a strategy to mitigate hallucinations in large language models driven by In-Context Learning (ICL) they propose TA-SQL, a text-to-Sql framework that encourages LLMs to take advantage of similar tasks rather than starting from scratch.
Outcome: The proposed framework improves the performance of the GPT-4 model by 21.23% on BIRD dev.
On Exposure Bias, Hallucination and Domain Shift in Neural Machine Translation (2020.acl-main)

Copied to clipboard

Challenge: Neural machine translation suffers from exposure bias, and alternative approaches to mitigate this are under debate.
Approach: They propose to reduce exposure bias by using minimum risk training to mitigate hallucinations . they find that exposure bias is more problematic under domain shift .
Outcome: The proposed methods can reduce exposure bias even on in-domain test sets.
K-COMP: Retrieval-Augmented Medical Domain Question Answering With Knowledge-Injected Compressor (2025.naacl-long)

Copied to clipboard

Challenge: Documents retrieved for closed domains require high expertise, so reader model may have difficulty comprehending the text.
Approach: They propose a system which augments the prior knowledge required to answer correctly by adding thousands of tokens to the retrieved documents.
Outcome: The proposed system provides the knowledge required to answer correctly and generates prior knowledge to facilitate the answer process prior to compression of the retrieved passages.
INVITE: a Testbed of Automatically Generated Invalid Questions to Evaluate Large Language Models for Hallucinations (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in Large language models have enabled them to hold free form conversations over multiple turns, but they exhibit a tendency to make unfounded and incorrect statements, commonly labeled as hallucinations.
Approach: They propose a framework to test large language models for hallucinations using automatically generated INValId questions.
Outcome: The proposed framework is based on a testbed of automatically generated INValId questions to evaluate large language models for hallucinations.
Mutual Information Alleviates Hallucinations in Abstractive Summarization (2022.emnlp-main)

Copied to clipboard

Challenge: Abstractive summarization models exhibit the tendency to hallucinate, i.e., output content not supported by the source document.
Approach: They propose a decoding strategy that optimizes for pointwise mutual information of source and target tokens when models exhibit uncertainty.
Outcome: The proposed method decreases the probability of hallucinated tokens while maintaining the Rouge and BERT-S scores of top-performing decoding strategies.
Improving Faithfulness of Large Language Models in Summarization via Sliding Generation and Self-Consistency (2024.lrec-main)

Copied to clipboard

Challenge: Abstractive summarization models (LLMs) have demonstrated impressive performance in various tasks, but they are still suffering from factual inconsistency problem called hallucination.
Approach: They propose to improve the faithfulness of large language models by impelling them to process the entire article more fairly and faithfully.
Outcome: The proposed strategy improves the faithfulness of large language models in summarization while maintaining their fluency and informativeness.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations