Papers with hallucinations
ALOHa: A New Measure for Hallucination in Captioning Models (2024.naacl-short)
Copied to clipboard
Suzanne Petryk, David Chan, Anish Kachinthaya, Haodi Zou, John Canny, Joseph Gonzalez, Trevor Darrell
| Challenge: | Existing metric for object hallucination, CHAIR, is limited to MS COCO objects and synonyms. |
| Approach: | They propose a new open-vocabulary metric, ALOHa, which leverages large language models to measure object hallucinations. |
| Outcome: | The proposed metric correctly identifies 13.6% more hallucinated objects than CHAIR on HAT and 30.8% more on nocaps. |
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination (2025.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) have a tendency to hallucinate false or misleading information, limiting their reliability. |
| Approach: | They examine how architecture-based inductive biases affect the propensity to hallucinate . they find that the models are more reliable and more reliable than traditional models . |
| Outcome: | The proposed models can be used to train and train large language models that are factual or able to explain themselves through their knowledge. |
Before Generation, Align it! A Novel and Effective Strategy for Mitigating Hallucinations in Text-to-SQL Generation (2024.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) driven by In-Context Learning (ICL) have improved performance of text-to-SQL. |
| Approach: | They propose a strategy to mitigate hallucinations in large language models driven by In-Context Learning (ICL) they propose TA-SQL, a text-to-Sql framework that encourages LLMs to take advantage of similar tasks rather than starting from scratch. |
| Outcome: | The proposed framework improves the performance of the GPT-4 model by 21.23% on BIRD dev. |
On Exposure Bias, Hallucination and Domain Shift in Neural Machine Translation (2020.acl-main)
Copied to clipboard
| Challenge: | Neural machine translation suffers from exposure bias, and alternative approaches to mitigate this are under debate. |
| Approach: | They propose to reduce exposure bias by using minimum risk training to mitigate hallucinations . they find that exposure bias is more problematic under domain shift . |
| Outcome: | The proposed methods can reduce exposure bias even on in-domain test sets. |
K-COMP: Retrieval-Augmented Medical Domain Question Answering With Knowledge-Injected Compressor (2025.naacl-long)
Copied to clipboard
| Challenge: | Documents retrieved for closed domains require high expertise, so reader model may have difficulty comprehending the text. |
| Approach: | They propose a system which augments the prior knowledge required to answer correctly by adding thousands of tokens to the retrieved documents. |
| Outcome: | The proposed system provides the knowledge required to answer correctly and generates prior knowledge to facilitate the answer process prior to compression of the retrieved passages. |
INVITE: a Testbed of Automatically Generated Invalid Questions to Evaluate Large Language Models for Hallucinations (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in Large language models have enabled them to hold free form conversations over multiple turns, but they exhibit a tendency to make unfounded and incorrect statements, commonly labeled as hallucinations. |
| Approach: | They propose a framework to test large language models for hallucinations using automatically generated INValId questions. |
| Outcome: | The proposed framework is based on a testbed of automatically generated INValId questions to evaluate large language models for hallucinations. |
Mutual Information Alleviates Hallucinations in Abstractive Summarization (2022.emnlp-main)
Copied to clipboard
| Challenge: | Abstractive summarization models exhibit the tendency to hallucinate, i.e., output content not supported by the source document. |
| Approach: | They propose a decoding strategy that optimizes for pointwise mutual information of source and target tokens when models exhibit uncertainty. |
| Outcome: | The proposed method decreases the probability of hallucinated tokens while maintaining the Rouge and BERT-S scores of top-performing decoding strategies. |
Improving Faithfulness of Large Language Models in Summarization via Sliding Generation and Self-Consistency (2024.lrec-main)
Copied to clipboard
| Challenge: | Abstractive summarization models (LLMs) have demonstrated impressive performance in various tasks, but they are still suffering from factual inconsistency problem called hallucination. |
| Approach: | They propose to improve the faithfulness of large language models by impelling them to process the entire article more fairly and faithfully. |
| Outcome: | The proposed strategy improves the faithfulness of large language models in summarization while maintaining their fluency and informativeness. |