Challenge: Previously available quality assessments do not distinguish between hallucinations and omissions.
Approach: They propose to annotate hallucinations and omissions in machine translation using a single language pair.
Outcome: The proposed dataset covers 18 translation directions with varying resource levels and scripts.

Similar Papers

HAT: Hallucination Annotation for Translation (2026.acl-long)

Copied to clipboard

Challenge: Hallucinations in machine translation (MT) outputs are prone to hallucination, authors say . lack of high-quality benchmarks for halluciation detection has hindered MT deployments .
Approach: They propose a dataset that provides annotated hallucination distributions and benchmarks . they use 350,959 span-level annotations across 38 language pairs to analyze hallucis a MT output .
Outcome: The proposed dataset provides high-quality benchmarks for hallucination detection in machine translation . the dataset includes 350,959 span-level annotated samples across 38 language pairs .
HalluLens: LLM Hallucination Benchmark (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) generate responses that deviate from user input or training data, a phenomenon known as "hallucination" .
Approach: They propose a hallucination benchmark HalluLens that includes both extrinsic and intrinsic evaluation tasks to distinguish between extrindic and intrinsic hallucines.
Outcome: The proposed framework disentangles LLM hallucination from "factuality" and distinguishes between extrinsic and intrinsic hallucines to promote consistency and facilitate research.
HALoGEN: Fantastic LLM Hallucinations and Where to Find Them (2025.acl-long)

Copied to clipboard

Challenge: generative large language models produce hallucinations that are not aligned with world knowledge or input context.
Approach: They propose a hallucination benchmark framework that measures hallucinism in large language models . they evaluate 150,000 generations from 14 language models and find they are riddled with hallucinos .
Outcome: The proposed framework evaluates 150,000 generations from 14 language models.
On the Hallucination in Simultaneous Machine Translation (2024.acl-short)

Copied to clipboard

Challenge: Currently, there are no studies which systematically analyze hallucination in SiMT.
Approach: They conduct a comprehensive analysis of hallucination in simultaneous machine translation (SiMT) they find that halluciation is extremely severe, especially as latency increases .
Outcome: The results show that it is possible to alleviate hallucination by decreasing the over usage of target-side information for SiMT.
Machine Translation Hallucination Detection for Low and High Resource Languages using Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for detecting hallucinations in machine translation are limited for low-resource languages.
Approach: They evaluate sentence-level hallucination detection approaches using Large Language Models (LLMs) they find that the choice of model is essential for performance.
Outcome: The proposed models outperform the existing models in HRLs and LRLs on average by 0.16 MCC.
Detecting and Mitigating Hallucinations in Multilingual Summarisation (2023.emnlp-main)

Copied to clipboard

Challenge: Existing faithfulness metrics for abstractive summarisation models focus on English . metric mFACT is best suited to detect hallucinations in cross-lingual transfer .
Approach: They propose a method to evaluate the faithfulness of non-English summaries by translation-based transfer from multiple English faithfulness metrics.
Outcome: The proposed method reduces hallucinations in cross-lingual transfer by weighing the loss of each training example by its faithfulness score.
Detecting and Mitigating Hallucinations in Machine Translation: Model Internal Workings Alone Do Well, Sentence Similarity Even Better (2023.acl-long)

Copied to clipboard

Challenge: a recent study shows that without artificially encouraging models to hallucinate, existing methods fall short . hallucinations are cases when the model generates output that is partially or fully unrelated to the source sentence.
Approach: They propose a method that evaluates the percentage of the source contribution to a generated translation.
Outcome: The proposed method improves detection accuracy for the most severe hallucinations by a factor of 2.
The Troubling Emergence of Hallucination in Large Language Models - An Extensive Definition, Quantification, and Prescriptive Remediations (2023.emnlp-main)

Copied to clipboard

Challenge: Recent advances in Large Language Models have generated widespread acclaim, but hallucination has also emerged as a by-product.
Approach: They propose a fine-grained discourse on profiling hallucination based on its degree, orientation, and category . they categorize hallucines into six types: acronym ambiguity, generated golem, virtual voice, geographic erratum, time wrap .
Outcome: The proposed method categorizes hallucination into six types based on their degree, orientation, and category .
HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Existing large language models (LLMs) are prone to generate hallucinations . a recent study shows that LLMs are able to generate content that conflicts with the source or cannot be verified by factual knowledge.
Approach: They propose a framework to evaluate the performance of large language models (LLMs) they propose to use a sample of generated and human-annotated hallucinated samples to evaluate their performance .
Outcome: The proposed framework generates and annotates hallucinated samples from ChatGPT . the results show that existing LLMs face great challenges in recognizing hallucines .
HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing studies on hallucination focus on text or vision, while few audio-oriented studies are limited in scale, modality coverage, and diagnostic depth.
Approach: They propose a large-scale benchmark for evaluating hallucinations across speech, sound, and music.
Outcome: The proposed model improves hallucination rate, yes/no bias, error-type analysis, and refusal rate.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations