Challenge: Existing studies on hallucination focus on text or vision, while few audio-oriented studies are limited in scale, modality coverage, and diagnostic depth.
Approach: They propose a large-scale benchmark for evaluating hallucinations across speech, sound, and music.
Outcome: The proposed model improves hallucination rate, yes/no bias, error-type analysis, and refusal rate.

Similar Papers

HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Existing large language models (LLMs) are prone to generate hallucinations . a recent study shows that LLMs are able to generate content that conflicts with the source or cannot be verified by factual knowledge.
Approach: They propose a framework to evaluate the performance of large language models (LLMs) they propose to use a sample of generated and human-annotated hallucinated samples to evaluate their performance .
Outcome: The proposed framework generates and annotates hallucinated samples from ChatGPT . the results show that existing LLMs face great challenges in recognizing hallucines .
AHA: Aligning Large Audio-Language Models for Reasoning Hallucinations via Counterfactual Hard Negatives (2026.findings-acl)

Copied to clipboard

Challenge: Large Audio-Language Models suffer from hallucinations, e.g., generating text not grounded in the audio input.
Approach: They propose a framework to address hallucination problems in large audio-language models . they use a preference dataset to test the model's accuracy .
Outcome: The proposed model outperforms the latest SOTA methods in terms of performance and generalization.
DiaHalu: A Dialogue-level Hallucination Evaluation Benchmark for Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmarks for hallucination detection are intentionally generated by large language models (LLMs) however, many focus on factuality while ignoring faithfulness.
Approach: They propose a dialogue-level hallucination evaluation benchmark for large language models . they integrate the topic into prompts and facilitate a dialog between two LLMs .
Outcome: The proposed benchmark covers four common multi-turn dialogue domains and five hallucination subtypes, extended from factuality and faithfulness hallucines.
The Troubling Emergence of Hallucination in Large Language Models - An Extensive Definition, Quantification, and Prescriptive Remediations (2023.emnlp-main)

Copied to clipboard

Challenge: Recent advances in Large Language Models have generated widespread acclaim, but hallucination has also emerged as a by-product.
Approach: They propose a fine-grained discourse on profiling hallucination based on its degree, orientation, and category . they categorize hallucines into six types: acronym ambiguity, generated golem, virtual voice, geographic erratum, time wrap .
Outcome: The proposed method categorizes hallucination into six types based on their degree, orientation, and category .
Evaluating Evaluation Metrics – The Mirage of Hallucination Detection (2025.findings-emnlp)

Copied to clipboard

Challenge: a large-scale empirical evaluation of hallucination detection metrics is conducted . hallucinosity is a significant obstacle to the reliability and widespread adoption of language models .
Approach: They conduct large-scale empirical evaluation of hallucination detection metrics . they compare hallucinian language models, language models and decoding methods .
Outcome: The results show that the evaluations of hallucination detection metrics fail to align with human judgments, they say . they also show that evaluations with LLM-based evaluation yield the best overall results .
InterrogateLLM: Zero-Resource Hallucination Detection in LLM-Generated Answers (2024.acl-long)

Copied to clipboard

Challenge: Existing methods for detecting hallucinations in large language models are limited due to their high frequency and high accuracy.
Approach: They propose a method to detect hallucinations in large language models by repeating model-generated responses from its generated answer.
Outcome: The proposed method achieves 87% hallucinations in a specific experiment without external knowledge.
Fine-Grained Detection of Context-Grounded Hallucinations Using LLMs (2026.findings-acl)

Copied to clipboard

Challenge: Existing representations of hallucinations limit the types of errors that can be expressed, so we propose a new representation based on free-form textual descriptions, capturing the full range of possible errors.
Approach: They propose a benchmark for localizing hallucinations using LLMs with a human annotation of over 1,000 examples and a protocol to verify its quality in a humans evaluation.
Outcome: The proposed representation captures the full range of possible errors, and the best model achieves an F1 score of 0.67.
MHALO: Evaluating MLLMs as Fine-grained Hallucination Detectors (2025.findings-acl)

Copied to clipboard

Challenge: Hallucination remains a critical challenge for multimodal large language models, undermining their reliability in real-world applications.
Approach: They propose a benchmark specifically designed for evaluating MLLMs’ capability in performing token-level hallucination detection (FHD) . they use curated training data to train a specialized model that significantly outperforms existing models.
Outcome: The proposed model outperforms existing models in the evaluation of 9 MLLMs and reaches an average F1IoU of 40.59%.
Detecting Hallucinations in SpeechLLMs at Inference Time Using Attention Maps (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for hallucination detection for text-based LLMs do not capture audio-specific signals.
Approach: They propose to capture pathological attention patterns associated with hallucination using four attention-derived metrics to train lightweight logistic regression classifiers.
Outcome: The proposed approach outperforms baselines on in-domain data and generalises to out-of-domain ASR settings.
HALoGEN: Fantastic LLM Hallucinations and Where to Find Them (2025.acl-long)

Copied to clipboard

Challenge: generative large language models produce hallucinations that are not aligned with world knowledge or input context.
Approach: They propose a hallucination benchmark framework that measures hallucinism in large language models . they evaluate 150,000 generations from 14 language models and find they are riddled with hallucinos .
Outcome: The proposed framework evaluates 150,000 generations from 14 language models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations