Challenge: sycophancy is a type of hallucination in Large Language Models, which can lead to false information being presented.
Approach: They explore the sycophantic tendencies of Large Language Models where models provide accurate answers even if they are not entirely correct.
Outcome: The proposed models generate factually correct statements even when they are not completely correct.

Similar Papers

Challenging the Evaluator: LLM Sycophancy Under User Rebuttal (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) often exhibit sycophancy, distorting responses to align with user beliefs.
Approach: They investigate why LLMs exhibit sycophancy when challenged in subsequent conversational turns, yet perform well when evaluating conflicting arguments presented simultaneously?
Outcome: The proposed models are more likely to endorse a user’s counterargument when framed as a follow-up from a users, rather than when both responses are presented simultaneously for evaluation.
Addressing Bias and Hallucination in Large Language Models (2024.lrec-tutorials)

Copied to clipboard

Challenge: This tutorial provides a comprehensive overview of two critical aspects of Large Language Models: bias and hallucination.
Approach: This tutorial provides an overview of two critical aspects of Large Language Models: bias and hallucination.
Outcome: This tutorial delves into the complex dimensions of Large Language Models (LLMs) it outlines ethical considerations pertinent to their development and discusses hallucination, a prevalent issue in generative AI systems such as LLMs.
A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have gained significant attention due to their capabilities in performing diverse tasks across domains.
Approach: They review the primary challenges and limitations causing inconsistencies in evaluations . early models could generate coherent text but limited to simple tasks .
Outcome: The proposed evaluations are reproducible, reliable, and robust.
A Survey of Uncertainty Estimation Methods on Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated remarkable capabilities but could produce biased, hallucinated, or non-factual responses.
Approach: They propose to conduct extensive experimental evaluations of LLM uncertainty estimation methods . large language models have demonstrated remarkable capabilities across tasks .
Outcome: The proposed method could produce biased, hallucinated, or non-factual responses . a lack of comprehensive surveys on LLM uncertainty estimation is a problem .
Echoes of Agreement: Argument Driven Sycophancy in Large Language models (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluations of political biases in Large Language Models outline the high sensitivity to prompt formulation.
Approach: They investigate how argumentative prompts induce sycophantic behaviour in Large Language Models in a political context.
Outcome: The proposed model sycophancy is observed in single and multi-turn interactions and its intensity correlates with argument strength.
Tutorial Proposal: Hallucination in Large Language Models (2024.lrec-tutorials)

Copied to clipboard

Challenge: Grasping the intricacies of hallucination in LLMs can be daunting, especially for those new to the field.
Approach: This tutorial aims to bridge the gap between the field and the field of hallucination . it will explore the key aspects of hallucinonation, including benchmarking, detection, and mitigation techniques .
Outcome: This tutorial will explore the key aspects of hallucination in LLMs . it will also explore the specific constraints and shortcomings of current approaches .
Pointing to a Llama and Call it a Camel: On the Sycophancy of Multimodal Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Multimodal large language models exhibit a pronounced form of visual sycophantic behavior when they process image inputs.
Approach: They propose a technique that allows multimodal large language models to engage in reflective reasoning and determine whether a user’s instruction is misleading or corrective.
Outcome: The proposed model resists misleading instructions but is stubborn even if it is wrong.
The Troubling Emergence of Hallucination in Large Language Models - An Extensive Definition, Quantification, and Prescriptive Remediations (2023.emnlp-main)

Copied to clipboard

Challenge: Recent advances in Large Language Models have generated widespread acclaim, but hallucination has also emerged as a by-product.
Approach: They propose a fine-grained discourse on profiling hallucination based on its degree, orientation, and category . they categorize hallucines into six types: acronym ambiguity, generated golem, virtual voice, geographic erratum, time wrap .
Outcome: The proposed method categorizes hallucination into six types based on their degree, orientation, and category .
The Dawn After the Dark: An Empirical Study on Factuality Hallucination in Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: a growing number of researchers are studying the hallucination issue in large language models.
Approach: They propose a hallucination detection benchmark and a method to detect hallucines in LLMs.
Outcome: The proposed method detects hallucinations and mitigates them using different training stages.
Systematic Biases in LLM Simulations of Debates (2024.emnlp-main)

Copied to clipboard

Challenge: Current research suggests that LLM-based agents become increasingly human-like in their performance, sparking interest in using these AI agents as substitutes for human participants in behavioral studies.
Approach: They propose to use LLMs to simulate political debates on topics that are important aspects of people’s day-to-day lives and decision-making processes.
Outcome: The proposed model can simulate political debates on topics that are important aspects of people’s day-to-day lives and decision-making processes.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations