Challenge: Existing work has demonstrated the presence of anchoring bias in large language models . Existing research does not predict when a model will be most susceptible to anchoring .
Approach: They analyze anchoring bias as a function of model confidence and accuracy . they find that incorrect models resist anchoring as effectively as accurate ones .
Outcome: The findings suggest that anchoring resistance is a structural property of uncertainty rather than knowledge correctness.

Similar Papers

How Does Cognitive Bias Affect Large Language Models? A Case Study on the Anchoring Effect in Price Negotiation Simulations (2025.findings-emnlp)

Copied to clipboard

Challenge: Cognitive biases can be observed in LLMs, affecting their reliability in real-world applications.
Approach: They investigate the anchoring effect in LLM-driven price negotiations . reasoning models are less prone to the anchor effect, they find .
Outcome: The proposed study shows that LLMs are influenced by the anchoring effect like humans . reasoning models are less prone to the anchor effect, but personality traits are not affected .
On the Calibration of Large Language Models and Alignment (2023.findings-emnlp)

Copied to clipboard

Challenge: Large language models are becoming more popular and are proving to be reliable . however, their reliability is often understudied due to their uncertainty and complex structure .
Approach: They conduct a systematic examination of the calibration of aligned language models throughout the entire construction process including pretraining and alignment training.
Outcome: The results shed light on whether popular large language models are well-calibrated and how the training process influences model calibration.
Language Models Resist Alignment: Evidence From Data Compression (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) may exhibit undesirable behaviors due to the inevitable biases and harmful content present in training.
Approach: They propose to investigate the elasticity of large language models by examining their performance.
Outcome: The proposed model performance declines rapidly before reverting to the pre-training distribution, the authors show . the proposed model weight and code are available at pku-lm-res ist-alignment.github.io.
Conformity in Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Conformity is a form of social influence that affects the way people respond to information.
Approach: They adapt psychological experiments to examine the extent of conformity in large language models.
Outcome: The proposed interventions mitigate conformity by reducing the naturalness of majority tones and reducing instruction-tuned models.
When Quantization Affects Confidence of Large Language Models? (2024.findings-naacl)

Copied to clipboard

Challenge: Existing studies have shown that quantization compromises performance and exacerbates biases in Large Language Models.
Approach: They propose an explanation for quantization loss based on confidence levels . they propose a range of efficient compression and acceleration methods including quan-tization .
Outcome: The proposed methods show that quantization decreases confidence regarding true labels and that it exacerbates biases across different scales.
Bias in the Mirror : Are LLMs opinions robust to their own adversarial attacks (2025.acl-long)

Copied to clipboard

Challenge: Existing work on large language models lacks robustness, highlighting the limitations of such models.
Approach: They propose a novel approach where two LLMs engage in self-debate to persuade a neutral version of the model.
Outcome: The proposed approach examines whether large language models are robust during interactions and whether they are susceptible to reinforcing misinformation or shifting to harmful viewpoints.
Methods for Estimating and Improving Robustness of Language Models (2022.naacl-srw)

Copied to clipboard

Challenge: Large language models suffer from weak generalisation ability due to shallow textual relations over full semantic complexity of the problem.
Approach: They propose to incorporate some of these measures into training objectives to enhance distributional robustness of LLMs.
Outcome: The proposed models outperform human models on complex tasks and outperformed other models on deep networks.
Beyond Performance: Quantifying and Mitigating Label Bias in LLMs (2024.naacl-long)

Copied to clipboard

Challenge: Large language models exhibit undesirable preference toward predicting certain answers over others, despite their adaptability to diverse tasks.
Approach: They propose a label bias calibration method that outperforms recent calibration approaches for improving performance and mitigating label bias.
Outcome: The proposed method outperforms calibration approaches for improving performance and mitigating label bias.
Prior Beliefs Prejudice LLM-as-Judge: Evidence from Persuasion Evaluation (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models are increasingly used as judges to evaluate text quality, content and assess arguments.
Approach: They propose to exploit belief-conditioned rating inflation by using persuasion-based probing to examine persuasive arguments.
Outcome: The proposed model fails to evaluate persuasive arguments based on belief alignment . the model fails in three of the three tasks, with belief-conditioned rating inflation accounting for 88% of cases.
Evaluating the Instruction-Following Robustness of Large Language Models to Prompt Injection (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated exceptional proficiency in instruction-following, making them increasingly integral to various applications.
Approach: They establish a benchmark to evaluate the robustness of instruction-following LLMs against prompt injection attacks, assessing their ability to discern which instructions to follow and which to disregard.
Outcome: The proposed model is overly sensitive to prompt injection attacks, focusing on the latter part of the prompt without fully understanding the context.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations