Challenge: Current approaches to evaluation of large language models ignore high entropy of LLM responses.
Approach: They propose a Monte-Carlo evaluation framework for evaluating large language models . they test multiple LLMs to see if they are susceptible to cognitive biases .
Outcome: The proposed framework shows that LLMs are more human-like and less rational . it also shows that larger LLM models are more susceptible to cognitive biases .

Similar Papers

A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have gained significant attention due to their capabilities in performing diverse tasks across domains.
Approach: They review the primary challenges and limitations causing inconsistencies in evaluations . early models could generate coherent text but limited to simple tasks .
Outcome: The proposed evaluations are reproducible, reliable, and robust.
SCORE: Systematic COnsistency and Robustness Evaluation for Large Language Models (2025.naacl-industry)

Copied to clipboard

Challenge: Typical evaluations of Large Language Models (LLMs) report a single accuracy metric per dataset, often derived from an optimized setup.
Approach: They propose a framework for non-adversarial evaluation of large language models that evaluates models by repeatedly testing them on the same benchmarks in various setups.
Outcome: The proposed framework evaluates models by repeatedly testing them on the same benchmarks in various setups to give a realistic estimate of their accuracy and consistency.
Do LLMs Play Dice? Exploring Probability Distribution Sampling in Large Language Models for Behavioral Simulation (2025.coling-main)

Copied to clipboard

Challenge: LLMs are used to emulate sequential decision-making processes of humans . however, their ability to perform probabilistic sampling is limited .
Approach: They propose to use large language models (LLMs) as agents to emulate the sequential decision-making processes of humans represented as Markov decision-makers (MDPs).
Outcome: The proposed models can understand probabilities, but struggle with sampling precision . integrating coding tools can improve sampling precision, but this level of sampling precision still makes it difficult to simulate human behavior as agents.
ConSiDERS-The-Human Evaluation Framework: Rethinking Human Evaluation for Generative Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: In this position paper, we argue that human evaluation of generative large language models (LLMs) should be a multidisciplinary undertaking that draws upon the insights from disciplines such as user experience research and human behavioral psychology to ensure that the results are reliable.
Approach: They propose a framework for human evaluation of generative large language models that takes into account usability, aesthetics and cognitive biases.
Outcome: The proposed framework is based on the framework proposed by Deutsch and alnajjar . it is aimed at ensuring that human evaluation is accurate in the age of generative AI .
A Survey of Confidence Estimation and Calibration in Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated impressive capabilities across a wide range of tasks in various domains, but they can be unreliable due to factual errors in their generations.
Approach: They summarize recent advances in LLM confidence estimation and calibration and outline their main lessons learned.
Outcome: The proposed methods can be used to assess the reliability of models and to calibrate them across tasks.
A Survey of Uncertainty Estimation Methods on Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated remarkable capabilities but could produce biased, hallucinated, or non-factual responses.
Approach: They propose to conduct extensive experimental evaluations of LLM uncertainty estimation methods . large language models have demonstrated remarkable capabilities across tasks .
Outcome: The proposed method could produce biased, hallucinated, or non-factual responses . a lack of comprehensive surveys on LLM uncertainty estimation is a problem .
On Evaluating LLMs’ Capabilities as Functional Approximators: A Bayesian Evaluation Framework (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have revolutionized the way we can formulate tasks in text-in-text-out format.
Approach: They propose a new evaluation framework to comprehensively assess LLMs’ function modeling abilities by adopting a Bayesian perspective of function modeling.
Outcome: The proposed evaluation framework enables LLMs to excel in utilizing prior knowledge to develop a strong understanding of the underlying function.
LLM Evaluate: An Industry-Focused Evaluation Tool for Large Language Models (2025.coling-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated impressive capability to solve a wide range of tasks in recent years.
Approach: They propose to build an on-premise system for LLM evaluation to address the challenges in the evaluation of LLMs in real-world industrial settings.
Outcome: The proposed evaluation system protects customer privacy and protects data integrity in real-world industrial environments.
A Chinese Dataset for Evaluating the Safeguards in Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: a recent study has shown that large language models can produce harmful responses, exposing users to unexpected risks.
Approach: They propose a dataset for the safety evaluation of Chinese LLMs in Mandarin Chinese . they extend the dataset to better identify false negative and false positive examples .
Outcome: The proposed dataset is for the safety evaluation of Chinese LLMs, and is based on a Chinese dataset.
Are Large Language Model-based Evaluators the Solution to Scaling Up Multilingual Evaluation? (2024.findings-eacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) excel in various tasks, but their evaluation, especially in languages beyond the top 20, remains inadequate due to existing benchmarks and metrics limitations.
Approach: They propose to use Large Language Models as evaluators to rank or score other models’ outputs by calibrating them against 20K human judgments across three text-generation tasks, five metrics, and eight languages.
Outcome: The proposed evaluation methods can be used to improve multilingual evaluation by calibrating them against 20K human judgments across three text-generation tasks, five metrics, and eight languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations