Challenge: Existing definitions of behavioural consistency are inconsistent across many studies.
Approach: They propose a behavioural consistency model and propose behavioural taxonomy that classifies consistencies into several sub-categories.
Outcome: The proposed model performs poorly on 19 test cases while exhibiting high inconsistency in many cases.

Similar Papers

SCORE: Systematic COnsistency and Robustness Evaluation for Large Language Models (2025.naacl-industry)

Copied to clipboard

Challenge: Typical evaluations of Large Language Models (LLMs) report a single accuracy metric per dataset, often derived from an optimized setup.
Approach: They propose a framework for non-adversarial evaluation of large language models that evaluates models by repeatedly testing them on the same benchmarks in various setups.
Outcome: The proposed framework evaluates models by repeatedly testing them on the same benchmarks in various setups to give a realistic estimate of their accuracy and consistency.
Evaluating the Consistency of LLM Evaluators (2025.coling-main)

Copied to clipboard

Challenge: Large language models (LLMs) have shown potential as general evaluators with the benefits of speed and cost.
Approach: They conduct extensive studies on the two aspects of consistency in LLM evaluations, Self-Consistency (SC) and Inter-scale Consistency on different scoring scales and criterion granularity with open-source and proprietary models.
Outcome: The results show that strong proprietary models are not necessarily consistent evaluators, highlighting the importance of considering consistency in assessing the capability of LLM evalueators.
SaGE: Evaluating Moral Consistency in Large Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies on Large Language Models (LLMs) have focused on accuracy but lack universally agreed-upon answers for moral scenarios.
Approach: They propose a measure called Semantic Graph Entropy to measure a model's moral consistency grounded in "Rules of Thumb" they construct a moral Consistency Corpus (MCC) with 50K moral questions and the RoTs they followed to investigate LLM consistency on two popular datasets.
Outcome: The proposed measure measures moral consistency on two popular datasets .
The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models (2025.naacl-long)

Copied to clipboard

Challenge: a recent study evaluated language models using abstract evaluation criteria that lack the flexibility and granularity of human assessment.
Approach: They propose a benchmark to evaluate nine distinct language models' capabilities . they use instance-specific evaluation criteria to mirror human evaluation .
Outcome: The proposed benchmark evaluates nine distinct capabilities of language models across 77 tasks.
On the Consistency of Commonsense in Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Existing evaluations of commonsense for large language models focus on downstream knowledge tasks, failing to probe whether LLMs truly understand and utilize knowledge or merely memorize it.
Approach: They propose to automatically construct a large benchmark named CoCo which measures LLMs’ knowledge memorization, comprehension, and application and examines the consistency between these tasks.
Outcome: The proposed benchmark systematically assesses LLMs’ knowledge memorization, comprehension, and application and examines the consistency between these tasks.
Are Economists Always More Introverted? Analyzing Consistency in Persona-Assigned LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Personalized Large Language Models are increasingly used in diverse applications . prior research examined how well LLMs adhere to predefined personas in writing style . inconsistent responses are influenced by multiple factors, including the assigned persona, stereotypes, and model design choices.
Approach: They propose a standardized framework to analyze consistency in persona-assigned LLMs.
Outcome: The proposed framework evaluates personas across multiple tasks and runs.
A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have gained significant attention due to their capabilities in performing diverse tasks across domains.
Approach: They review the primary challenges and limitations causing inconsistencies in evaluations . early models could generate coherent text but limited to simple tasks .
Outcome: The proposed evaluations are reproducible, reliable, and robust.
Benchmarking and Improving LLM Robustness for Personalized Generation (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluations focus on whether a model’s responses align with a user’s preferences, but factuality is an important yet overlooked dimension.
Approach: They propose a scalable framework for evaluating robustness of large language models in personalization and a new dataset, PERGData.
Outcome: The proposed framework improves robustness by 25% across models.
Firm or Fickle? Evaluating Large Language Models Consistency in Sequential Interactions (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse tasks, but their deployment in high-stake domains requires consistent and coherent behavior across multiple rounds of user interaction.
Approach: They propose a framework for evaluating and improving LLM response consistency, and introduce a benchmark dataset to evaluate LLM consistency.
Outcome: The proposed framework improves response stability without sacrificing accuracy, and offers a practical path toward more dependable behavior in critical, real-world deployments.
Rethinking the Evaluation of Alignment Methods: Insights into Diversity, Generalisation, and Safety (2026.eacl-srw)

Copied to clipboard

Challenge: Existing studies focus on individual techniques or specific dimensions, lacking a holistic assessment of the inherent trade-offs.
Approach: They propose a framework that compares LLM alignment methods across five axes . they use a validated LLM-as-judge prompt to compare the results .
Outcome: The proposed framework compares LLM alignment methods across factuality, safety, conciseness, proactivity, diversity and safety axes . it provides insights into trade-offs of common alignment methods, guiding the development of more balanced and reliable LLMs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations