Challenge: Existing evaluation methods for large language models rely on static benchmarks and standardized evaluation protocols.
Approach: They propose an adaptive evaluation framework that integrates dynamic domain knowledge modeling with progressive reasoning assessment to improve evaluation fidelity.
Outcome: Empirical results show that the framework distinguishes LLMs in terms of domain knowledge coverage and reasoning chain completeness.

Similar Papers

ChessArena: A Chess Testbed for Evaluating Strategic Reasoning Capabilities of Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluation frameworks focus on isolated question-answering tasks that may not capture the essential aspects of strategic reasoning.
Approach: They evaluate 13 large language models across over 800 games in chess . they use a chessian-based framework to test strategic reasoning and pattern recognition .
Outcome: The proposed framework improves performance and basic understanding of large language models.
Adaptation with Self-Evaluation to Improve Selective Prediction in LLMs (2023.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have shown impressive capabilities in many tasks, including natural language understanding and generation.
Approach: They propose a framework for adaptation with self-evaluation to improve selective prediction performance of large language models.
Outcome: The proposed framework outperforms state-of-the-art selective prediction methods on QA datasets and improves the AUACC from 91.23% to 92.63% and AUROC from 74.61% to 80.25%.
CodeArena: A Collective Evaluation Platform for LLM Code Generation (2025.acl-demo)

Copied to clipboard

Challenge: Large Language Models (LLMs) have reshaped code generation, but persistent challenges impede accurate assessment.
Approach: They propose an online evaluation framework tailored for large language models to assess their coding capabilities.
Outcome: a new evaluation framework for large language models (LLMs) provides unbiased, unbiased evaluations and open access to solutions and test cases.
PersonaArena: Dynamic Simulation for Evaluating and Enhancing Persona-Level Role-Playing in Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing research focuses on character-level settings and static evaluation formats fail to capture the complexity of everyday social interactions.
Approach: They propose a dynamic simulation framework for evaluating and improving persona-level role-playing in large language models (LLMs).
Outcome: The proposed framework leverages user-generated social content to construct a nuanced persona bank and elicits multi-turn, context-rich interactions within simulated social environments.
Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation (2025.coling-main)

Copied to clipboard

Challenge: Recent advances in Large Language Models have demonstrated remarkable performance across tasks.
Approach: They propose a benchmark self-evolving framework to dynamically evaluate rapidly advancing Large Language Models.
Outcome: The proposed framework extends existing benchmarks to extend models across tasks and tasks.
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation (2026.findings-acl)

Copied to clipboard

Challenge: Current evaluation methods for large language models rely on static benchmarks . limited knowledge coverage and fixed difficulties hinder the targeted optimizations resulting in superficial evaluations of LLMs - a problem that has been addressed by JudgeAgent .
Approach: They propose a knowledge-driven and dynamic evaluation framework for large language models . judgeAgent leverages LLM agents equipped with context graphs to traverse knowledge structures .
Outcome: The proposed framework can achieve comprehensive evaluations and facilitate effective model iterations.
GuessingGame: Measuring the Informativeness of Open-Ended Questions in Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models excel at factual recall, arithmetic reasoning, multi-turn dialogue . their capacity as askers, formulating strategic, adaptive, and information-seeking questions, remains less explored .
Approach: They propose a protocol for evaluating large language models as strategic question-askers . they propose entropy-based methods that filter candidates via ConceptNet and Bayesian method that tracks belief updates over semantic concepts .
Outcome: The proposed method is model-agnostic and supports post hoc analysis.
ConSiDERS-The-Human Evaluation Framework: Rethinking Human Evaluation for Generative Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: In this position paper, we argue that human evaluation of generative large language models (LLMs) should be a multidisciplinary undertaking that draws upon the insights from disciplines such as user experience research and human behavioral psychology to ensure that the results are reliable.
Approach: They propose a framework for human evaluation of generative large language models that takes into account usability, aesthetics and cognitive biases.
Outcome: The proposed framework is based on the framework proposed by Deutsch and alnajjar . it is aimed at ensuring that human evaluation is accurate in the age of generative AI .
LLM Evaluate: An Industry-Focused Evaluation Tool for Large Language Models (2025.coling-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated impressive capability to solve a wide range of tasks in recent years.
Approach: They propose to build an on-premise system for LLM evaluation to address the challenges in the evaluation of LLMs in real-world industrial settings.
Outcome: The proposed evaluation system protects customer privacy and protects data integrity in real-world industrial environments.
Think Twice Before Trusting: Self-Detection for Large Language Models through Comprehensive Answer Reflection (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to self-detection only retrospectively evaluate LLM-generated answers, leading to over-trust in incorrectly generated answers.
Approach: They propose a self-detection paradigm that considers the comprehensive answer space beyond LLM-generated answers to mitigate the over-trust in LLM generated incorrect answers.
Outcome: The proposed framework can be integrated with existing approaches for superior self-detection.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations