Challenge: Large Language Models (LLMs) are increasingly serving as evaluators in Natural Language Generation (NLG) tasks.
Approach: They propose a framework that measures the discernment of Large Language Models (LLMs) across diverse NLG tasks.
Outcome: The proposed framework provides quantitative discernment scores for LLMs across four NLG tasks.

Similar Papers

Are Large Language Model-based Evaluators the Solution to Scaling Up Multilingual Evaluation? (2024.findings-eacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) excel in various tasks, but their evaluation, especially in languages beyond the top 20, remains inadequate due to existing benchmarks and metrics limitations.
Approach: They propose to use Large Language Models as evaluators to rank or score other models’ outputs by calibrating them against 20K human judgments across three text-generation tasks, five metrics, and eight languages.
Outcome: The proposed evaluation methods can be used to improve multilingual evaluation by calibrating them against 20K human judgments across three text-generation tasks, five metrics, and eight languages.
Leveraging Large Language Models for NLG Evaluation: Advances and Challenges (2024.emnlp-main)

Copied to clipboard

Challenge: introducing Large Language Models (LLMs) has opened new avenues for assessing generated content quality, e.g., coherence, creativity, and context relevance.
Approach: They propose a taxonomy for organizing existing LLM-based evaluation metrics and a structured framework to understand and compare them.
Outcome: The proposed taxonomy offers a framework to understand and compare LLM-based evaluation methods.
LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks (2025.acl-short)

Copied to clipboard

Challenge: Existing evaluations of NLP models with LLMs are based on human judgments . however, there are concerns about their validity and reproducibility in proprietary models .
Approach: They evaluate 11 current LLMs for their ability to replicate annotations. they show substantial variance across models and datasets.
Outcome: The proposed model can replicate human annotations on 20 NLP datasets and show substantial variance across models and datasets.
Exploring the Reliability of Large Language Models as Customized Evaluators for Diverse NLP Tasks (2025.coling-main)

Copied to clipboard

Challenge: Existing work uses large language models (LLMs) to evaluate natural language process tasks, but there are shortcomings in current LLMs.
Approach: They examine the alignment between LLM evaluators and human annotators by comparing conventional and alignment tasks with different evaluation criteria.
Outcome: The proposed models excel in general criteria, such as fluency, but face challenges with complex criteria, including numerical reasoning.
Benchmarking Generation and Evaluation Capabilities of Large Language Models for Instruction Controllable Summarization (2024.findings-naacl)

Copied to clipboard

Challenge: Recent studies have found that large language models (LLMs) can achieve state-of-the-art performance on generic summarization benchmarks, but their performance on more complex summarizing task settings is less studied.
Approach: They benchmark large language models on instruction controllable text summarization . they use 4 evaluation protocols and 11 LLMs to evaluate their performance .
Outcome: The proposed model performs well on instruction controllable text summarization tasks with 4 evaluation protocols and 11 LLMs.
In Benchmarks We Trust ... Or Not? (2025.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks for Large Language Models (LLMs) are inadequate and lack a clear solution.
Approach: They propose checklists to cover all aspects of benchmarking issues, both for benchmark creation and usage.
Outcome: The proposed checklists cover all aspects of benchmarking issues, both for benchmark creation and usage.
Evaluation Metrics in the Era of GPT-4: Reliably Evaluating Large Language Models on Sequence to Sequence Tasks (2023.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) evaluation is a patchy and inconsistent landscape . established automatic evaluation metrics are poor surrogates, correlating weakly with human judgement.
Approach: They propose to use both automatic and human evaluation to evaluate generative LLMs on three NLP benchmarks: text summarisation, text simplification and grammatical error correction.
Outcome: The proposed model outperforms many popular models according to human reviewers on the majority of metrics, while scoring much worse when using classic automatic evaluation metrics.
SCORE: Systematic COnsistency and Robustness Evaluation for Large Language Models (2025.naacl-industry)

Copied to clipboard

Challenge: Typical evaluations of Large Language Models (LLMs) report a single accuracy metric per dataset, often derived from an optimized setup.
Approach: They propose a framework for non-adversarial evaluation of large language models that evaluates models by repeatedly testing them on the same benchmarks in various setups.
Outcome: The proposed framework evaluates models by repeatedly testing them on the same benchmarks in various setups to give a realistic estimate of their accuracy and consistency.
From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) inspire the "LLM-as-a-judge" paradigm . traditional methods of assessment and evaluation fail in dynamic and open-ended scenarios .
Approach: They propose a paradigm where LLMs are leveraged to perform scoring, ranking, or selection for machine learning evaluation scenarios.
Outcome: The proposed model-based judgment and evaluation paradigms are based on large language models and are compared to the current model-driven evaluation paradigm.
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used to judge code, but their reliability remains poorly understood.
Approach: They propose a benchmark to evaluate Large Language Models as code judges . they find that small reasoning models outperform larger non-reasoning models .
Outcome: The proposed benchmark evaluates LLM-as-a-Judge models across three coding tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations