Challenge: Consensual Assessment Technique (CAT) for large language models is used to evaluate creativity, but is costly and time-consuming with non-experts.
Approach: They adapt the Consensual Assessment Technique (CAT) for Large Language Models to a 90-poem dataset with a ground truth based on publication venue.
Outcome: The proposed method outperforms the best human non-expert evaluations by significantly outperforming the best language models.

Similar Papers

Capabilities and Evaluation Biases of Large Language Models in Classical Chinese Poetry Generation: A Case Study on Tang Poetry (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly applied to creative domains, yet performance in classical Chinese poetry generation and evaluation remains poorly understood.
Approach: They propose a framework that combines computational metrics, LLM-as-a-judge assessment, and human expert validation to evaluate large language models.
Outcome: The proposed framework evaluates state-of-the-art LLMs across multiple dimensions of poetic quality in Tang poetry generation.
Sonnet or Not, Bot? Poetry Evaluation for Large Models and Datasets (2024.findings-emnlp)

Copied to clipboard

Challenge: a task evaluates how well LLMs recognize poetry, but performance varies by poetic form . performance varying by poetic forms; models struggle to identify unfixed poetic forms .
Approach: They use a benchmark dataset to evaluate how well LLMs recognize poetry . they find that the models can identify fixed poetic forms with high accuracy .
Outcome: The proposed task evaluates how well LLMs recognize poetry features . performance varies significantly by poetic form; models struggle to identify unfixed forms . authors urge more work that builds nuance and ambiguity into humanistic benchmarks .
Poller: Are LLMs Suitable for Evaluating Poetry Understanding Task? (2026.findings-acl)

Copied to clipboard

Challenge: Traditional methods for poetry evaluation are expensive and unsuitable for large-scale data.
Approach: They propose a method leveraging Large Language Models to evaluate poetry understanding tasks using Large Language models.
Outcome: The proposed method reduces the evaluation error between LLMs and humans by adopting the poet's perspective.
Language Model Council: Democratically Benchmarking Foundation Models on Highly Subjective Tasks (2025.naacl-long)

Copied to clipboard

Challenge: Existing evaluations of Large Language Models (LLMs) rely on a single large model to score outputs from other LLMs, but this is prone to intra-model bias and many tasks may be too subjective for a one model to judge fairly.
Approach: They propose a language model council where a group of LLMs collaborate to create tests, respond to them, and evaluate each other’s responses to produce a ranking in a democratic fashion.
Outcome: The proposed model produces rankings that are more separable and robust than any individual LLM judge.
Are Large Language Model-based Evaluators the Solution to Scaling Up Multilingual Evaluation? (2024.findings-eacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) excel in various tasks, but their evaluation, especially in languages beyond the top 20, remains inadequate due to existing benchmarks and metrics limitations.
Approach: They propose to use Large Language Models as evaluators to rank or score other models’ outputs by calibrating them against 20K human judgments across three text-generation tasks, five metrics, and eight languages.
Outcome: The proposed evaluation methods can be used to improve multilingual evaluation by calibrating them against 20K human judgments across three text-generation tasks, five metrics, and eight languages.
Leveraging Large Language Models for NLG Evaluation: Advances and Challenges (2024.emnlp-main)

Copied to clipboard

Challenge: introducing Large Language Models (LLMs) has opened new avenues for assessing generated content quality, e.g., coherence, creativity, and context relevance.
Approach: They propose a taxonomy for organizing existing LLM-based evaluation metrics and a structured framework to understand and compare them.
Outcome: The proposed taxonomy offers a framework to understand and compare LLM-based evaluation methods.
A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have gained significant attention due to their capabilities in performing diverse tasks across domains.
Approach: They review the primary challenges and limitations causing inconsistencies in evaluations . early models could generate coherent text but limited to simple tasks .
Outcome: The proposed evaluations are reproducible, reliable, and robust.
How Reliable is Multilingual LLM-as-a-Judge? (2025.findings-emnlp)

Copied to clipboard

Challenge: LLMs are a popular evaluation strategy, but their reliability in multilingual evaluation remains uncertain.
Approach: They evaluate five models from different model families across five diverse tasks involving 25 languages.
Outcome: The models perform poorly across languages and average Fleiss’ Kappa is 0.3 .
Towards A “Novel” Benchmark: Evaluating Literary Fiction with Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) context windows have enabled them to process inputs over 100K tokens and generate outputs of up to 10K token.
Approach: They propose a multi-level evaluation framework that incorporates ten metrics across the Macro, Meso, and Micro levels and an annotated fiction dataset.
Outcome: The proposed framework incorporates ten metrics across the Macro, Meso, and Micro levels and is based on a human-human-AI dataset.
Sentiment Analysis in the Era of Large Language Models: A Reality Check (2024.findings-naacl)

Copied to clipboard

Challenge: Sentiment analysis (SA) has been a long-standing research area in natural language processing.
Approach: They propose a benchmark to evaluate LLMs' SA abilities and propose 'sentiEval' benchmark to be used for a more comprehensive evaluation.
Outcome: The proposed benchmark outperforms small language models on 26 datasets on 13 tasks and compared them with LLMs trained on domain-specific datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations