Can Large Language Models Outperform Non-Experts in Poetry Evaluation? A Comparative Study Using the Consensual Assessment Technique (2025.emnlp-main)
Copied to clipboard
| Challenge: | Consensual Assessment Technique (CAT) for large language models is used to evaluate creativity, but is costly and time-consuming with non-experts. |
| Approach: | They adapt the Consensual Assessment Technique (CAT) for Large Language Models to a 90-poem dataset with a ground truth based on publication venue. |
| Outcome: | The proposed method outperforms the best human non-expert evaluations by significantly outperforming the best language models. |
Similar Papers
Capabilities and Evaluation Biases of Large Language Models in Classical Chinese Poetry Generation: A Case Study on Tang Poetry (2026.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly applied to creative domains, yet performance in classical Chinese poetry generation and evaluation remains poorly understood. |
| Approach: | They propose a framework that combines computational metrics, LLM-as-a-judge assessment, and human expert validation to evaluate large language models. |
| Outcome: | The proposed framework evaluates state-of-the-art LLMs across multiple dimensions of poetic quality in Tang poetry generation. |
Sonnet or Not, Bot? Poetry Evaluation for Large Models and Datasets (2024.findings-emnlp)
Copied to clipboard
| Challenge: | a task evaluates how well LLMs recognize poetry, but performance varies by poetic form . performance varying by poetic forms; models struggle to identify unfixed poetic forms . |
| Approach: | They use a benchmark dataset to evaluate how well LLMs recognize poetry . they find that the models can identify fixed poetic forms with high accuracy . |
| Outcome: | The proposed task evaluates how well LLMs recognize poetry features . performance varies significantly by poetic form; models struggle to identify unfixed forms . authors urge more work that builds nuance and ambiguity into humanistic benchmarks . |
Poller: Are LLMs Suitable for Evaluating Poetry Understanding Task? (2026.findings-acl)
Copied to clipboard
| Challenge: | Traditional methods for poetry evaluation are expensive and unsuitable for large-scale data. |
| Approach: | They propose a method leveraging Large Language Models to evaluate poetry understanding tasks using Large Language models. |
| Outcome: | The proposed method reduces the evaluation error between LLMs and humans by adopting the poet's perspective. |
Language Model Council: Democratically Benchmarking Foundation Models on Highly Subjective Tasks (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing evaluations of Large Language Models (LLMs) rely on a single large model to score outputs from other LLMs, but this is prone to intra-model bias and many tasks may be too subjective for a one model to judge fairly. |
| Approach: | They propose a language model council where a group of LLMs collaborate to create tests, respond to them, and evaluate each other’s responses to produce a ranking in a democratic fashion. |
| Outcome: | The proposed model produces rankings that are more separable and robust than any individual LLM judge. |
Are Large Language Model-based Evaluators the Solution to Scaling Up Multilingual Evaluation? (2024.findings-eacl)
Copied to clipboard
Rishav Hada, Varun Gumma, Adrian Wynter, Harshita Diddee, Mohamed Ahmed, Monojit Choudhury, Kalika Bali, Sunayana Sitaram
| Challenge: | Large Language Models (LLMs) excel in various tasks, but their evaluation, especially in languages beyond the top 20, remains inadequate due to existing benchmarks and metrics limitations. |
| Approach: | They propose to use Large Language Models as evaluators to rank or score other models’ outputs by calibrating them against 20K human judgments across three text-generation tasks, five metrics, and eight languages. |
| Outcome: | The proposed evaluation methods can be used to improve multilingual evaluation by calibrating them against 20K human judgments across three text-generation tasks, five metrics, and eight languages. |
Leveraging Large Language Models for NLG Evaluation: Advances and Challenges (2024.emnlp-main)
Copied to clipboard
| Challenge: | introducing Large Language Models (LLMs) has opened new avenues for assessing generated content quality, e.g., coherence, creativity, and context relevance. |
| Approach: | They propose a taxonomy for organizing existing LLM-based evaluation metrics and a structured framework to understand and compare them. |
| Outcome: | The proposed taxonomy offers a framework to understand and compare LLM-based evaluation methods. |
A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations (2024.emnlp-main)
Copied to clipboard
Md Tahmid Rahman Laskar, Sawsan Alqahtani, M Saiful Bari, Mizanur Rahman, Mohammad Abdullah Matin Khan, Haidar Khan, Israt Jahan, Amran Bhuiyan, Chee Wei Tan, Md Rizwan Parvez, Enamul Hoque, Shafiq Joty, Jimmy Huang
| Challenge: | Large Language Models (LLMs) have gained significant attention due to their capabilities in performing diverse tasks across domains. |
| Approach: | They review the primary challenges and limitations causing inconsistencies in evaluations . early models could generate coherent text but limited to simple tasks . |
| Outcome: | The proposed evaluations are reproducible, reliable, and robust. |
How Reliable is Multilingual LLM-as-a-Judge? (2025.findings-emnlp)
Copied to clipboard
| Challenge: | LLMs are a popular evaluation strategy, but their reliability in multilingual evaluation remains uncertain. |
| Approach: | They evaluate five models from different model families across five diverse tasks involving 25 languages. |
| Outcome: | The models perform poorly across languages and average Fleiss’ Kappa is 0.3 . |
Towards A “Novel” Benchmark: Evaluating Literary Fiction with Large Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) context windows have enabled them to process inputs over 100K tokens and generate outputs of up to 10K token. |
| Approach: | They propose a multi-level evaluation framework that incorporates ten metrics across the Macro, Meso, and Micro levels and an annotated fiction dataset. |
| Outcome: | The proposed framework incorporates ten metrics across the Macro, Meso, and Micro levels and is based on a human-human-AI dataset. |
Sentiment Analysis in the Era of Large Language Models: A Reality Check (2024.findings-naacl)
Copied to clipboard
| Challenge: | Sentiment analysis (SA) has been a long-standing research area in natural language processing. |
| Approach: | They propose a benchmark to evaluate LLMs' SA abilities and propose 'sentiEval' benchmark to be used for a more comprehensive evaluation. |
| Outcome: | The proposed benchmark outperforms small language models on 26 datasets on 13 tasks and compared them with LLMs trained on domain-specific datasets. |