Challenge: Existing evaluation methods focus predominantly on multiple-choice and question-answering tasks, leaving open-ended generation largely unaddressed.
Approach: They propose an evaluation framework that assesses LLM pluralism in open-ended generation by comparing outputs against free-form crowd responses.
Outcome: The proposed evaluation framework decomposes ground-truth responses into atomic, non-overlapping claims and evaluates whether LLMs adequately cover this diverse claim space.

Similar Papers

Too Open for Opinion? Embracing Open-Endedness in Large Language Models for Social Simulation (2026.eacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used to simulate public opinion and other social phenomena.
Approach: They argue that open-endedness is essential for realistic social simulations . they argue that it captures expressiveness and individuality .
Outcome: The proposed frameworks can improve measurement and design, support exploration of unanticipated views, and reduce researcher-imposed directive bias.
Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration (2024.emnlp-main)

Copied to clipboard

Challenge: Existing alignment paradigms for large language models learn an averaged human preference and struggle to model diverse preferences across cultures, demographics, and communities.
Approach: They propose a modular framework that "plugs" into a base LLM a pool of smaller but specialized community LMs where models collaborate in distinct modes to support three modes of pluralism: Overton, steerable, and distributional.
Outcome: The proposed framework “plugs into” a base LLM a pool of smaller but specialized community LMs, where models collaborate in distinct modes to support three modes of pluralism: Overton, steerable, and distributional.
WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback (2026.acl-long)

Copied to clipboard

Challenge: Traditional alignment methods rely on human annotations and are subjective and misalignment with real-world user preferences.
Approach: They propose a framework that leverages in-situ user feedback during conversations with LLMs to create preference datasets automatically.
Outcome: The proposed framework identifies and classifies user feedback to LLM responses between conversation turns and creates examples of preferred and dispreferred responses according to user preferences.
Can LLMs Really Judge? A Progressive Argumentation-Mining Framework for Distinguishing Understanding from Aggregation (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluations of large language models rely on dataset-based generation accuracy . however, generative correctness does not guarantee discriminative capability to verify solutions .
Approach: They propose a diagnostic framework that explicitly controls context and isolates discriminative behaviors.
Outcome: The proposed framework explicitly controls context and isolates discriminative behaviors.
Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-Judge (2025.acl-long)

Copied to clipboard

Challenge: Existing methods rely on majority voting or criteria expansion to capture detailed and detailed details, often leading to incomplete outcomes.
Approach: They propose a method which introduces additional crowd responses to compare with the candidate responses, thereby exposing deeper and more comprehensive details within the candidate answers.
Outcome: Experiments show that the proposed method improves evaluation reliability and achieves an average gain of 6.7% across five benchmarks.
Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) produce incomplete or selectively omit key information . omissions of key information or misrepresentation of conflicting evidence can cause harm .
Approach: They propose a method that decomposes texts into atomic statements and uses natural language inference to identify missing facts and a Q A-based metric that extracts question-answer pairs and compares responses across sources.
Outcome: The proposed evaluation metrics show they perform better than more complex metrics, but at a cost.
DebateQA: Evaluating Question Answering on Debatable Knowledge (2026.findings-eacl)

Copied to clipboard

Challenge: Existing QA benchmarks that provide fixed answers to debatable questions are inadequate for evaluating their performance.
Approach: They propose to use a dataset of 2,941 debatable questions to assess their ability to provide comprehensive answers to inherently debatably asked questions.
Outcome: The proposed model performs well on 2,941 debatable questions accompanied by human-annotated partial answers that capture a variety of perspectives.
MAPLE: Multi-Aspect Panels of LLM Evaluators for Open-Ended Questions (2026.findings-acl)

Copied to clipboard

Challenge: LLM-as-a-Judge uses LLMs to evaluate open-ended questions . however, the discrepancy between LLM generated evaluations and human evaluations remains a critical problem in this field .
Approach: They propose a framework that orchestrates evaluations across multiple criteria using multiple LLMs.
Outcome: The proposed framework achieves superior alignment with human evaluations compared to baselines.
Rethinking Pragmatics in Large Language Models: Towards Open-Ended Evaluation and Preference Tuning (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods to assess social-pragmatic inference in large language models are inadequacy, and preferential tuning is the best approach.
Approach: They propose to use free-form models' responses as a measure to assess social-pragmatic reasoning and advocate for preference optimization over supervised finetuning (SFT).
Outcome: The proposed model outperforms supervised finetuning (SFT) and offers a near-free launch in pragmatic abilities without compromising general capabilities.
Co-Eval: Augmenting LLM-based Evaluation with Machine Metrics (2025.emnlp-main)

Copied to clipboard

Challenge: Existing LLMs suffer from biases and misalignment due to limited functional understanding and knowledge gaps.
Approach: They introduce a framework that leverages a criteria planner model and optimized machine metrics to enhance the scalability and fairness of LLM-based evaluation.
Outcome: The proposed framework reduces biases and improves alignment with human preferences, with gains of up to 0.324 in Spearman correlation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations