Challenge: Many commonsense reasoning questions require a hard selection of a single correct answer . ambiguity and semantic mismatches are common in many MCQs .
Approach: They collect plausibility judgments on 5 000 commonsense reasoning questions . they find that the answer rated most plausible does not match the benchmark gold answers .
Outcome: Experiments with LLMS reveal low accuracy and high variation in performance on the subset . high plausibility rating for the most plausible answer is highlighted in bold .

Similar Papers

LOGICAL-COMMONSENSEQA: A Benchmark for Logical Commonsense Reasoning (2026.acl-short)

Copied to clipboard

Challenge: LOGICAL-COMMONSENSEQA benchmarks evaluate commonsense reasoning as logical composition over pairs of atomic statements . commonsensible reasoning is central to human cognition and a long-standing challenge in artificial intelligence and natural language understanding.
Approach: They propose a benchmark that reframes commonsense reasoning as logical composition over pairs of atomic statements using plausibility-level operators.
Outcome: LOGICAL-COMMONSENSEQA exposes fundamental reasoning limitations and provides a framework for advancing compositional commonsense reasoning.
BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks (2026.acl-long)

Copied to clipboard

Challenge: Multiple-choice question answering (MCQ) is standard in NLP, but benchmarks lack rigorous quality control.
Approach: They propose an education-inspired toolkit that uses LLM judges to flag flaws in MCQs . they validate the tool with annotations and run it to audit 12 benchmarks based on 19-rule education rubric .
Outcome: The proposed toolkit flags three common MCQ flaws based on a 19-rule education rubric . contaminated MCqs tend to inflate accuracy, while writing errors lower it and change rankings .
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above (2025.acl-long)

Copied to clipboard

Challenge: Multiple choice question answering (MCQA) is popular for LLM evaluation due to its simplicity and human-like testing.
Approach: They argue for a reform of multiple choice question answering (MCQA) they argue for more generative formats based on human testing .
Outcome: The proposed reforms improve the quality of MCQA, the authors argue . they show that even when MCQ is a useful format, its datasets suffer from leakage, unanswerability, shortcuts and saturation.
ReMedQA: Are We Done With Medical Multiple-Choice Benchmarks? (2026.eacl-long)

Copied to clipboard

Challenge: Multiple-choice question answering (MCQA) benchmarks show near-human accuracy . but a single accuracy score is a poor proxy for competence .
Approach: They propose a medical multiple-choice question answering (MCQA) benchmark that augments three standard medical MCQA datasets with open-ended answers and systematically perturbed options.
Outcome: The proposed benchmarks show that high MCQA accuracy masks low reliability . MCQ is the dominant paradigm for assessing medical knowledge in large language models .
Revisiting the Self-Consistency Challenges in Multi-Choice Question Formats for Large Language Model Evaluation (2024.lrec-main)

Copied to clipboard

Challenge: Multi-choice questions (MCQs) are a common method for assessing the world knowledge of large language models.
Approach: They propose three knowledge-equivalent question variants to assess LLMs' world knowledge . they propose option position shuffle, option label replacement, and conversion to a True/False format .
Outcome: The proposed questions are shuffle, label replacement, and True/False format.
Can Multiple-choice Questions Really Be Useful in Detecting the Abilities of LLMs? (2024.lrec-main)

Copied to clipboard

Challenge: Multiple-choice questions (MCQs) are widely used in the evaluation of large language models (LLMs) however, there are concerns about whether MCQ can truly measure LLM’s capabilities.
Approach: They propose to use multiple choice questions to evaluate large language models (LLMs) to assess their capabilities.
Outcome: The proposed methods show that MCQs are less reliable than LFGQs in terms of expected calibration error.
Arguments that Alter Minds: LLM Rationales Sway Human (and LLM) Notions of Plausibility (2026.acl-long)

Copied to clipboard

Challenge: Experiments with LLMs reveal similar patterns of influence on human plausibility judgments of commonsense benchmark answers.
Approach: They find that human plausibility judgments of commonsense benchmark answers are affected by implausibility arguments for or against an answer.
Outcome: The results show that human judges find LLM rationales convincing and that human annotators agree on the most plausible answer when the plausibility gap is wide.
HellaSwag-Pro: A Large-Scale Bilingual Benchmark for Evaluating the Robustness of LLMs in Commonsense Reasoning (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies show that large language models are robust in commonsense reasoning . however, some variations in questions can lead to incorrect responses .
Approach: They propose a large-scale bilingual benchmark consisting of 11,200 cases . they conduct extensive experiments on 41 representative LLMs .
Outcome: The proposed benchmark systematically evaluates the robustness of large language models in commonsense reasoning.
LLMs May Perform MCQA by Selecting the Least Incorrect Option (2025.coling-main)

Copied to clipboard

Challenge: Multiple Choice Question Answering (MCQA) is a fundamental format for various tasks in NLP, such as commonsense reasoning.
Approach: They propose a method to increase the number of correct options in a dataset.
Outcome: The proposed method improves the performance of multiple choice question answering (MCQA) and improves its accuracy.
Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It (2026.findings-eacl)

Copied to clipboard

Challenge: Using social reasoning benchmarks, we uncover pervasive flaws in both benchmark items and evaluation methodology.
Approach: They audit three widely used social reasoning benchmarks and identify flaws in their design and evaluation methodology.
Outcome: The results challenge the validity of current benchmark-based claims about social reasoning in large language models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations