Challenge: Large language models (LLMs) give reasoning before answering, excelling in multiple-choice question answering (MCQA) . but, some studies find that LLMs sans reasoning fail in MCQA without using the question, i.e., choices-only.
Approach: They propose to use reasoning LLMs to separate problematic data from less problematic strategies by examining reasoning traces.
Outcome: The proposed models perform well in multiple-choice question answering without the question, but they fail to use the question.

Similar Papers

Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question? (2024.acl-long)

Copied to clipboard

Challenge: Multiple-choice question answering (MCQA) is often used to evaluate large language models . a recent study found that LLMs perform MCQA with choices-only prompts .
Approach: They investigate whether LLMs can perform multiple-choice question answering (MCQA) with choices-only prompts . they find no evidence that the choices- only accuracy stems from memorization alone .
Outcome: The results show that LLMs perform MCQA with choices-only prompts with 0.33 accuracy gain.
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above (2025.acl-long)

Copied to clipboard

Challenge: Multiple choice question answering (MCQA) is popular for LLM evaluation due to its simplicity and human-like testing.
Approach: They argue for a reform of multiple choice question answering (MCQA) they argue for more generative formats based on human testing .
Outcome: The proposed reforms improve the quality of MCQA, the authors argue . they show that even when MCQ is a useful format, its datasets suffer from leakage, unanswerability, shortcuts and saturation.
Is Large Language Model Performance on Reasoning Tasks Impacted by Different Ways Questions Are Asked? (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies on Large Language Models (LLMs) have not investigated the impact of question types on LLM performance.
Approach: They evaluate the performance of five Large Language Models on reasoning tasks . they use quantitative reasoning tasks and deductive reasoning tasks to evaluate the models .
Outcome: The results show that Reasoning accuracy does not correlate with final selection accuracy.
There’s No Such Thing as Simple Reasoning for LLMs (2025.findings-acl)

Copied to clipboard

Challenge: Existing work has focused on relatively complex “many-hop” reasoning problems.
Approach: They analyse the performance of fine-tuned LLMs on simple reasoning problems . they find the models remain highly brittle, being susceptible to seemingly innocent perturbations .
Outcome: The proposed models fail on simple reasoning problems, but are highly brittle . they are susceptible to seemingly innocent perturbations, such as adding duplicates to the set of premises and shuffling the order in which the premises are presented.
When Models Decide and When They Bind: A Two-Stage Computation for Multiple-Choice Question Answering (2026.findings-acl)

Copied to clipboard

Challenge: Multiple-choice question answering (MCQA) is easy to evaluate but adds a meta-task . prior work has shown that language models exhibit selection biases for particular option identifiers such as the label "A"
Approach: They find that option-boundary residual states contain strong linearly decodable signals . winning content position becomes decoded after final option is processed .
Outcome: The proposed model solves the problem and outputs the symbol that represents the answer.
ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answering (2026.acl-long)

Copied to clipboard

Challenge: Recent work in language modeling has led to effective SLMs with impressive performance levels across various benchmarks.
Approach: They propose a benchmark that introduces process-level evaluation for commonsense reasoning tasks.
Outcome: The proposed benchmarks show that large language models provide correct answers despite flawed reasoning processes in a substantial portion of cases.
An Empirical Study of LLM Reasoning Ability Under Strict Output Length Constraint (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are a powerful tool for test-time scaling, but they are often used under time constraints.
Approach: They propose to use LLMs to make models think before answering questions . they also use self-correction and best-of-N decoding to encourage deeper thinking .
Outcome: The proposed models are able to achieve higher inference accuracy with extra inference computation under time constraints.
LLMs May Perform MCQA by Selecting the Least Incorrect Option (2025.coling-main)

Copied to clipboard

Challenge: Multiple Choice Question Answering (MCQA) is a fundamental format for various tasks in NLP, such as commonsense reasoning.
Approach: They propose a method to increase the number of correct options in a dataset.
Outcome: The proposed method improves the performance of multiple choice question answering (MCQA) and improves its accuracy.
Logic Haystacks: Probing LLMs’ Long-Context Logical Reasoning (Without Easily Identifiable Unrelated Padding) (2026.eacl-short)

Copied to clipboard

Challenge: Recent large language models claim long context windows, but evaluations often involve simple retrieval tasks or synthetic tasks padded with irrelevant text.
Approach: They use grammars to generate simplified English with logical representations to create long input text while controlling its semantics.
Outcome: The proposed model performs better with realistic distractors than with standard models.
Exploring Automated Distractor Generation for Math Multiple-choice Questions via Large Language Models (2024.findings-naacl)

Copied to clipboard

Challenge: Multiple-choice questions (MCQs) are easy to administer and grade . but crafting high-quality distractors remains labor-intensive and limited scalability .
Approach: They propose to automate the generation of distractors in math MCQs by using large language models to generate distractors.
Outcome: The proposed methods can generate valid distractors, but they are less adept at anticipating common errors or misconceptions among real students.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations