| Challenge: | Large language models (LLMs) give reasoning before answering, excelling in multiple-choice question answering (MCQA) . but, some studies find that LLMs sans reasoning fail in MCQA without using the question, i.e., choices-only. |
| Approach: | They propose to use reasoning LLMs to separate problematic data from less problematic strategies by examining reasoning traces. |
| Outcome: | The proposed models perform well in multiple-choice question answering without the question, but they fail to use the question. |
Similar Papers
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question? (2024.acl-long)
Copied to clipboard
| Challenge: | Multiple-choice question answering (MCQA) is often used to evaluate large language models . a recent study found that LLMs perform MCQA with choices-only prompts . |
| Approach: | They investigate whether LLMs can perform multiple-choice question answering (MCQA) with choices-only prompts . they find no evidence that the choices- only accuracy stems from memorization alone . |
| Outcome: | The results show that LLMs perform MCQA with choices-only prompts with 0.33 accuracy gain. |
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above (2025.acl-long)
Copied to clipboard
| Challenge: | Multiple choice question answering (MCQA) is popular for LLM evaluation due to its simplicity and human-like testing. |
| Approach: | They argue for a reform of multiple choice question answering (MCQA) they argue for more generative formats based on human testing . |
| Outcome: | The proposed reforms improve the quality of MCQA, the authors argue . they show that even when MCQ is a useful format, its datasets suffer from leakage, unanswerability, shortcuts and saturation. |
Is Large Language Model Performance on Reasoning Tasks Impacted by Different Ways Questions Are Asked? (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing studies on Large Language Models (LLMs) have not investigated the impact of question types on LLM performance. |
| Approach: | They evaluate the performance of five Large Language Models on reasoning tasks . they use quantitative reasoning tasks and deductive reasoning tasks to evaluate the models . |
| Outcome: | The results show that Reasoning accuracy does not correlate with final selection accuracy. |
There’s No Such Thing as Simple Reasoning for LLMs (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing work has focused on relatively complex “many-hop” reasoning problems. |
| Approach: | They analyse the performance of fine-tuned LLMs on simple reasoning problems . they find the models remain highly brittle, being susceptible to seemingly innocent perturbations . |
| Outcome: | The proposed models fail on simple reasoning problems, but are highly brittle . they are susceptible to seemingly innocent perturbations, such as adding duplicates to the set of premises and shuffling the order in which the premises are presented. |
When Models Decide and When They Bind: A Two-Stage Computation for Multiple-Choice Question Answering (2026.findings-acl)
Copied to clipboard
| Challenge: | Multiple-choice question answering (MCQA) is easy to evaluate but adds a meta-task . prior work has shown that language models exhibit selection biases for particular option identifiers such as the label "A" |
| Approach: | They find that option-boundary residual states contain strong linearly decodable signals . winning content position becomes decoded after final option is processed . |
| Outcome: | The proposed model solves the problem and outputs the symbol that represents the answer. |
ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answering (2026.acl-long)
Copied to clipboard
| Challenge: | Recent work in language modeling has led to effective SLMs with impressive performance levels across various benchmarks. |
| Approach: | They propose a benchmark that introduces process-level evaluation for commonsense reasoning tasks. |
| Outcome: | The proposed benchmarks show that large language models provide correct answers despite flawed reasoning processes in a substantial portion of cases. |
An Empirical Study of LLM Reasoning Ability Under Strict Output Length Constraint (2025.emnlp-main)
Copied to clipboard
Yi Sun, Han Wang, Jiaqiang Li, Jiacheng Liu, Xiangyu Li, Hao Wen, Yizhen Yuan, Huiwen Zheng, Yan Liang, Yuanchun Li, Yunxin Liu
| Challenge: | Large Language Models (LLMs) are a powerful tool for test-time scaling, but they are often used under time constraints. |
| Approach: | They propose to use LLMs to make models think before answering questions . they also use self-correction and best-of-N decoding to encourage deeper thinking . |
| Outcome: | The proposed models are able to achieve higher inference accuracy with extra inference computation under time constraints. |
LLMs May Perform MCQA by Selecting the Least Incorrect Option (2025.coling-main)
Copied to clipboard
| Challenge: | Multiple Choice Question Answering (MCQA) is a fundamental format for various tasks in NLP, such as commonsense reasoning. |
| Approach: | They propose a method to increase the number of correct options in a dataset. |
| Outcome: | The proposed method improves the performance of multiple choice question answering (MCQA) and improves its accuracy. |
Logic Haystacks: Probing LLMs’ Long-Context Logical Reasoning (Without Easily Identifiable Unrelated Padding) (2026.eacl-short)
Copied to clipboard
| Challenge: | Recent large language models claim long context windows, but evaluations often involve simple retrieval tasks or synthetic tasks padded with irrelevant text. |
| Approach: | They use grammars to generate simplified English with logical representations to create long input text while controlling its semantics. |
| Outcome: | The proposed model performs better with realistic distractors than with standard models. |
Exploring Automated Distractor Generation for Math Multiple-choice Questions via Large Language Models (2024.findings-naacl)
Copied to clipboard
Wanyong Feng, Jaewook Lee, Hunter McNichols, Alexander Scarlatos, Digory Smith, Simon Woodhead, Nancy Ornelas, Andrew Lan
| Challenge: | Multiple-choice questions (MCQs) are easy to administer and grade . but crafting high-quality distractors remains labor-intensive and limited scalability . |
| Approach: | They propose to automate the generation of distractors in math MCQs by using large language models to generate distractors. |
| Outcome: | The proposed methods can generate valid distractors, but they are less adept at anticipating common errors or misconceptions among real students. |