Papers with WiCkeD
WiCkeD: A Simple Method to Make Multiple Choice Benchmarks More Challenging (2025.acl-short)
Copied to clipboard
| Challenge: | Multiple choice question (MCQ) benchmarks are widely used to evaluate Large Language Models (LLMs). |
| Approach: | They propose a method to increase the complexity of existing multiple-choice benchmarks by randomly replacing a choice with “None of the above”. |
| Outcome: | The proposed method can be applied to 6 popular benchmarks and evaluate 18 open-weight LLMs. |