Challenge: Multiple choice questions (MCQs) are often used in employee selection and training, but their creation is resource-intensive and requires significant effort and investment.
Approach: They propose to use large language models and prompt engineering techniques to automate the generation and validation of MCQs.
Outcome: The proposed system reduces the burden on human resources and enables scalable, cost-effective MCQ generation.

Similar Papers

MCQG-SRefine: Multiple Choice Question Generation and Evaluation with Iterative Self-Critique, Correction, and Comparison Feedback (2025.naacl-long)

Copied to clipboard

Challenge: Generating multiple-choice questions (MCQG) for professional exams is challenging due to outdated knowledge, hallucination issues, and prompt sensitivity.
Approach: They propose a framework for converting medical cases into high-quality USMLE-style questions using a self-refine-based framework.
Outcome: The proposed framework improves human expert satisfaction regarding quality and difficulty of medical questions.
Can Multiple-choice Questions Really Be Useful in Detecting the Abilities of LLMs? (2024.lrec-main)

Copied to clipboard

Challenge: Multiple-choice questions (MCQs) are widely used in the evaluation of large language models (LLMs) however, there are concerns about whether MCQ can truly measure LLM’s capabilities.
Approach: They propose to use multiple choice questions to evaluate large language models (LLMs) to assess their capabilities.
Outcome: The proposed methods show that MCQs are less reliable than LFGQs in terms of expected calibration error.
Exploring Automated Distractor Generation for Math Multiple-choice Questions via Large Language Models (2024.findings-naacl)

Copied to clipboard

Challenge: Multiple-choice questions (MCQs) are easy to administer and grade . but crafting high-quality distractors remains labor-intensive and limited scalability .
Approach: They propose to automate the generation of distractors in math MCQs by using large language models to generate distractors.
Outcome: The proposed methods can generate valid distractors, but they are less adept at anticipating common errors or misconceptions among real students.
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above (2025.acl-long)

Copied to clipboard

Challenge: Multiple choice question answering (MCQA) is popular for LLM evaluation due to its simplicity and human-like testing.
Approach: They argue for a reform of multiple choice question answering (MCQA) they argue for more generative formats based on human testing .
Outcome: The proposed reforms improve the quality of MCQA, the authors argue . they show that even when MCQ is a useful format, its datasets suffer from leakage, unanswerability, shortcuts and saturation.
From Generation to Selection: Findings of Converting Analogical Problem-Solving into Multiple-Choice Questions (2024.findings-emnlp)

Copied to clipboard

Challenge: Abstract and Reasoning Corpus (ARC) is a benchmark designed to evaluate reasoning abilities alone by reducing the amount of prior knowledge and data required to solve the tasks.
Approach: They propose a multiple-choice format suitable for assessing stages like Understand and Apply in Large Language Models (LLMs).
Outcome: The proposed model supports analogical reasoning and evidence analysis, but LLMs use shortcuts in the MC-LARC format.
Revisiting the Self-Consistency Challenges in Multi-Choice Question Formats for Large Language Model Evaluation (2024.lrec-main)

Copied to clipboard

Challenge: Multi-choice questions (MCQs) are a common method for assessing the world knowledge of large language models.
Approach: They propose three knowledge-equivalent question variants to assess LLMs' world knowledge . they propose option position shuffle, option label replacement, and conversion to a True/False format .
Outcome: The proposed questions are shuffle, label replacement, and True/False format.
Selecting Better Samples from Pre-trained LLMs: A Case Study on Question Generation (2023.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated impressive prowess in natural language generation.
Approach: They propose a method to select high-quality questions from LLM-generated candidates using round-trip and prompt-based scoring.
Outcome: The proposed approach can select high-quality questions from a set of LLM-generated candidates without modification of the underlying model nor rely on human annotations.
Benchmarking Critical Questions Generation: A Challenging Reasoning Task for Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Progress in the task of Critical Questions Generation has been hindered by the lack of suitable datasets and automatic evaluation standards.
Approach: They propose a comprehensive approach to support the development and benchmarking of systems for this task.
Outcome: The proposed approach supports the development and benchmarking of systems for this task.
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering (2025.findings-acl)

Copied to clipboard

Challenge: Multiple-choice question answering tasks are one of the most commonly used tasks for evaluating Large Language Models (LLMs).
Approach: They analyze whether existing answer extraction methods are aligned with human judgment and how they are influenced by answer constraints in the prompt across different domains.
Outcome: The proposed evaluation strategies can be inconsistent with human judgment, and can lead to inaccurate and misleading comparisons.
EducationQ: Evaluating LLMs’ Teaching Capabilities Through Multi-Agent Dialogue Framework (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used as educational tools, yet evaluating their teaching capabilities remains challenging due to the resource-intensive nature of teacher-student interactions.
Approach: They propose a multi-agent dialogue framework that efficiently assesses teaching capabilities through simulated dynamic educational scenarios.
Outcome: The proposed framework outperforms open-source models on 1,498 questions across 13 disciplines and 10 difficulty levels on 1,400 questions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations