Challenge: evaluating the effectiveness of new technology requires real students, which is time-consuming and hard to scale up.
Approach: They propose to define the student simulation task and benchmark a wide range of student simulation methods on these metrics.
Outcome: The proposed evaluation metrics show that prompting strategies perform poorly on a real-world tutoring dialogue dataset.

Similar Papers

Real or Robotic? Assessing Whether LLMs Accurately Simulate Qualities of Human Responses in Human-LLM Dialogue (2026.findings-acl)

Copied to clipboard

Challenge: Recent work has sought to use large language models to simulate human-human and human-LLM interactions.
Approach: They use a large-scale dataset to generate a paired LLM-LLM and human-LLm dialogues from the WildChat dataset and quantify how well they align with their human counterparts.
Outcome: The proposed models perform similarly in simulating English, Chinese, and Russian dialogues.
Using LLMs to simulate students’ responses to exam questions (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have used Large Language Models to simulate students answering exam questions . a proposed prompt for GPT-3.5 is not suitable for all LLMs, and there is no correlation between the quality of the rationales obtained with the model and the accuracy of the student simulation task.
Approach: They propose a large language model prompt engineered for GPT-3.5 that can be used to answer exam questions simulating students of different skill levels.
Outcome: The proposed prompt is robust to different educational domains and generalise to data unseen during prompt engineering phase.
Student Data Paradox and Curious Case of Single Student-Tutor Model: Regressive Side Effects of Training LLMs for Personalized Learning (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are being developed to provide personalized tutoring systems that can understand and adapt to individual student needs.
Approach: They propose to train large language models on student-tutor dialogue datasets to understand student behavior and evaluate their performance across multiple benchmarks.
Outcome: The proposed model performance declines across multiple benchmarks, indicating a broad impact on their capabilities when trained to model student behavior.
MathDial: A Dialogue Tutoring Dataset with Rich Pedagogical Properties Grounded in Math Reasoning Problems (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing models for automatic dialogue tutoring fail to provide accurate feedback or reveal solutions to students too early.
Approach: They propose a framework to generate one-to-one teacher-student tutoring dialogues by pairing human teachers with a Large Language Model (LLM) they use scaffolding questions and annotations to fine-tune models to be more effective tutors .
Outcome: The proposed framework can generate 3k one-to-one teacher-student tutoring dialogues grounded in multi-step math reasoning problems.
Stepwise Verification and Remediation of Student Reasoning Errors with Large Language Model Tutors (2024.emnlp-main)

Copied to clipboard

Challenge: Existing models for dialog tutoring fail to detect student errors and tailor their feedback to them.
Approach: They propose to build dialog tutoring models to scaffold students' problem-solving and verify student solutions by using automatic and human evaluation.
Outcome: The proposed model improves the quality of the tutor response generation by detecting student errors and adjusting the feedback to the errors.
Embracing Imperfection: Simulating Students with Diverse Cognitive Levels Using LLM-based Agents (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are becoming increasingly popular in education, enabling researchers to simulate students' learning patterns and learning patterns.
Approach: They propose a training-free framework for student simulation that takes into account student cognitive diversity and realism.
Outcome: The proposed model outperforms baseline models and achieves 100% improvement in simulation accuracy and realism.
Unifying AI Tutor Evaluation: An Evaluation Taxonomy for Pedagogical Ability Assessment of LLM-Powered AI Tutors (2025.naacl-long)

Copied to clipboard

Challenge: Existing evaluations of large language models have been limited to subjective protocols and benchmarks.
Approach: They propose a unified evaluation taxonomy with eight pedagogical dimensions based on key learning sciences principles to assess the pedagical value of LLM-powered AI tutor responses grounded in student mistakes or confusions in the mathematical domain.
Outcome: The proposed taxonomy, benchmark, and human-annotated labels will streamline the evaluation process and help track the progress in AI tutors’ development.
Closing the Loop: Learning to Generate Writing Feedback via Language Model Simulated Student Revisions (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in language models (LMs) have made it possible to automatically generate feedback that is actionable and well-aligned with human-specified attributes.
Approach: They propose a tool that PROduces Feedback via learning from LM simulated student revisions and propose to iteratively optimize the feedback generator by directly maximizing the effectiveness of students’ overall revising performance.
Outcome: The proposed approach surpasses baseline methods in effectiveness of improving students’ writing and demonstrates enhanced pedagogical values, even though it was not explicitly trained for this aspect.
Position: LLMs Can be Good Tutors in English Education (2025.emnlp-main)

Copied to clipboard

Challenge: Recent efforts to integrate large language models into English education lack adaptability to language learning.
Approach: They argue that large language models can be effective tutors in English education . they encourage interdisciplinary research to explore these roles, fostering innovation and risks .
Outcome: The proposed models can play three critical roles: 1) as data enhancers, 2) as task predictors, 3) as agents, enabling personalized and inclusive education.
One LLM Does Not Simulate All Students: Ability-Aware Student Simulation via Cognitive Diagnosis Guided LLM Assignment (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods rely on a single high-capacity LLM to represent an entire population of diverse learners.
Approach: They propose an ability-aware student simulation framework that matches students with appropriate LLM backbones through cognitive alignment.
Outcome: The proposed framework significantly reduces simulation bias and outperforms single-model baselines across the entire proficiency spectrum.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations