Simulated Students in Tutoring Dialogues: Substance or Illusion? (2026.acl-long)
Copied to clipboard
| Challenge: | evaluating the effectiveness of new technology requires real students, which is time-consuming and hard to scale up. |
| Approach: | They propose to define the student simulation task and benchmark a wide range of student simulation methods on these metrics. |
| Outcome: | The proposed evaluation metrics show that prompting strategies perform poorly on a real-world tutoring dialogue dataset. |
Similar Papers
Real or Robotic? Assessing Whether LLMs Accurately Simulate Qualities of Human Responses in Human-LLM Dialogue (2026.findings-acl)
Copied to clipboard
Jonathan Ivey, Shivani Kumar, Jiayu Liu, Hua Shen, Sushrita Rakshit, Rohan Raju, Haotian Zhang, Aparna Ananthasubramaniam, Junghwan Kim, Bowen Yi, Dustin Wright, Abraham Israeli, Anders Giovanni Møller, Lechen Zhang, David Jurgens
| Challenge: | Recent work has sought to use large language models to simulate human-human and human-LLM interactions. |
| Approach: | They use a large-scale dataset to generate a paired LLM-LLM and human-LLm dialogues from the WildChat dataset and quantify how well they align with their human counterparts. |
| Outcome: | The proposed models perform similarly in simulating English, Chinese, and Russian dialogues. |
Using LLMs to simulate students’ responses to exam questions (2024.findings-emnlp)
Copied to clipboard
Luca Benedetto, Giovanni Aradelli, Antonia Donvito, Alberto Lucchetti, Andrea Cappelli, Paula Buttery
| Challenge: | Existing studies have used Large Language Models to simulate students answering exam questions . a proposed prompt for GPT-3.5 is not suitable for all LLMs, and there is no correlation between the quality of the rationales obtained with the model and the accuracy of the student simulation task. |
| Approach: | They propose a large language model prompt engineered for GPT-3.5 that can be used to answer exam questions simulating students of different skill levels. |
| Outcome: | The proposed prompt is robust to different educational domains and generalise to data unseen during prompt engineering phase. |
Student Data Paradox and Curious Case of Single Student-Tutor Model: Regressive Side Effects of Training LLMs for Personalized Learning (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are being developed to provide personalized tutoring systems that can understand and adapt to individual student needs. |
| Approach: | They propose to train large language models on student-tutor dialogue datasets to understand student behavior and evaluate their performance across multiple benchmarks. |
| Outcome: | The proposed model performance declines across multiple benchmarks, indicating a broad impact on their capabilities when trained to model student behavior. |
MathDial: A Dialogue Tutoring Dataset with Rich Pedagogical Properties Grounded in Math Reasoning Problems (2023.findings-emnlp)
Copied to clipboard
Jakub Macina, Nico Daheim, Sankalan Chowdhury, Tanmay Sinha, Manu Kapur, Iryna Gurevych, Mrinmaya Sachan
| Challenge: | Existing models for automatic dialogue tutoring fail to provide accurate feedback or reveal solutions to students too early. |
| Approach: | They propose a framework to generate one-to-one teacher-student tutoring dialogues by pairing human teachers with a Large Language Model (LLM) they use scaffolding questions and annotations to fine-tune models to be more effective tutors . |
| Outcome: | The proposed framework can generate 3k one-to-one teacher-student tutoring dialogues grounded in multi-step math reasoning problems. |
Stepwise Verification and Remediation of Student Reasoning Errors with Large Language Model Tutors (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing models for dialog tutoring fail to detect student errors and tailor their feedback to them. |
| Approach: | They propose to build dialog tutoring models to scaffold students' problem-solving and verify student solutions by using automatic and human evaluation. |
| Outcome: | The proposed model improves the quality of the tutor response generation by detecting student errors and adjusting the feedback to the errors. |
Embracing Imperfection: Simulating Students with Diverse Cognitive Levels Using LLM-based Agents (2025.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) are becoming increasingly popular in education, enabling researchers to simulate students' learning patterns and learning patterns. |
| Approach: | They propose a training-free framework for student simulation that takes into account student cognitive diversity and realism. |
| Outcome: | The proposed model outperforms baseline models and achieves 100% improvement in simulation accuracy and realism. |
Unifying AI Tutor Evaluation: An Evaluation Taxonomy for Pedagogical Ability Assessment of LLM-Powered AI Tutors (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing evaluations of large language models have been limited to subjective protocols and benchmarks. |
| Approach: | They propose a unified evaluation taxonomy with eight pedagogical dimensions based on key learning sciences principles to assess the pedagical value of LLM-powered AI tutor responses grounded in student mistakes or confusions in the mathematical domain. |
| Outcome: | The proposed taxonomy, benchmark, and human-annotated labels will streamline the evaluation process and help track the progress in AI tutors’ development. |
Closing the Loop: Learning to Generate Writing Feedback via Language Model Simulated Student Revisions (2024.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in language models (LMs) have made it possible to automatically generate feedback that is actionable and well-aligned with human-specified attributes. |
| Approach: | They propose a tool that PROduces Feedback via learning from LM simulated student revisions and propose to iteratively optimize the feedback generator by directly maximizing the effectiveness of students’ overall revising performance. |
| Outcome: | The proposed approach surpasses baseline methods in effectiveness of improving students’ writing and demonstrates enhanced pedagogical values, even though it was not explicitly trained for this aspect. |
Position: LLMs Can be Good Tutors in English Education (2025.emnlp-main)
Copied to clipboard
Jingheng Ye, Shen Wang, Deqing Zou, Yibo Yan, Kun Wang, Hai-Tao Zheng, Ruitong Liu, Zenglin Xu, Irwin King, Philip S. Yu, Qingsong Wen
| Challenge: | Recent efforts to integrate large language models into English education lack adaptability to language learning. |
| Approach: | They argue that large language models can be effective tutors in English education . they encourage interdisciplinary research to explore these roles, fostering innovation and risks . |
| Outcome: | The proposed models can play three critical roles: 1) as data enhancers, 2) as task predictors, 3) as agents, enabling personalized and inclusive education. |
One LLM Does Not Simulate All Students: Ability-Aware Student Simulation via Cognitive Diagnosis Guided LLM Assignment (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods rely on a single high-capacity LLM to represent an entire population of diverse learners. |
| Approach: | They propose an ability-aware student simulation framework that matches students with appropriate LLM backbones through cognitive alignment. |
| Outcome: | The proposed framework significantly reduces simulation bias and outperforms single-model baselines across the entire proficiency spectrum. |