Papers with MRBench
Unifying AI Tutor Evaluation: An Evaluation Taxonomy for Pedagogical Ability Assessment of LLM-Powered AI Tutors (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing evaluations of large language models have been limited to subjective protocols and benchmarks. |
| Approach: | They propose a unified evaluation taxonomy with eight pedagogical dimensions based on key learning sciences principles to assess the pedagical value of LLM-powered AI tutor responses grounded in student mistakes or confusions in the mathematical domain. |
| Outcome: | The proposed taxonomy, benchmark, and human-annotated labels will streamline the evaluation process and help track the progress in AI tutors’ development. |
Memory-Driven Role-Playing: Evaluation and Enhancement of Persona Knowledge Utilization in LLMs (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing models fail to recall and accurately apply designated persona knowledge without explicit cues . memory-driven role-playing paradigms are attracting significant interest . |
| Approach: | They propose a memory-driven role-playing paradigm that frames persona knowledge as the LLM's internal memory store and a prompting architecture that guides structured memory retrieval and response generation. |
| Outcome: | The proposed paradigm provides a comprehensive diagnostic for four-stage role-playing abilities across 12 LLMs. |