Papers by Cesare Aloisi
Two Heads Are Better Than One: Dual-Model Verbal Reflection at Inference-Time (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing LLMs struggle to reliably detect subtle reasoning errors in ASAS tasks. |
| Approach: | They propose a dual-model framework with a dedicated Critic model trained for effective reflection that generates precise verbal feedback. |
| Outcome: | The proposed framework outperforms existing ASAS benchmarks and provides valuable insights into the performance of the proposed framework. |
Distilling ChatGPT for Explainable Automated Student Answer Assessment (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing automated student answer assessment models lack explainable and faithful feedback. |
| Approach: | They propose a framework that leverages ChatGPT for student answer scoring and rationale generation. |
| Outcome: | The proposed method improves the overall QWK score by 11% compared to ChatGPT. |
Calibrating LLMs with Preference Optimization on Thought Trees for Generating Rationale in Science Question Scoring (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for generating rationales that justify scoring decisions are not accurate and often contain hallucinated information. |
| Approach: | They propose a framework capable of generating more faithful rationales and matching performance with classifier-based scoring systems. |
| Outcome: | The proposed framework achieves 38% improvement in QWK score compared to prior work . it can be used to match performance with classifier-based scoring systems . |
LearnLens: LLM-Enabled Personalised, Curriculum-Grounded Feedback with Educators in the Loop (2025.emnlp-demos)
Copied to clipboard
| Challenge: | Existing systems that provide personalised, curriculum-aligned feedback are time-intensive and time-consuming. |
| Approach: | They propose a modular, LLM-based system that generates personalised, curriculum-aligned feedback in science education. |
| Outcome: | The proposed system generates personalised, curriculum-aligned feedback in science education. |
AERA Chat: An Interactive Platform for Automated Explainable Student Answer Assessment (2025.emnlp-demos)
Copied to clipboard
| Challenge: | Existing systems that use pretrained language models to score student answers are noisy and unreliable. |
| Approach: | They propose a visualization platform for automated student answer assessment that leverages multiple LLMs to generate rationales. |
| Outcome: | The proposed platform enables educators to mark tasks and researchers to evaluate rationale quality from different models. |