Papers by Lui Yoshida
Are the Reasoning Models Good at Automated Essay Scoring? (2025.findings-emnlp)
Copied to clipboard
| Challenge: | o3-mini and o4-mini reasoning models perform poorly on automated essay scoring tasks, despite excellent performance on many benchmarks. |
| Approach: | They evaluated OpenAI’s o3-mini and o4-mini reasoning models in automated essay scoring tasks by measuring agreement with expert ratings and consistency in repeated evaluations. |
| Outcome: | The models’ performance on the TOEFL11 dataset is evaluated by measuring agreement with expert ratings and consistency in repeated evaluations. |