Papers by Lui Yoshida

1 papers
Are the Reasoning Models Good at Automated Essay Scoring? (2025.findings-emnlp)

Copied to clipboard

Challenge: o3-mini and o4-mini reasoning models perform poorly on automated essay scoring tasks, despite excellent performance on many benchmarks.
Approach: They evaluated OpenAI’s o3-mini and o4-mini reasoning models in automated essay scoring tasks by measuring agreement with expert ratings and consistency in repeated evaluations.
Outcome: The models’ performance on the TOEFL11 dataset is evaluated by measuring agreement with expert ratings and consistency in repeated evaluations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations