| Challenge: | Automated essay scoring (AES) is a key application of natural language processing . it is based on a holistic score that summarizes the essay's overall quality . |
| Approach: | aaron carroll: automated essay scoring is one of the most important applications in NLP . carroll says the task is still far from being solved, but it's still progressing steadily . he says it'll be interesting to see how researchers can improve performance numbers . |
| Outcome: | a new neural model can beat existing models on a standard evaluation dataset, authors say . authors: the current model is not enough to improve performance numbers . they say it could spark discussion among researchers on how to move forward . |
Similar Papers
Conundrums in Cross-Prompt Automated Essay Scoring: Making Sense of the State of the Art (2024.acl-long)
Copied to clipboard
| Challenge: | Automated essay scoring (AES) is a task of assigning a single score to an essay . authors abandon sophisticated neural architectures and develop a simple feature-based approach . |
| Approach: | a team of researchers develop a feature-based approach to cross-prompt automated essay scoring that adopts a simple neural architecture. |
| Outcome: | a new approach to cross-prompt automated essay scoring can achieve state-of-the-art results. |
Beyond the Gold Standard in Analytic Automated Essay Scoring (2025.acl-srw)
Copied to clipboard
| Challenge: | Automated Essay Scoring (AES) is a new approach to assessing writing practice . traditional holistic scoring methods are not reliable and lack formative feedback in the classroom. |
| Approach: | They propose to combine analytic and holistic AES to create a system that learns from individual raters instead of gold standard labels. |
| Outcome: | The proposed system learns from individual raters instead of gold standard labels. |
Cross-Prompt Automated Essay Scoring of Multiple Traits: Making Sense of the State of the Art (2026.acl-long)
Copied to clipboard
| Challenge: | despite recent progress in cross-prompt essay scoring, there is little analysis of what makes a state-of-the-art cross-propert scorer work well. |
| Approach: | They propose to apply transductive learning to cross-prompt scoring for the first time . they propose to train a model that can offer good performance when applied to unseen prompts . |
| Outcome: | The proposed model could be used in the rarely-studied classroom setting without additional training data. |
Neural Automated Essay Scoring and Coherence Modeling for Adversarially Crafted Input (N18-1)
Copied to clipboard
| Challenge: | Existing approaches to Automated Essay Scoring (AES) are not well-suited to capture adversarially crafted input of grammatical but incoherent sequences of sentences. |
| Approach: | They propose a neural model of local coherence that can effectively learn connectedness features between sentences. |
| Outcome: | The proposed approach strengthens the validity of neural essay scoring models. |
Enhancing Automated Essay Scoring Performance via Fine-tuning Pre-trained Language Models with Combination of Regression and Ranking (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Recent work on sentence prediction tasks uses shallow neural networks to learn essay representations and constrain calculated scores with regression loss or ranking loss. |
| Approach: | They propose to use a pre-trained language model to learn text representations first and then to constrain the scores with regression loss or ranking loss. |
| Outcome: | The proposed model outperforms state-of-the-art models on the Automated Student Assessment Prize dataset. |
Automated Essay Scoring in the Presence of Biased Ratings (N18-1)
Copied to clipboard
| Challenge: | Existing studies on rater effects in general settings have not investigated how rater bias affects automated essay scoring. |
| Approach: | They propose to model rater bias by removing essays associated with potentially biased scores from annotated corpus. |
| Outcome: | The proposed model is based on comments provided by raters and is compared with existing corpus. |
EssayJudge: A Multi-Granular Benchmark for Assessing Automated Essay Scoring Capabilities of Multimodal Large Language Models (2025.findings-acl)
Copied to clipboard
Jiamin Su, Yibo Yan, Fangteng Fu, Zhang Han, Jingheng Ye, Xiang Liu, Jiahao Huo, Huiyu Zhou, Xuming Hu
| Challenge: | Automated Essay Scoring (AES) systems face three major challenges: reliance on handcrafted features that limit generalizability, difficulty in capturing fine-grained traits like coherence and argumentation, and inability to handle multimodal contexts. |
| Approach: | They propose a multimodal benchmark to evaluate AES capabilities across lexical-, sentence-, and discourse-level traits without manual feature engineering. |
| Outcome: | The proposed system can evaluate AES capabilities across lexical-, sentence-, and discourse-level traits without manual feature engineering. |
Reproduction and Replication: A Case Study with Automatic Essay Scoring (2020.lrec-1)
Copied to clipboard
| Challenge: | reproducibility of experiments has gained more attention in the NLP community . recent negative reproduction results indicate that published results are not verifiable . |
| Approach: | They propose to reproduce an earlier study of automatic essay scoring for determining the proficiency of second language learners in a multilingual setting. |
| Outcome: | The proposed reproduction of an AES system for determining the proficiency of second language learners in a multilingual setting is compared with the original. |
Automated Scoring: Beyond Natural Language Processing (C18-1)
Copied to clipboard
| Challenge: | In this paper, we argue that building operational automated scoring systems is a task that has disciplinary complexity above and beyond competitive shared tasks. |
| Approach: | They argue that building operational automated scoring systems is a task that has disciplinary complexity above and beyond standard competitive shared tasks . they argue that it is essential for us as NLP researchers to understand and incorporate these perspectives in our research and work towards a mutually satisfactory solution . |
| Outcome: | The proposed approach is based on the findings of a recent conference on automated scoring. |
Analytic Automated Essay Scoring Based on Deep Neural Networks Integrating Multidimensional Item Response Theory (2022.coling-1)
Copied to clipboard
| Challenge: | Essay exams have two drawbacks in that grading them is expensive and raises questions about fairness. |
| Approach: | They propose to use a multidimensional item response theory model to improve interpretability while maintaining scoring accuracy. |
| Outcome: | The proposed model improves interpretability while maintaining accuracy while preserving cost and accuracy. |