| Challenge: | Existing studies on rater effects in general settings have not investigated how rater bias affects automated essay scoring. |
| Approach: | They propose to model rater bias by removing essays associated with potentially biased scores from annotated corpus. |
| Outcome: | The proposed model is based on comments provided by raters and is compared with existing corpus. |
Similar Papers
Automated Essay Scoring: A Reflection on the State of the Art (2024.emnlp-main)
Copied to clipboard
| Challenge: | Automated essay scoring (AES) is a key application of natural language processing . it is based on a holistic score that summarizes the essay's overall quality . |
| Approach: | aaron carroll: automated essay scoring is one of the most important applications in NLP . carroll says the task is still far from being solved, but it's still progressing steadily . he says it'll be interesting to see how researchers can improve performance numbers . |
| Outcome: | a new neural model can beat existing models on a standard evaluation dataset, authors say . authors: the current model is not enough to improve performance numbers . they say it could spark discussion among researchers on how to move forward . |
Beyond the Gold Standard in Analytic Automated Essay Scoring (2025.acl-srw)
Copied to clipboard
| Challenge: | Automated Essay Scoring (AES) is a new approach to assessing writing practice . traditional holistic scoring methods are not reliable and lack formative feedback in the classroom. |
| Approach: | They propose to combine analytic and holistic AES to create a system that learns from individual raters instead of gold standard labels. |
| Outcome: | The proposed system learns from individual raters instead of gold standard labels. |
Neural Automated Essay Scoring and Coherence Modeling for Adversarially Crafted Input (N18-1)
Copied to clipboard
| Challenge: | Existing approaches to Automated Essay Scoring (AES) are not well-suited to capture adversarially crafted input of grammatical but incoherent sequences of sentences. |
| Approach: | They propose a neural model of local coherence that can effectively learn connectedness features between sentences. |
| Outcome: | The proposed approach strengthens the validity of neural essay scoring models. |
ASAP++: Enriching the ASAP Automated Essay Grading Dataset with Essay Attribute Scores (L18-1)
Copied to clipboard
| Challenge: | Automated essay grading (AEG) is one of the most challenging activities in natural language processing (NLP). |
| Approach: | They propose to annotate the ASAP AEG dataset and use it to score different attributes of the essays. |
| Outcome: | The proposed resource is based on the ASAP++ dataset, which contains scores for different attributes of the essays, such as content, word choice, organization, sentence fluency, etc. |
Enhancing Automated Essay Scoring Performance via Fine-tuning Pre-trained Language Models with Combination of Regression and Ranking (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Recent work on sentence prediction tasks uses shallow neural networks to learn essay representations and constrain calculated scores with regression loss or ranking loss. |
| Approach: | They propose to use a pre-trained language model to learn text representations first and then to constrain the scores with regression loss or ranking loss. |
| Outcome: | The proposed model outperforms state-of-the-art models on the Automated Student Assessment Prize dataset. |
Can Large Language Models Differentiate Harmful from Argumentative Essays? Steps Toward Ethical Essay Scoring (2025.coling-main)
Copied to clipboard
| Challenge: | Existing automated essay scoring systems overlook ethical and moral aspects of content, erroneously assigning high scores to essays that propagate harmful opinions. |
| Approach: | They introduce a Harmful Essay Detection benchmark to test the effectiveness of various Large Language Models (LLMs) they find that current AES systems overlook ethically and morally problematic elements in essays . |
| Outcome: | The proposed benchmark compared LLMs and AES models to identify and score harmful essays. |
Analytic Automated Essay Scoring Based on Deep Neural Networks Integrating Multidimensional Item Response Theory (2022.coling-1)
Copied to clipboard
| Challenge: | Essay exams have two drawbacks in that grading them is expensive and raises questions about fairness. |
| Approach: | They propose to use a multidimensional item response theory model to improve interpretability while maintaining scoring accuracy. |
| Outcome: | The proposed model improves interpretability while maintaining accuracy while preserving cost and accuracy. |
Cross-Prompt Automated Essay Scoring of Multiple Traits: Making Sense of the State of the Art (2026.acl-long)
Copied to clipboard
| Challenge: | despite recent progress in cross-prompt essay scoring, there is little analysis of what makes a state-of-the-art cross-propert scorer work well. |
| Approach: | They propose to apply transductive learning to cross-prompt scoring for the first time . they propose to train a model that can offer good performance when applied to unseen prompts . |
| Outcome: | The proposed model could be used in the rarely-studied classroom setting without additional training data. |
It’s All Relative: Learning Interpretable Models for Scoring Subjective Bias in Documents from Pairwise Comparisons (2024.eacl-long)
Copied to clipboard
| Challenge: | a new model to score subjective bias in documents is developed to perform pairwise comparisons . a recent study shows that the model can be explained and validated for other domains based on the training data. |
| Approach: | They propose an interpretable model to score subjective bias in Wikipedia articles . they train the model on pairs of revisions of the same Wikipedia article . |
| Outcome: | The proposed model can interpret parameters to discover words most indicative of bias . it compares legal texts, news media and law amendments in three settings . |
Automatic Essay Scoring Incorporating Rating Schema via Reinforcement Learning (D18-1)
Copied to clipboard
| Challenge: | Existing systems for automatic essay scoring are trained to predict the score of each essay at a time without considering rating schema. |
| Approach: | They propose a reinforcement learning framework that incorporates quadratic weighted kappa as guidance to optimize the scoring system. |
| Outcome: | Experiments on benchmark datasets show the proposed framework is effective. |