Give Me More Feedback II: Annotating Thesis Strength and Related Attributes in Student Essays (P19-1)
Copied to clipboard
| Challenge: | Existing work on automated essay scoring has focused on holistic scoring, but there is limited annotated corpus of essays with thesis strength scores. |
| Approach: | They propose a scoring rubric for persuasive essay quality and annotate corpus of essays with thesis strength scores. |
| Outcome: | The proposed scoring rubric could provide feedback to students on why essay gets thesis strength score . the rubric can be used to score persuasive essay quality, thesis strength, and organization . |
Similar Papers
Give Me More Feedback: Annotating Argument Persuasiveness and Related Attributes in Student Essays (P18-1)
Copied to clipboard
| Challenge: | Existing work on automated essay scoring has focused on holistic scoring, which summarizes the quality of an essay with a single score. |
| Approach: | They present a corpus of essays simultaneously annotated with argument components, argument persuasiveness scores, and attributes of argument components that impact an argument’s persuasiveness. |
| Outcome: | The proposed corpus could trigger the development of novel computational models that provide useful feedback to students on why their arguments are (un)persuasive . |
ASAP++: Enriching the ASAP Automated Essay Grading Dataset with Essay Attribute Scores (L18-1)
Copied to clipboard
| Challenge: | Automated essay grading (AEG) is one of the most challenging activities in natural language processing (NLP). |
| Approach: | They propose to annotate the ASAP AEG dataset and use it to score different attributes of the essays. |
| Outcome: | The proposed resource is based on the ASAP++ dataset, which contains scores for different attributes of the essays, such as content, word choice, organization, sentence fluency, etc. |
FeedEval: Pedagogically Aligned Evaluation of LLM-Generated Essay Feedback (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent research emphasizes the generation of high-quality feedback that provides justification and actionable guidance. |
| Approach: | They propose an LLM-based framework for evaluating LLM feedback along three dimensions: specificity, helpfulness, and validity. |
| Outcome: | The proposed framework evaluates LLM-generated feedback along three dimensions: specificity, helpfulness, and validity. |
Automated Essay Scoring: A Reflection on the State of the Art (2024.emnlp-main)
Copied to clipboard
| Challenge: | Automated essay scoring (AES) is a key application of natural language processing . it is based on a holistic score that summarizes the essay's overall quality . |
| Approach: | aaron carroll: automated essay scoring is one of the most important applications in NLP . carroll says the task is still far from being solved, but it's still progressing steadily . he says it'll be interesting to see how researchers can improve performance numbers . |
| Outcome: | a new neural model can beat existing models on a standard evaluation dataset, authors say . authors: the current model is not enough to improve performance numbers . they say it could spark discussion among researchers on how to move forward . |
Cross-Prompt Automated Essay Scoring of Multiple Traits: Making Sense of the State of the Art (2026.acl-long)
Copied to clipboard
| Challenge: | despite recent progress in cross-prompt essay scoring, there is little analysis of what makes a state-of-the-art cross-propert scorer work well. |
| Approach: | They propose to apply transductive learning to cross-prompt scoring for the first time . they propose to train a model that can offer good performance when applied to unseen prompts . |
| Outcome: | The proposed model could be used in the rarely-studied classroom setting without additional training data. |
Dataset and Baseline for Automatic Student Feedback Analysis (2022.lrec-1)
Copied to clipboard
| Challenge: | Currently, student feedback is collected manually, but it does not indicate the student's opinion on different aspects of the teaching/learning process. |
| Approach: | They propose to annotate student feedback corpus which contains 3000 instances . they propose a hierarchical taxonomy for aspect categorization, which covers all areas . |
| Outcome: | The proposed model can be used for aspects analysis, document level sentiment analysis and document level analysis. |
ICLE++: Modeling Fine-Grained Traits for Holistic Essay Scoring (2024.naacl-long)
Copied to clipboard
| Challenge: | Recent advances in automated essay scoring have limited the generalizability of models trained on ASAP. |
| Approach: | They propose to annotate persuasive student essays with holistic and trait-specific scores in a corpus of persuasive student essay annotated with ICLE++. |
| Outcome: | The proposed model can be used to evaluate models for newer AES problems such as multi-trait scoring and cross-prompt scoring. |
Score It All Together: A Multi-Task Learning Study on Automatic Scoring of Argumentative Essays (2023.findings-acl)
Copied to clipboard
| Challenge: | a multi-task learning approach outperforms sequential approaches for scoring argumentative essays . segmentation and classification of argumentative elements are important steps towards providing feedback on writing structure, but assessing the quality of arguments is less researched . |
| Approach: | They use a student essay dataset to study how argumentative essays are scored . they use automated span detection, type and quality prediction to combine these tasks . |
| Outcome: | The proposed method outperforms sequential approaches for segmentation and quality prediction. |
Let’s discuss! Quality Dimensions and Annotated Datasets for Computational Argument Quality Assessment (2024.emnlp-main)
Copied to clipboard
| Challenge: | Argumentation is a key competence and an important cultural technique in democratic societies. |
| Approach: | They propose to create domain-specific datasets and methods to assess argument quality. |
| Outcome: | The proposed methods address gaps in the literature and aid future research in the domain. |
Automated Essay Scoring in the Presence of Biased Ratings (N18-1)
Copied to clipboard
| Challenge: | Existing studies on rater effects in general settings have not investigated how rater bias affects automated essay scoring. |
| Approach: | They propose to model rater bias by removing essays associated with potentially biased scores from annotated corpus. |
| Outcome: | The proposed model is based on comments provided by raters and is compared with existing corpus. |