LAILA: A Large Trait-Based Dataset for Arabic Automated Essay Scoring (2026.eacl-long)
Copied to clipboard
May Bashendy, Walid Massoud, Sohaila Eltanbouly, Salam Albatarni, Marwan Sayed, Abrar Abir, Houda Bouamor, Tamer Elsayed
| Challenge: | Existing Arabic resources are small in scale and lack trait-specific annotations. |
| Approach: | They propose to use LAILA to build a large Arabic AES dataset with holistic and trait-specific annotations of seven writing proficiency traits. |
| Outcome: | The LAILA dataset comprises 7,859 essays annotated with holistic and trait-specific scores on seven dimensions: relevance, organization, vocabulary, style, development, mechanics, and grammar. |
Similar Papers
Automated Essay Scoring: A Reflection on the State of the Art (2024.emnlp-main)
Copied to clipboard
| Challenge: | Automated essay scoring (AES) is a key application of natural language processing . it is based on a holistic score that summarizes the essay's overall quality . |
| Approach: | aaron carroll: automated essay scoring is one of the most important applications in NLP . carroll says the task is still far from being solved, but it's still progressing steadily . he says it'll be interesting to see how researchers can improve performance numbers . |
| Outcome: | a new neural model can beat existing models on a standard evaluation dataset, authors say . authors: the current model is not enough to improve performance numbers . they say it could spark discussion among researchers on how to move forward . |
EssayJudge: A Multi-Granular Benchmark for Assessing Automated Essay Scoring Capabilities of Multimodal Large Language Models (2025.findings-acl)
Copied to clipboard
Jiamin Su, Yibo Yan, Fangteng Fu, Zhang Han, Jingheng Ye, Xiang Liu, Jiahao Huo, Huiyu Zhou, Xuming Hu
| Challenge: | Automated Essay Scoring (AES) systems face three major challenges: reliance on handcrafted features that limit generalizability, difficulty in capturing fine-grained traits like coherence and argumentation, and inability to handle multimodal contexts. |
| Approach: | They propose a multimodal benchmark to evaluate AES capabilities across lexical-, sentence-, and discourse-level traits without manual feature engineering. |
| Outcome: | The proposed system can evaluate AES capabilities across lexical-, sentence-, and discourse-level traits without manual feature engineering. |
Automated Chinese Essay Scoring from Multiple Traits (2022.coling-1)
Copied to clipboard
| Challenge: | Current research on AES focuses on scoring the overall quality or single trait of prompt-specific essays. |
| Approach: | They propose a hierarchical multi-task trait scorer to evaluate quality of writing . they propose an inter-sequence attention mechanism to enhance information interaction . |
| Outcome: | The proposed model outperforms several strong models on ACEA and outperformed other models. |
ICLE++: Modeling Fine-Grained Traits for Holistic Essay Scoring (2024.naacl-long)
Copied to clipboard
| Challenge: | Recent advances in automated essay scoring have limited the generalizability of models trained on ASAP. |
| Approach: | They propose to annotate persuasive student essays with holistic and trait-specific scores in a corpus of persuasive student essay annotated with ICLE++. |
| Outcome: | The proposed model can be used to evaluate models for newer AES problems such as multi-trait scoring and cross-prompt scoring. |
Beyond the Gold Standard in Analytic Automated Essay Scoring (2025.acl-srw)
Copied to clipboard
| Challenge: | Automated Essay Scoring (AES) is a new approach to assessing writing practice . traditional holistic scoring methods are not reliable and lack formative feedback in the classroom. |
| Approach: | They propose to combine analytic and holistic AES to create a system that learns from individual raters instead of gold standard labels. |
| Outcome: | The proposed system learns from individual raters instead of gold standard labels. |
DREsS: Dataset for Rubric-based Essay Scoring on EFL Writing (2025.acl-long)
Copied to clipboard
| Challenge: | Automated essay scoring (AES) is a useful tool in English as a foreign language (EFL) writing education. |
| Approach: | They propose a large-scale, standard dataset for rubric-based automated essay scoring with 48.9K samples in total. |
| Outcome: | The proposed system improves the baseline scores by 45.44%. |
Qayyem: A Real-time Platform for Scoring Proficiency of Arabic Essays (2026.acl-demo)
Copied to clipboard
| Challenge: | Existing Arabic writing technologies primarily use a single quality score for essays, but there is limited support for Arabic AES. |
| Approach: | They propose a Web-based platform that integrates Arabic AES workflows with a user-friendly interface. |
| Outcome: | The proposed system integrates with existing Arabic scoring systems and provides a user-friendly interface. |
Automated Essay Scoring System for Nonnative Japanese Learners (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing systems only provide a holistic score that summarizes the quality of an essay, which provides little feedback for a language learner. |
| Approach: | They developed an automated essay scoring system for Japanese as a second language learners using an essay dataset with annotations for a holistic score and multiple trait scores. |
| Outcome: | The proposed system achieves the highest accuracy in various natural language processing tasks. |
Conundrums in Cross-Prompt Automated Essay Scoring: Making Sense of the State of the Art (2024.acl-long)
Copied to clipboard
| Challenge: | Automated essay scoring (AES) is a task of assigning a single score to an essay . authors abandon sophisticated neural architectures and develop a simple feature-based approach . |
| Approach: | a team of researchers develop a feature-based approach to cross-prompt automated essay scoring that adopts a simple neural architecture. |
| Outcome: | a new approach to cross-prompt automated essay scoring can achieve state-of-the-art results. |
Beyond the Score: Uncertainty-Calibrated LLMs for Automated Essay Assessment (2025.emnlp-main)
Copied to clipboard
| Challenge: | Automated Essay Scoring (AES) systems attain near–human agreement on some public benchmarks, but real-world adoption is limited. |
| Approach: | They propose a distribution-free wrapper that equips any classifier with set-valued outputs enjoying formal coverage guarantees. |
| Outcome: | The proposed model achieves coverage targets while keeping prediction sets compact. |