Assessing Online Writing Feedback Resources: Generative AI vs. Good Samaritans (2024.lrec-main)
Copied to clipboard
| Challenge: | Providing constructive feedback on student essays presents significant challenges . large language models (LLMs) such as ChatGPT can facilitate this process . |
| Approach: | They compare essayforum.com and large language models such as ChatGPT for students . they argue that both can mutually reinforce each other and provide constructive feedback . |
| Outcome: | The findings highlight the potential of AI in advancing the field of automated essay evaluation. |
Similar Papers
Can Large Language Models Automatically Score Proficiency of Written Essays? (2024.lrec-main)
Copied to clipboard
| Challenge: | Automated essay scoring (AES) is one of the earliest research problems in natural language processing. |
| Approach: | They propose to use large language models to analyze and score written essays using four different prompts. |
| Outcome: | The proposed models show comparable performance on four different prompts and a slight advantage over the state-of-the-art models. |
LEAF: Language Learners’ English Essays and Feedback Corpus (2024.naacl-short)
Copied to clipboard
| Challenge: | Current automated essay scoring models lack the granularity desired by learners and instructors seeking more detailed insights. |
| Approach: | They present a corpus of English essays and their corresponding feedback from the “essayforum” website. |
| Outcome: | The LEAF corpus provides valuable feedback for students and teachers . it provides insights on argumentative aspects and organizational coherence . |
LLMs can Perform Multi-Dimensional Analytic Writing Assessments: A Case Study of L2 Graduate-Level Academic English Writing (2025.acl-long)
Copied to clipboard
| Challenge: | a growing number of studies have indicated the general usefulness of LLMs for automated writing assessments. |
| Approach: | They propose a framework that evaluates LLMs' ability to provide scores and comments based on multiple assessment criteria. |
| Outcome: | The proposed framework is interpretable, cost-efficient, scalable, and reproducible . it is compared to existing methods that rely on manual judgments . |
Credible without Credit: Domain Experts Assess Generative Language Models (2023.acl-short)
Copied to clipboard
| Challenge: | ChatGPT has been criticized for its lack of accuracy and coherence . authors argue that language models could replace search engines and make college essays obsolete . |
| Approach: | a team of 10 domain experts conducts an initial assessment of language models using 100 expert-written questions. |
| Outcome: | The results show that language models are mixed in their accuracy. |
Help Me Write a Story: Evaluating LLMs’ Ability to Generate Writing Feedback (2025.acl-long)
Copied to clipboard
| Challenge: | Current models provide specific and mostly accurate writing feedback, but they fail to identify the biggest writing issue in the story and to correctly decide when to offer critical vs. positive feedback. |
| Approach: | They propose a task that corrupts 1,300 stories to intentionally introduce writing issues to study model performance. |
| Outcome: | The proposed model performs well in a controlled task with human and automatic evaluation metrics. |
Unleashing Large Language Models’ Proficiency in Zero-shot Essay Scoring (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in automated essay scoring (AES) have relied on labeled essays, requiring tremendous cost and expertise for their acquisition. |
| Approach: | They propose a zero-shot prompting framework that automatically decomposes writing proficiency into distinct traits and generates scoring criteria for each trait. |
| Outcome: | The proposed framework outperforms straightforward prompting (Vanilla) on TOEFL11 and ASAP, while the small-sized Llama2-13b-chat significantly outperformed ChatGPT. |
FEAT-writing: An Interactive Training System for Argumentative Writing (2025.coling-demos)
Copied to clipboard
| Challenge: | Argumentative writing is a critical skill for academic success, but many students struggle to develop these skills. |
| Approach: | They developed an online system that provides students with automated feedback and exercises for argumentative writing. |
| Outcome: | The proposed system improves argumentative writing quality among native English speakers and english-as-a-foreign-language learners. |
IFlyEA: A Chinese Essay Assessment System with Automated Rating, Review Generation, and Recommendation (2021.acl-demo)
Copied to clipboard
| Challenge: | Automated Essay Assessment (AEA) aims to judge students’ writing proficiency in an automatic way. |
| Approach: | They propose to use Chinese AEA system IFlyEssayAssess to evaluate essays written by native Chinese students from primary and junior schools. |
| Outcome: | The proposed system provides application services for essay scoring, review generation, recommendation, and explainable analytical visualization. |
Machine Learning–Driven Language Assessment (2020.tacl-1)
Copied to clipboard
| Challenge: | Language proficiency tests are cumbersome to create and maintain, and items may be copied and leaked or simply used too often. |
| Approach: | They propose a method that uses machine learning and natural language processing to induce proficiency scales and linguistic models to estimate item difficulty directly for computer-adaptive testing. |
| Outcome: | The proposed method produces scores that are reliable and reliable while generating item banks large enough to satisfy security requirements. |
FeedEval: Pedagogically Aligned Evaluation of LLM-Generated Essay Feedback (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent research emphasizes the generation of high-quality feedback that provides justification and actionable guidance. |
| Approach: | They propose an LLM-based framework for evaluating LLM feedback along three dimensions: specificity, helpfulness, and validity. |
| Outcome: | The proposed framework evaluates LLM-generated feedback along three dimensions: specificity, helpfulness, and validity. |