| Challenge: | atypical characteristics of some responses make it difficult for an automated scoring system to assign a valid score . a typical spoken response with a lot of background noise may suffer from frequent errors in automated speech recognition . |
| Approach: | They propose a pipeline that detects and processes non-scorable responses at run-time . they also propose linguistic filtering models for spoken responses in language tests . |
| Outcome: | The proposed pipeline detects and processes non-scorable responses at run-time and evaluates them for spoken responses in language proficiency assessment. |
Similar Papers
Automated Essay Scoring: A Reflection on the State of the Art (2024.emnlp-main)
Copied to clipboard
| Challenge: | Automated essay scoring (AES) is a key application of natural language processing . it is based on a holistic score that summarizes the essay's overall quality . |
| Approach: | aaron carroll: automated essay scoring is one of the most important applications in NLP . carroll says the task is still far from being solved, but it's still progressing steadily . he says it'll be interesting to see how researchers can improve performance numbers . |
| Outcome: | a new neural model can beat existing models on a standard evaluation dataset, authors say . authors: the current model is not enough to improve performance numbers . they say it could spark discussion among researchers on how to move forward . |
The Promises and Pitfalls of Using Language Models to Measure Instruction Quality in Education (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing methods to assess instruction quality require trained raters to observe classrooms based on established criteria. |
| Approach: | They propose to use Natural Language Processing techniques to assess multiple high-inference instructional practices in in-person K-12 classrooms and simulated performance tasks for pre-service teachers. |
| Outcome: | The proposed method is able to assess multiple high-inference instructional practices in two educational settings: in-person K-12 classrooms and simulated performance tasks for pre-service teachers. |
Automated Scoring: Beyond Natural Language Processing (C18-1)
Copied to clipboard
| Challenge: | In this paper, we argue that building operational automated scoring systems is a task that has disciplinary complexity above and beyond competitive shared tasks. |
| Approach: | They argue that building operational automated scoring systems is a task that has disciplinary complexity above and beyond standard competitive shared tasks . they argue that it is essential for us as NLP researchers to understand and incorporate these perspectives in our research and work towards a mutually satisfactory solution . |
| Outcome: | The proposed approach is based on the findings of a recent conference on automated scoring. |
ASAP++: Enriching the ASAP Automated Essay Grading Dataset with Essay Attribute Scores (L18-1)
Copied to clipboard
| Challenge: | Automated essay grading (AEG) is one of the most challenging activities in natural language processing (NLP). |
| Approach: | They propose to annotate the ASAP AEG dataset and use it to score different attributes of the essays. |
| Outcome: | The proposed resource is based on the ASAP++ dataset, which contains scores for different attributes of the essays, such as content, word choice, organization, sentence fluency, etc. |
Beyond the Score: Uncertainty-Calibrated LLMs for Automated Essay Assessment (2025.emnlp-main)
Copied to clipboard
| Challenge: | Automated Essay Scoring (AES) systems attain near–human agreement on some public benchmarks, but real-world adoption is limited. |
| Approach: | They propose a distribution-free wrapper that equips any classifier with set-valued outputs enjoying formal coverage guarantees. |
| Outcome: | The proposed model achieves coverage targets while keeping prediction sets compact. |
Incremental Natural Language Processing: Challenges, Strategies, and Evaluation (C18-1)
Copied to clipboard
| Challenge: | In this survey, I consolidate and categorize the approaches, identifying similarities and differences in computation and data, and show trade-offs that have to be considered. |
| Approach: | They consolidate and categorize approaches to incremental processing and show trade-offs that have to be considered. |
| Outcome: | The proposed approaches show that they have similarities and differences in computation and data and that they are not trivial. |
ICLE++: Modeling Fine-Grained Traits for Holistic Essay Scoring (2024.naacl-long)
Copied to clipboard
| Challenge: | Recent advances in automated essay scoring have limited the generalizability of models trained on ASAP. |
| Approach: | They propose to annotate persuasive student essays with holistic and trait-specific scores in a corpus of persuasive student essay annotated with ICLE++. |
| Outcome: | The proposed model can be used to evaluate models for newer AES problems such as multi-trait scoring and cross-prompt scoring. |
Beyond the Gold Standard in Analytic Automated Essay Scoring (2025.acl-srw)
Copied to clipboard
| Challenge: | Automated Essay Scoring (AES) is a new approach to assessing writing practice . traditional holistic scoring methods are not reliable and lack formative feedback in the classroom. |
| Approach: | They propose to combine analytic and holistic AES to create a system that learns from individual raters instead of gold standard labels. |
| Outcome: | The proposed system learns from individual raters instead of gold standard labels. |
From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge (2025.emnlp-main)
Copied to clipboard
Dawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi, Chengshuai Zhao, Zhen Tan, Amrita Bhattacharjee, Yuxuan Jiang, Canyu Chen, Tianhao Wu, Kai Shu, Lu Cheng, Huan Liu
| Challenge: | Recent advances in Large Language Models (LLMs) inspire the "LLM-as-a-judge" paradigm . traditional methods of assessment and evaluation fail in dynamic and open-ended scenarios . |
| Approach: | They propose a paradigm where LLMs are leveraged to perform scoring, ranking, or selection for machine learning evaluation scenarios. |
| Outcome: | The proposed model-based judgment and evaluation paradigms are based on large language models and are compared to the current model-driven evaluation paradigm. |
Explaining Pre-Trained Language Models with Attribution Scores: An Analysis in Low-Resource Settings (2024.lrec-main)
Copied to clipboard
| Challenge: | Currently, prompt-based models are gaining popularity due to their easier adaptability in low-resource settings. |
| Approach: | They analyze attribution scores extracted from prompt-based models w.r.t. plausibility and faithfulness and compare them with attribution score extracted from fine-tuned models and large language models. |
| Outcome: | The proposed model outperforms attention and Integrated Gradients in plausibility and faithfulness, while fine-tuning models are harder to explain in low-resource settings. |