| Challenge: | Existing systems for writing evaluation focus on argumentative texts . a précis is a written text that provides a coherent summary of main points . |
| Approach: | They propose to use a corpus of English précis texts to train a machine learning model . they find it is able to predict the grade of précis texts with only a moderate error margin . |
| Outcome: | The proposed model predicts the grade of précis texts with only a moderate error margin. |
Similar Papers
Enhancing Automated Essay Scoring Performance via Fine-tuning Pre-trained Language Models with Combination of Regression and Ranking (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Recent work on sentence prediction tasks uses shallow neural networks to learn essay representations and constrain calculated scores with regression loss or ranking loss. |
| Approach: | They propose to use a pre-trained language model to learn text representations first and then to constrain the scores with regression loss or ranking loss. |
| Outcome: | The proposed model outperforms state-of-the-art models on the Automated Student Assessment Prize dataset. |
Neural Text Summarization: A Critical Evaluation (D19-1)
Copied to clipboard
| Challenge: | Current approaches to text summarization use advanced attention and copying mechanisms, multi-task and multi-reward training techniques. |
| Approach: | They evaluate datasets, evaluation metrics, and models for text summarization . they highlight three primary shortcomings: 1) datasets leave task underconstrained; 2) models overfit layout biases . |
| Outcome: | The current evaluation protocol is weakly correlated with human judgment and does not account for factual correctness. |
Automatic Focus Annotation: Bringing Formal Pragmatics Alive in Analyzing the Information Structure of Authentic Data (N18-1)
Copied to clipboard
| Challenge: | Using focus-background dichotomy, discourse and information structure of sentences are being studied in context. |
| Approach: | They propose to automate the analysis of focus in authentic written data by using a range of lexical, syntactic, and semantic features to achieve an accuracy of 78.1%. |
| Outcome: | The proposed approach achieves 78.1% accuracy for identifying focus in authentic written data. |
Beyond the Tip of the Iceberg: Assessing Coherence of Text Classifiers (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Large-scale, pre-trained language models achieve human-level and superhuman accuracy on existing language understanding tasks, but statistical bias in benchmark data and probing studies has recently called into question their true capabilities. |
| Approach: | They propose to evaluate systems through a measure of prediction coherence by using two existing language understanding benchmarks with different properties to demonstrate its versatility. |
| Outcome: | The proposed evaluation framework is quick, effective, and versatile to provide insight into the coherence of machines’ predictions. |
Neural Automated Essay Scoring and Coherence Modeling for Adversarially Crafted Input (N18-1)
Copied to clipboard
| Challenge: | Existing approaches to Automated Essay Scoring (AES) are not well-suited to capture adversarially crafted input of grammatical but incoherent sequences of sentences. |
| Approach: | They propose a neural model of local coherence that can effectively learn connectedness features between sentences. |
| Outcome: | The proposed approach strengthens the validity of neural essay scoring models. |
APPLS: Evaluating Evaluation Metrics for Plain Language Summarization (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing evaluation metrics for plain language summarization (PLS) lack a dedicated assessment metric and the suitability of text generation evaluation metrics is unclear due to unique transformations. |
| Approach: | They propose a granular meta-evaluation testbed to evaluate PLS metrics . they identify four PLS criteria and define perturbations that sensitive metrics should be able to detect . |
| Outcome: | The proposed testbed assesses performance of 14 existing metrics including scores, features, and prompt-based evaluations. |
A Dataset for Investigating the Impact of Feedback on Student Revision Outcome (2020.lrec-1)
Copied to clipboard
| Challenge: | Despite numerous studies on the kinds of feedback that can best promote learning, this question remains an open debate in the area of Second Language Acquisition (SLA). |
| Approach: | They annotate a corpus of student-written sentences with teacher feedback provided for the errors. |
| Outcome: | The proposed annotation scheme and the teacher feedback dataset are based on student-written sentences in their original and revised versions with teacher feedback provided for the errors. |
LEAF: Language Learners’ English Essays and Feedback Corpus (2024.naacl-short)
Copied to clipboard
| Challenge: | Current automated essay scoring models lack the granularity desired by learners and instructors seeking more detailed insights. |
| Approach: | They present a corpus of English essays and their corresponding feedback from the “essayforum” website. |
| Outcome: | The LEAF corpus provides valuable feedback for students and teachers . it provides insights on argumentative aspects and organizational coherence . |
Automated Essay Scoring in the Presence of Biased Ratings (N18-1)
Copied to clipboard
| Challenge: | Existing studies on rater effects in general settings have not investigated how rater bias affects automated essay scoring. |
| Approach: | They propose to model rater bias by removing essays associated with potentially biased scores from annotated corpus. |
| Outcome: | The proposed model is based on comments provided by raters and is compared with existing corpus. |
Beyond Agreement: Diagnosing the Rationale Alignment of Automated Essay Scoring Methods based on Linguistically-informed Counterfactuals (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing Automated Essay Scoring (AES) methods focus on sentence-level features, whereas Large Language Models (LLMs) are sensitive to conventions & accuracy, language complexity, and organization. |
| Approach: | They propose to use large language models to aid in decision-making . they propose to analyze the reasoning of neural models by analyzing sentence-level features. |
| Outcome: | The proposed method improves understanding of neural approaches to Automated Essay Scoring (AES) and can also apply to other domains seeking transparency in model-driven decisions. |