Challenge: Existing studies have evaluated grammatical error correction models on a single corpus, but the evaluation is incomplete because the task difficulty varies depending on the corpus and conditions such as proficiency levels of the writers and essay topics.
Approach: They evaluate the performance of several GEC models against various learner corpora and compare their rankings against the corpus.
Outcome: The evaluation of several models against learner corpora shows that the models’ rankings vary depending on the corpus, indicating that single-corpus evaluation is insufficient for GEC models.

Similar Papers

Grammatical Error Correction: Are We There Yet? (2022.coling-1)

Copied to clipboard

Challenge: grammatical error correction (GEC) systems outperform humans on the CoNLL-2014 test set, but there are still classes of errors that they fail to correct.
Approach: They found that state-of-the-art GEC systems outperform humans by a wide margin on the CoNLL-2014 test set . however, they found that there are still classes of errors that they fail to correct .
Outcome: The F0.5 evaluation metric outperforms the CoNLL-2014 test set, but there are still classes of errors that they fail to correct.
Targeted Syntactic Evaluation for Grammatical Error Correction (2025.acl-long)

Copied to clipboard

Challenge: Existing evaluation datasets based on learner-produced texts are insufficient for evaluating models . Currently, sequence-to-sequence models and sequence tagging models perform well on beginner-level grammar items .
Approach: They propose a new evaluation paradigm that assesses GEC models using minimal pairs of ungrammatical and grammatically paired sentences for each grammar item.
Outcome: The proposed evaluation paradigm assesses models using minimal pairs of ungrammatical and grammatically-spaced sentences for each grammar item.
Grammatical Error Correction in Low Error Density Domains: A New Benchmark and Analyses (2020.emnlp-main)

Copied to clipboard

Challenge: CWEB is a new benchmark for grammatical error correction (GEC) systems . website data contains far fewer grammamatical errors than learner essays .
Approach: They propose to broaden the target domain of grammatical error correction (GEC) systems . website data contains far fewer grammamatical errors than learner essays .
Outcome: The proposed model can't rely on a strong internal language model in low error density domains.
FCGEC: Fine-Grained Corpus for Chinese Grammatical Error Correction (2022.findings-emnlp)

Copied to clipboard

Challenge: grammatical error correction (GEC) is a complex task that requires high-quality data from native speakers.
Approach: They propose a human-annotated corpus to detect, identify and correct grammatical errors in Chinese examinations.
Outcome: The proposed model outperforms other models in low-resource settings, but there is a significant gap between the models and humans that encourages future models to bridge it.
Rethinking Evaluation Metrics for Grammatical Error Correction: Why Use a Different Evaluation Process than Human? (2025.acl-short)

Copied to clipboard

Challenge: Existing automatic evaluation metrics are based on procedures that diverge from human evaluation.
Approach: They propose to aggregate automatic evaluation metrics to bridge this gap . they propose to use edit-based metrics, -gram based metrics and sentence-level metrics to find the best ranking system.
Outcome: The proposed method outperforms existing metrics on the SEEDA benchmark and improves edit-based metrics, -gram based metrics and sentence-level metrics.
Near Human-Level Performance in Grammatical Error Correction with Hybrid Machine Translation (N18-2)

Copied to clipboard

Challenge: Currently, most effective GEC systems are based on phrase-based statistical machine translation.
Approach: They combine two of the most popular approaches to automated Grammatical Error Correction (GEC) they create a hybrid GEC system that preserves the accuracy of SMT output and generates more fluent sentences .
Outcome: The proposed system achieves state-of-the-art on the CoNLL-2014 and JFLEG benchmarks.
Do Grammatical Error Correction Models Realize Grammatical Generalization? (2021.findings-acl)

Copied to clipboard

Challenge: Existing models for grammatical error correction use pseudo data, but they are inconvenient for realworld deployment due to large amounts of training data.
Approach: They propose a method to evaluate whether GEC models can generalize to unseen errors by using synthetic and real GEC datasets with controlled vocabularies.
Outcome: The proposed model fails to realize grammatical generalization even in simple settings with limited vocabulary and syntax, suggesting it lacks the generalization ability required to correct errors from provided training examples.
Evaluation of Really Good Grammatical Error Correction (2024.lrec-main)

Copied to clipboard

Challenge: emergence of large language models has highlighted the shortcomings of evaluation methods . evaluators often use grammatical error correction (GEC) to correct language errors at multiple levels .
Approach: They perform a comprehensive evaluation of various GEC systems using Swedish learner texts . they suggest using human post-editing to analyze amount of change required to reach native-level human performance .
Outcome: The proposed evaluations outperform existing methods for grammatical error correction in Swedish . the results highlight the shortcomings of existing evaluation methods .
Cross-Sentence Grammatical Error Correction (P19-1)

Copied to clipboard

Challenge: Existing approaches to automatic grammatical error correction (GEC) ignore cross-sentence context . existing approaches only correct one sentence at a time and ignore useful contextual information .
Approach: They propose to use an auxiliary encoder that encodes previous sentences and incorporates the encoding in the decoder via attention and gating mechanisms.
Outcome: The proposed model improves over strong baselines on a synthetic dataset showing high performance in verb tense corrections that require cross-sentence context.
IMPARA-GED: Grammatical Error Detection is Boosting Reference-free Grammatical Error Quality Estimator (2025.findings-acl)

Copied to clipboard

Challenge: Existing reference-free automatic grammatical error correction methods do not correlate with human evaluation.
Approach: They propose a reference-free automatic grammatical error correction evaluation method with enhanced gramma-ed capabilities.
Outcome: The proposed method achieves highest correlation with human evaluations on a meta-evaluation dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations