Challenge: Existing studies on automatic evaluation of grammatical error correction (GEC) have shown that quality estimation models built from manual evaluation can achieve high performance in automatic evaluation in English.
Approach: They used a dataset with manual evaluation to build an automatic evaluation model for Japanese GEC.
Outcome: The proposed model is based on a Japanese dataset with manual evaluation and meta-evaluation.

Similar Papers

System Combination via Quality Estimation for Grammatical Error Correction (2023.emnlp-main)

Copied to clipboard

Challenge: Existing quality estimation models are not good enough to distinguish good corrections from bad ones, resulting in low F0.5 scores when used for system combination.
Approach: They propose a new quality estimation model that gives a better estimate of the quality of a corrected sentence.
Outcome: The proposed model outperforms the state-of-the-art on the CoNLL-2014 and BEA-2019 test sets, and achieves the highest F0.5 scores published to date.
ProQE: Proficiency-wise Quality Estimation dataset for Grammatical Error Correction (2022.lrec-1)

Copied to clipboard

Challenge: Prior work has shown that QE models of grammatical error correction are biased toward data by learners with relatively high proficiency levels.
Approach: They investigated whether learners' proficiency affects supervised quality estimation models of grammatical error correction (GEC) . they created a QE dataset that includes multiple proficiency levels and explored the necessity of performing proficiency-wise evaluation for QE of GEC.
Outcome: The proposed model is based on multiple proficiency levels and can be performed in real-world scenarios.
Neural Quality Estimation of Grammatical Error Correction (D18-1)

Copied to clipboard

Challenge: Grammatical error correction systems are expected to correct most learners’ writing errors, but in practice they often produce spurious corrections and fail to correct many errors, thereby misleading learners.
Approach: They propose to use supervised learning to estimate the quality of GEC output sentences to help instructors decide whether to correct the errors or ignore them altogether.
Outcome: The proposed model improves on a feature-based baseline and shows that the state-of-the-art system can be improved when quality scores are used as features for re-ranking the N-best candidates.
IMPARA-GED: Grammatical Error Detection is Boosting Reference-free Grammatical Error Quality Estimator (2025.findings-acl)

Copied to clipboard

Challenge: Existing reference-free automatic grammatical error correction methods do not correlate with human evaluation.
Approach: They propose a reference-free automatic grammatical error correction evaluation method with enhanced gramma-ed capabilities.
Outcome: The proposed method achieves highest correlation with human evaluations on a meta-evaluation dataset.
Towards standardizing Korean Grammatical Error Correction: Datasets and Annotation (2023.acl-long)

Copied to clipboard

Challenge: Despite the growing number of Korean learners, little research has been conducted on Korean grammatical error correction (GEC) despite the difficulties of the Korean language, there is no evaluation benchmark for Korean GEC.
Approach: They propose to use Korean grammar error correction datasets to train a machine learning model that can automatically annotate Korean errors from parallel corpora.
Outcome: The proposed model outperforms the currently used statistical Korean GEC system on a wider range of error types.
gec-metrics: A Unified Library for Grammatical Error Correction Evaluation (2025.acl-demo)

Copied to clipboard

Challenge: a library for using and developing grammatical error correction (GEC) evaluation metrics is released under the MIT license .
Approach: They propose a library for using and developing grammatical error correction (GEC) evaluation metrics through a unified interface.
Outcome: The proposed method is based on a unified evaluation framework with a strong focus on API usage and extensible.
SOME: Reference-less Sub-Metrics Optimized for Manual Evaluations of Grammatical Error Correction (2020.coling-main)

Copied to clipboard

Challenge: Existing reference-less metrics are not optimized for manual evaluations of system outputs because no dataset exists for manual analysis.
Approach: They propose a reference-less metric trained on manual evaluations of system outputs for grammatical error correction.
Outcome: The proposed metric improves correlation with manual evaluation in system- and sentence-level meta-evaluation.
Enhancing Grammatical Error Correction Systems with Explanations (2023.acl-long)

Copied to clipboard

Challenge: To help language learners better understand why the GEC system makes a correction, the causes of errors and the corresponding error types are two key factors.
Approach: They propose to annotate large dataset with evidence words and grammatical error types to help language learners better understand corrections.
Outcome: The proposed model can be validated by human evaluation and can be used to help second-language learners decide whether to accept a correction suggestion and understand the associated grammar rule.
UnifiedGEC: Integrating Grammatical Error Correction Approaches for Multi-languages with a Unified Framework (2025.coling-demos)

Copied to clipboard

Challenge: Existing tools for GEC have been developed to support research on grammatical errors, but there is no comprehensive evaluation on these models.
Approach: They propose an open-source framework for Grammatical Error Correction that integrates 5 widely-used GEC models and compares their performance on 7 datasets in different languages.
Outcome: The proposed framework compares 5 widely-used models on 7 datasets in different languages.
Evaluation of Really Good Grammatical Error Correction (2024.lrec-main)

Copied to clipboard

Challenge: emergence of large language models has highlighted the shortcomings of evaluation methods . evaluators often use grammatical error correction (GEC) to correct language errors at multiple levels .
Approach: They perform a comprehensive evaluation of various GEC systems using Swedish learner texts . they suggest using human post-editing to analyze amount of change required to reach native-level human performance .
Outcome: The proposed evaluations outperform existing methods for grammatical error correction in Swedish . the results highlight the shortcomings of existing evaluation methods .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations