Challenge: Existing evaluation metrics for Grammatical error correction lack explainability . lack of explainability hinders researchers from analyzing strengths and weaknesses of models .
Approach: They propose to assign sentence-level scores to individual edits to improve GEC performance . they use Shapley values, from cooperative game theory, to compute contribution of each edit .
Outcome: The proposed method shows that the evaluation metrics are consistent across edits and human evaluations.

Similar Papers

Rethinking Evaluation Metrics for Grammatical Error Correction: Why Use a Different Evaluation Process than Human? (2025.acl-short)

Copied to clipboard

Challenge: Existing automatic evaluation metrics are based on procedures that diverge from human evaluation.
Approach: They propose to aggregate automatic evaluation metrics to bridge this gap . they propose to use edit-based metrics, -gram based metrics and sentence-level metrics to find the best ranking system.
Outcome: The proposed method outperforms existing metrics on the SEEDA benchmark and improves edit-based metrics, -gram based metrics and sentence-level metrics.
A Reassessment of Reference-Based Grammatical Error Correction Metrics (C18-1)

Copied to clipboard

Challenge: Existing studies on the correlation of GEC metrics with human judgments were inconclusive . a recent study found that GLEU produces counter-intuitive scores in common test examples .
Approach: They propose to use GLEU to evaluate grammatical error correction (GEC) systems . they also use statistical significance tests to assess their agreement with human judgments .
Outcome: The proposed metrics show no significant advantage over MaxMatch (GLEU) the results contradict previous studies that claim GLEU superior .
Is this the end of the gold standard? A straightforward reference-less grammatical error correction metric (2021.emnlp-main)

Copied to clipboard

Challenge: Existing evaluations of grammatical error correction systems use reference-based metrics, but they are limited because of multiple correct outputs.
Approach: They propose a system that uses commonly available tools to evaluate grammatical error correction (GEC) systems.
Outcome: The proposed system solves the issues related to the use of a reference and does not need another annotated dataset for fine-tuning.
CLEME2.0: Towards Interpretable Evaluation by Disentangling Edits for Grammatical Error Correction (2025.acl-long)

Copied to clipboard

Challenge: Existing studies have focused on the interpretability of Grammatical Error Correction (GEC) evaluation metrics, but the interpretabilty of these metrics has been neglected.
Approach: They propose a reference-based metric that describes four aspects of GEC systems: hit-correction, wrong-corrections, under-correcties, and over-corrects.
Outcome: The proposed metric reveals critical qualities and locates drawbacks of GEC systems.
SOME: Reference-less Sub-Metrics Optimized for Manual Evaluations of Grammatical Error Correction (2020.coling-main)

Copied to clipboard

Challenge: Existing reference-less metrics are not optimized for manual evaluations of system outputs because no dataset exists for manual analysis.
Approach: They propose a reference-less metric trained on manual evaluations of system outputs for grammatical error correction.
Outcome: The proposed metric improves correlation with manual evaluation in system- and sentence-level meta-evaluation.
gec-metrics: A Unified Library for Grammatical Error Correction Evaluation (2025.acl-demo)

Copied to clipboard

Challenge: a library for using and developing grammatical error correction (GEC) evaluation metrics is released under the MIT license .
Approach: They propose a library for using and developing grammatical error correction (GEC) evaluation metrics through a unified interface.
Outcome: The proposed method is based on a unified evaluation framework with a strong focus on API usage and extensible.
Automatic Metric Validation for Grammatical Error Correction (P18-1)

Copied to clipboard

Challenge: Existing methods for metric validation in GEC suffer from low inter-rater agreement.
Approach: They propose an automatic method for GEC metric validation that overcomes many of the difficulties in the existing method.
Outcome: The proposed method sheds new light on metric quality and shows valid edits are penalized by existing metrics.
Revisiting Grammatical Error Correction Evaluation and Beyond (2022.emnlp-main)

Copied to clipboard

Challenge: Pretraining-based (PT) evaluation metrics are not effective for training grammatical error correction systems.
Approach: They propose a pretraining-based GEC evaluation metric which only uses PT-based metrics to score the corrected parts of the system.
Outcome: The proposed evaluation metric outperforms existing methods on a CoNLL14 evaluation task.
Reliability Crisis of Reference-free Metrics for Grammatical Error Correction (2025.findings-emnlp)

Copied to clipboard

Challenge: Reference-free evaluation metrics for grammatical error correction have high correlation with human judgments, but they are not designed to evaluate adversarial systems that aim to obtain unjustifiably high scores.
Approach: They propose adversarial attack strategies for four reference-free metrics . they propose SOME, Scribendi, IMPARA, and LLM-based metrics based on these metrics a .
Outcome: The proposed attacks outperform the current state-of-the-art for four reference-free metrics .
Adapting Sentence-level Automatic Metrics for Document-level Simplification Evaluation (2025.naacl-long)

Copied to clipboard

Challenge: Existing studies on text simplification have focused on sentence simplification, but these metrics often underperform on longer texts.
Approach: They propose to adapt existing sentence-level metrics for paragraph- or document-level simplification by incorporating a new approach to the evaluation of text simplification metrics.
Outcome: The proposed approach outperforms existing sentence-level metrics in terms of correlation with human judgment and the sensitivity and robustness of various metrics to different types of errors produced by existing systems.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations