Challenge: BLEU is widely considered to be an informative metric for text-to-text generation . Xu et al. (2016) found that BLUE is not suitable for evaluation of sentence splitting .
Approach: They propose to use BLEU to evaluate sentence splitting as a metric for machine translation . they propose to compare BLUE with a corpus containing multiple structural paraphrases .
Outcome: The proposed BLEU is not suitable for evaluation of sentence splitting . a correlation analysis with human judgments shows low correlation with BLUE .

Similar Papers

Investigating Text Simplification Evaluation (2021.findings-acl)

Copied to clipboard

Challenge: Existing studies show that parallel TS corpora contain inaccurate simplifications and incorrect alignments.
Approach: They propose to improve the distribution of parallel text simplification corpora to build more robust TS models.
Outcome: The proposed models can be improved by improving the distribution of TS datasets.
Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation Metrics (2020.acl-main)

Copied to clipboard

Challenge: Existing methods for judging metrics are sensitive to the translations used for evaluation, leading to falsely confident conclusions about a metric’s efficacy.
Approach: They propose a method for thresholding performance improvement under an automatic metric against human judgements by using a pairwise system ranking method.
Outcome: The proposed method allows quantification of type I versus type II errors incurred, i.e., insignificant human differences in system quality that are accepted, and significant human differences that are rejected.
Adapting Sentence-level Automatic Metrics for Document-level Simplification Evaluation (2025.naacl-long)

Copied to clipboard

Challenge: Existing studies on text simplification have focused on sentence simplification, but these metrics often underperform on longer texts.
Approach: They propose to adapt existing sentence-level metrics for paragraph- or document-level simplification by incorporating a new approach to the evaluation of text simplification metrics.
Outcome: The proposed approach outperforms existing sentence-level metrics in terms of correlation with human judgment and the sensitivity and robustness of various metrics to different types of errors produced by existing systems.
Semantic Structural Evaluation for Text Simplification (N18-1)

Copied to clipboard

Challenge: Current measures for evaluating text simplification systems focus on lexical aspects, neglecting its structural aspects.
Approach: They propose to use a reference-less automatic evaluation procedure to assess simplification quality by decomposing the input based on its semantic structure and comparing it to the output.
Outcome: The proposed measure has a significant correlation with human judgments and is highly comparable with existing measures.
Document-Level Text Simplification: Dataset, Criteria and Baseline (2021.emnlp-main)

Copied to clipboard

Challenge: Text simplification is a valuable technique, but research on it is limited.
Approach: They propose a document-level simplification task using Wikipedia dumps as a dataset and propose an automatic evaluation metric called D-SARI.
Outcome: The proposed metric is more suitable for document-level simplification task.
A Study in Improving BLEU Reference Coverage with Diverse Automatic Paraphrasing (2020.findings-emnlp)

Copied to clipboard

Challenge: Using neural paraphrasing techniques, we investigate whether automatically generating additional *diverse* references can provide better coverage of the space of valid translations.
Approach: They propose to use neural paraphrasing techniques to generate additional references that provide better coverage of the space of valid translations.
Outcome: The proposed approach beats human paraphrases in the BLEU evaluation.
Are BLEU and Meaning Representation in Opposition? (P18-1)

Copied to clipboard

Challenge: Empirical evaluation suggests that the better the translation quality, the worse the learned sentence representations serve in a wide range of classification and similarity tasks.
Approach: They propose several variations of the attentive NMT architecture to bring this meeting point back . they propose to use a structured fixed-size representation of the input to produce static representations of input sentences.
Outcome: The proposed architecture improves translation quality and performance in a range of tasks.
A Detailed Evaluation of Neural Sequence-to-Sequence Models for In-domain and Cross-domain Text Simplification (L18-1)

Copied to clipboard

Challenge: Xu et al., 2016) show that a simple neural architecture can be efficiently used for in-domain and cross-domain text simplification.
Approach: They evaluate neural sequence-to-sequence models for text simplification on Wikipedia and Newsela datasets.
Outcome: The proposed model can generalize across corpora and overcome challenges when tested on Wikipedia and Newsela datasets.
Deconstructing NLG Evaluation: Evaluation Practices, Assumptions, and Their Implications (2022.naacl-main)

Copied to clipboard

Challenge: Evaluating natural language generation systems is difficult, as there are many ways to express similar things in text.
Approach: They combine interviews with NLG practitioners to examine ethical considerations and their implications for NLG evaluation.
Outcome: The findings of the study surface goals, community practices, assumptions, and constraints that shape NLG evaluations, and examine their implications and how they embody ethical considerations.
What Have We Achieved on Non-autoregressive Translation? (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies have shown that non-autoregressive (NAT) methods underperform autoregressive methods (AT) however, their evaluation using BLEU has been shown to weakly correlate with human annotations.
Approach: They propose to evaluate four representative NAT methods using BLEU to narrow the performance gap between autoregressive and autoregressive translations.
Outcome: The proposed methods underperform NAT and autoregressive methods under more reliable evaluation metrics.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations