Evaluating Style Transfer for Text (N19-1)

Copied to clipboard

Challenge: Existing studies on style transfer for text are lacking a standard set of evaluation practices.
Approach: They propose a set of metrics for automated evaluation that are more strongly correlated with human judgment and show tradeoffs between aspects of interest.
Outcome: The proposed models exhibit tradeoffs between aspects of interest and human judgment, demonstrating the importance of evaluating them at specific points of their tradeoff plots.

Similar Papers

Rethinking Sentiment Style Transfer (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluation methods for text style transfer are unsatisfactory.
Approach: They propose to use a graph-based method to extract attribute content from sentences . they propose an efficient regularization to leverage attribute-dependent content as guiding signals.
Outcome: The proposed method is based on a YELP and IMDB dataset and it is able to detect errors in the human evaluation.
Evaluating Text Style Transfer Evaluation: Are There Any Reliable Metrics? (2025.naacl-srw)

Copied to clipboard

Challenge: Text style transfer (TST) is a multidimensional task requiring the assessment of style transfer accuracy, content preservation, and naturalness.
Approach: They propose to use text style transfer metrics to evaluate outputs of text editors . they also investigate the potential of large language models as tools for TST evaluation .
Outcome: The proposed methods provide better insights than existing metrics, the authors show . their meta-evaluation through correlation with hu-man judgments shows they are effective .
Mind the Style Gap: Meta-Evaluation of Style and Attribute Transfer Metrics (2025.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) make it easy to rewrite a text in any style, but they are not straightforward when evaluating content preservation.
Approach: They propose a large meta-evaluation of metrics for evaluating style and attribute transfer, focusing on content preservation.
Outcome: The proposed method achieves higher alignment with human judgements than prompting a model of a similar size as an autorater.
Towards Actual (Not Operational) Textual Style Transfer Auto-Evaluation (D19-55)

Copied to clipboard

Challenge: elucidates the dangerous current state of style transfer auto-evaluation research.
Approach: They propose ways to aggregate the three metrics into one evaluator.
Outcome: The proposed method could be used to aggregate the three metrics into one evaluator.
Style versus Content: A distinction without a (learnable) difference? (2020.coling-main)

Copied to clipboard

Challenge: Textual style transfer assumes that it is possible to separate style from content . however, style transfer can provide insight into language more generally .
Approach: They propose to use sentiment transfer to examine whether style transfer is possible . they employ adversarial encoder-decoder networks to analyze style-related features .
Outcome: The proposed method combines style transfer with content preservation and fluency to show that style cannot be usefully separated from content within style transfer systems.
Style Transfer for Texts: Retrain, Report Errors, Compare with Rewrites (D19-1)

Copied to clipboard

Challenge: Currently, standard methods for style transfer have several significant problems.
Approach: They propose to take BLEU between input and human-written reformulations into consideration for benchmarks.
Outcome: The proposed architectures outperform state-of-the-art in style transfer metric on human-written reformulations and take BLEU between input and output into consideration for benchmarks.
Evaluating the Evaluation Metrics for Style Transfer: A Case Study in Multilingual Formality Transfer (2021.emnlp-main)

Copied to clipboard

Challenge: a lack of standardized and reliable methods for automatic evaluation hinders ST . prior work has employed as many as nine different automatic systems to rate formality alone .
Approach: They evaluate automatic metrics on the oft-researched task of formality style transfer . they outline best practices for automatic evaluation in (formality) style transfer and identify models that correlate well with human judgments.
Outcome: The proposed models correlate well with human judgments and are robust across languages.
A Call for Standardization and Validation of Text Style Transfer Evaluation (2023.findings-acl)

Copied to clipboard

Challenge: Text style transfer (TST) evaluation is inconsistent in practice.
Approach: They conduct a meta-analysis on human and automated TST evaluation and experimentation . they find a standardization gap and a validation gap in the field .
Outcome: The authors find that evaluation procedures are inconsistent and that they need to improve on them.
Text Style Transfer Evaluation Using Large Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have demonstrated their capacity to match and even exceed average human performance across diverse, unseen tasks.
Approach: They compare the results of different LLMs in TST evaluation using multiple input prompts and introduce the concept of prompt ensembling.
Outcome: The proposed model outperforms human evaluations on multiple input prompts.
Unsupervised Evaluation Metrics and Learning Criteria for Non-Parallel Textual Transfer (D19-56)

Copied to clipboard

Challenge: Existing methods for textual transfer with no parallel corpora are insufficient to evaluate textual paraphrases with modified attributes or properties.
Approach: They propose to add a metric for post-transfer classification accuracy and a method to combine them into a single overall score.
Outcome: The proposed metrics correlate well with human judgments, at both the sentence-level and system-level.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations