| Challenge: | Existing studies on style transfer for text are lacking a standard set of evaluation practices. |
| Approach: | They propose a set of metrics for automated evaluation that are more strongly correlated with human judgment and show tradeoffs between aspects of interest. |
| Outcome: | The proposed models exhibit tradeoffs between aspects of interest and human judgment, demonstrating the importance of evaluating them at specific points of their tradeoff plots. |
Similar Papers
Rethinking Sentiment Style Transfer (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing evaluation methods for text style transfer are unsatisfactory. |
| Approach: | They propose to use a graph-based method to extract attribute content from sentences . they propose an efficient regularization to leverage attribute-dependent content as guiding signals. |
| Outcome: | The proposed method is based on a YELP and IMDB dataset and it is able to detect errors in the human evaluation. |
Evaluating Text Style Transfer Evaluation: Are There Any Reliable Metrics? (2025.naacl-srw)
Copied to clipboard
| Challenge: | Text style transfer (TST) is a multidimensional task requiring the assessment of style transfer accuracy, content preservation, and naturalness. |
| Approach: | They propose to use text style transfer metrics to evaluate outputs of text editors . they also investigate the potential of large language models as tools for TST evaluation . |
| Outcome: | The proposed methods provide better insights than existing metrics, the authors show . their meta-evaluation through correlation with hu-man judgments shows they are effective . |
Mind the Style Gap: Meta-Evaluation of Style and Attribute Transfer Metrics (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models (LLMs) make it easy to rewrite a text in any style, but they are not straightforward when evaluating content preservation. |
| Approach: | They propose a large meta-evaluation of metrics for evaluating style and attribute transfer, focusing on content preservation. |
| Outcome: | The proposed method achieves higher alignment with human judgements than prompting a model of a similar size as an autorater. |
Towards Actual (Not Operational) Textual Style Transfer Auto-Evaluation (D19-55)
Copied to clipboard
| Challenge: | elucidates the dangerous current state of style transfer auto-evaluation research. |
| Approach: | They propose ways to aggregate the three metrics into one evaluator. |
| Outcome: | The proposed method could be used to aggregate the three metrics into one evaluator. |
Style versus Content: A distinction without a (learnable) difference? (2020.coling-main)
Copied to clipboard
| Challenge: | Textual style transfer assumes that it is possible to separate style from content . however, style transfer can provide insight into language more generally . |
| Approach: | They propose to use sentiment transfer to examine whether style transfer is possible . they employ adversarial encoder-decoder networks to analyze style-related features . |
| Outcome: | The proposed method combines style transfer with content preservation and fluency to show that style cannot be usefully separated from content within style transfer systems. |
Style Transfer for Texts: Retrain, Report Errors, Compare with Rewrites (D19-1)
Copied to clipboard
| Challenge: | Currently, standard methods for style transfer have several significant problems. |
| Approach: | They propose to take BLEU between input and human-written reformulations into consideration for benchmarks. |
| Outcome: | The proposed architectures outperform state-of-the-art in style transfer metric on human-written reformulations and take BLEU between input and output into consideration for benchmarks. |
Evaluating the Evaluation Metrics for Style Transfer: A Case Study in Multilingual Formality Transfer (2021.emnlp-main)
Copied to clipboard
| Challenge: | a lack of standardized and reliable methods for automatic evaluation hinders ST . prior work has employed as many as nine different automatic systems to rate formality alone . |
| Approach: | They evaluate automatic metrics on the oft-researched task of formality style transfer . they outline best practices for automatic evaluation in (formality) style transfer and identify models that correlate well with human judgments. |
| Outcome: | The proposed models correlate well with human judgments and are robust across languages. |
A Call for Standardization and Validation of Text Style Transfer Evaluation (2023.findings-acl)
Copied to clipboard
| Challenge: | Text style transfer (TST) evaluation is inconsistent in practice. |
| Approach: | They conduct a meta-analysis on human and automated TST evaluation and experimentation . they find a standardization gap and a validation gap in the field . |
| Outcome: | The authors find that evaluation procedures are inconsistent and that they need to improve on them. |
Text Style Transfer Evaluation Using Large Language Models (2024.lrec-main)
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have demonstrated their capacity to match and even exceed average human performance across diverse, unseen tasks. |
| Approach: | They compare the results of different LLMs in TST evaluation using multiple input prompts and introduce the concept of prompt ensembling. |
| Outcome: | The proposed model outperforms human evaluations on multiple input prompts. |
Unsupervised Evaluation Metrics and Learning Criteria for Non-Parallel Textual Transfer (D19-56)
Copied to clipboard
| Challenge: | Existing methods for textual transfer with no parallel corpora are insufficient to evaluate textual paraphrases with modified attributes or properties. |
| Approach: | They propose to add a metric for post-transfer classification accuracy and a method to combine them into a single overall score. |
| Outcome: | The proposed metrics correlate well with human judgments, at both the sentence-level and system-level. |