| Challenge: | Recent advances in large language models (LLMs) have demonstrated their capacity to match and even exceed average human performance across diverse, unseen tasks. |
| Approach: | They compare the results of different LLMs in TST evaluation using multiple input prompts and introduce the concept of prompt ensembling. |
| Outcome: | The proposed model outperforms human evaluations on multiple input prompts. |
Similar Papers
Evaluating Text Style Transfer Evaluation: Are There Any Reliable Metrics? (2025.naacl-srw)
Copied to clipboard
| Challenge: | Text style transfer (TST) is a multidimensional task requiring the assessment of style transfer accuracy, content preservation, and naturalness. |
| Approach: | They propose to use text style transfer metrics to evaluate outputs of text editors . they also investigate the potential of large language models as tools for TST evaluation . |
| Outcome: | The proposed methods provide better insights than existing metrics, the authors show . their meta-evaluation through correlation with hu-man judgments shows they are effective . |
A Call for Standardization and Validation of Text Style Transfer Evaluation (2023.findings-acl)
Copied to clipboard
| Challenge: | Text style transfer (TST) evaluation is inconsistent in practice. |
| Approach: | They conduct a meta-analysis on human and automated TST evaluation and experimentation . they find a standardization gap and a validation gap in the field . |
| Outcome: | The authors find that evaluation procedures are inconsistent and that they need to improve on them. |
Prompt-and-Rerank: A Method for Zero-Shot and Few-Shot Arbitrary Textual Style Transfer with Small Language Models (2022.emnlp-main)
Copied to clipboard
| Challenge: | a new method for textual style transfer is proposed for text with a limited set of style choices . textual styles are a complex task that requires specialized models to perform . |
| Approach: | They propose a method for arbitrary textual style transfer using pre-trained language models . they use a mathematical formulation of the TST task, decomposing it into three components . |
| Outcome: | The proposed method performs on par with state-of-the-art large-scale models while using less compute and memory. |
Towards Actual (Not Operational) Textual Style Transfer Auto-Evaluation (D19-55)
Copied to clipboard
| Challenge: | elucidates the dangerous current state of style transfer auto-evaluation research. |
| Approach: | They propose ways to aggregate the three metrics into one evaluator. |
| Outcome: | The proposed method could be used to aggregate the three metrics into one evaluator. |
Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models (2024.findings-acl)
Copied to clipboard
| Challenge: | Recent research shows that large language models (LLMs) perform poorly at segment level. |
| Approach: | They propose a new prompting method that emulates the commonly accepted human evaluation framework . they will release their code and scripts to facilitate the community . |
| Outcome: | The proposed method is based on the human evaluation framework MQM and produces explainable and reliable MT evaluations at both the system and segment level. |
PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization Evaluation (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) are useful for low-resource scenarios and time-restricted applications. |
| Approach: | They propose a large-scale evaluation tool for large language models that uses prompts . they evaluate 720 prompt templates for open-source LLM-based metrics on MT and summarization datasets a 6.6M evaluations. |
| Outcome: | The proposed model evaluates 720 prompt templates on machine translation and summarization datasets. |
Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles (2025.emnlp-main)
Copied to clipboard
| Challenge: | Current benchmarks for evaluating Large Language Models do not capture the rich variety of communication patterns exhibited by humans. |
| Approach: | They propose a low-cost method to emulate diverse writing styles by rewriting evaluation prompts using persona-based LLM prompting. |
| Outcome: | The proposed method improves the external validity of the benchmarks for Large Language Models (LLMs) based on persona-based prompting. |
Evaluation Metrics in the Era of GPT-4: Reliably Evaluating Large Language Models on Sequence to Sequence Tasks (2023.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) evaluation is a patchy and inconsistent landscape . established automatic evaluation metrics are poor surrogates, correlating weakly with human judgement. |
| Approach: | They propose to use both automatic and human evaluation to evaluate generative LLMs on three NLP benchmarks: text summarisation, text simplification and grammatical error correction. |
| Outcome: | The proposed model outperforms many popular models according to human reviewers on the majority of metrics, while scoring much worse when using classic automatic evaluation metrics. |
Evaluating the Evaluation Metrics for Style Transfer: A Case Study in Multilingual Formality Transfer (2021.emnlp-main)
Copied to clipboard
| Challenge: | a lack of standardized and reliable methods for automatic evaluation hinders ST . prior work has employed as many as nine different automatic systems to rate formality alone . |
| Approach: | They evaluate automatic metrics on the oft-researched task of formality style transfer . they outline best practices for automatic evaluation in (formality) style transfer and identify models that correlate well with human judgments. |
| Outcome: | The proposed models correlate well with human judgments and are robust across languages. |
Mind the Style Gap: Meta-Evaluation of Style and Attribute Transfer Metrics (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models (LLMs) make it easy to rewrite a text in any style, but they are not straightforward when evaluating content preservation. |
| Approach: | They propose a large meta-evaluation of metrics for evaluating style and attribute transfer, focusing on content preservation. |
| Outcome: | The proposed method achieves higher alignment with human judgements than prompting a model of a similar size as an autorater. |