Papers with post-editing

21 papers
IntelliCAT: Intelligent Machine Translation Post-Editing with Quality Estimation and Translation Suggestion (2021.acl-demo)

Copied to clipboard

Challenge: Existing computer-aided translation tools require the translator to edit incorrect parts of a document, while ITP tools require fewer edits.
Approach: They propose an interactive translation interface with neural models that streamline the post-editing process on machine translation output.
Outcome: The proposed interface can significantly improve translation quality and a user study shows that it speeds up the post-editing process by 52.9% compared to translating from scratch.
Improving Cross-Domain Low-Resource Text Generation through LLM Post-Editing: A Programmer-Interpreter Approach (2024.findings-eacl)

Copied to clipboard

Challenge: Large pre-trained language models such as GPT-3.5 and GPT-4 have gained significant attention in natural language research due to limited computational resources or inaccessible parameters.
Approach: They propose a neural programmer-interpreter approach that preserves the domain generalization ability of LLMs while editing their output.
Outcome: The proposed framework significantly improves GPT-3.5’s performance in logical form-to-text conversion and low-resource machine translation, surpassing other state-of-the-art (SOTA) LLM post-editing methods in cross-domain settings.
An Exploration of Post-Editing Effectiveness in Text Summarization (2022.naacl-main)

Copied to clipboard

Challenge: Automated summarization methods are efficient but can suffer from low quality.
Approach: They conducted an experiment with 72 participants to compare post-editing provided summaries with manual summarization for summary quality, human efficiency, and user experience.
Outcome: The results show that post-editing improves summary quality, human efficiency, and user experience on formal (XSum news) and informal (Reddit posts) text.
MMPE: A Multi-Modal Interface using Handwriting, Touch Reordering, and Speech Commands for Post-Editing Machine Translation (2020.acl-demos)

Copied to clipboard

Challenge: a shift from traditional translation to post-editing (PE) of machine-translated text can save time and reduce errors, but it also affects the design of translation interfaces.
Approach: They propose a prototype that combines traditional input modes with pen, touch, and speech modalities for post-editing of machine-translated (MT) they propose to use these modalités to cross out or hand-write new text, drag and drop words for reordering, or use spoken commands to update the text in place.
Outcome: The proposed interfaces can be used to cross out or hand-write new text, drag and drop words for reordering, or use spoken commands to update the text in place.
Global Optimization under Length Constraint for Neural Text Summarization (P19-1)

Copied to clipboard

Challenge: GOLC increases the probabilities of generating summaries that have high evaluation scores within a desired length.
Approach: They propose a global optimization method under length constraint for neural text summarization models.
Outcome: The proposed method generates fewer overlength summaries while maintaining the fastest processing speed.
Test-Time Scaling of Reasoning Models for Machine Translation (2026.eacl-long)

Copied to clipboard

Challenge: Using TTS, Reasoning Models (RMs) are able to perform tasks such as math and coding with limited results.
Approach: They evaluate 12 Reasoning Models across a diverse suite of MT benchmarks, examining three scenarios: direct translation, forced-reasoning extrapolation, and post-editing.
Outcome: The proposed approach improves translation quality on three domains, with inconsistent results for general-purpose RMs and performance plateauing.
MMPE: A Multi-Modal Interface for Post-Editing Machine Translation (2020.acl-main)

Copied to clipboard

Challenge: Current advances in machine translation (MT) increase the need for translators to switch from traditional translation to post-editing (PE) of machine-translated text.
Approach: They propose to combine traditional input modes with pen, touch, and speech modalities for post-editing of machine-translated text.
Outcome: The proposed interfaces are designed to reduce errors and save time.
Computer Assisted Translation with Neural Quality Estimation and Automatic Post-Editing (2020.findings-emnlp)

Copied to clipboard

Challenge: Using neural machine translation to approximate human parity is difficult due to the lack of parallel training corpora.
Approach: They propose an end-to-end deep learning framework for quality estimation and automatic post-editing of machine translation output.
Outcome: The proposed framework achieves state-of-the-art performance on the English–German dataset and human translators can significantly expedite their post-editing processing with the model.
What Does LLM Refinement Actually Improve? A Systematic Study on Document-Level Literary Translation (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have made document-level machine translation increasingly practical, enabled by long-context modeling and strong generation quality.
Approach: They propose to use document-level MT followed by segment-level refinement to find the strongest and most stable improvements across six LLMs and seven language pairs.
Outcome: The proposed method outperforms error-specific prompting and evaluate-then-refine schemes in document-level translation.
Automatic Post-Editing of Machine Translation: A Neural Programmer-Interpreter Approach (D18-1)

Copied to clipboard

Challenge: Existing approaches to inducing APE have suffered from over-correction, where the APE system tends to keep the machine translated text without any modification.
Approach: They propose a neural programmer-interpreter approach to automated post-editing (APE) that mimics human perform post- editing using discrete edit operations . their model outperforms previous neural models for inducing PE programs on the WMT17 APE task for German-English up to +1 BLEU score and -0.7 TER scores.
Outcome: The proposed model outperforms previous neural models for inducing PE programs on the WMT17 APE task for German-English up to +1 BLEU score and -0.7 TER scores.
WeTS: A Benchmark for Translation Suggestion (2022.emnlp-main)

Copied to clipboard

Challenge: Existing studies focus on overall performance of machine translation but ignore TS performance, authors say . if TS is applied into post-editing, it will reduce the time and cost of post-production.
Approach: They propose to use a golden corpus annotated by experts to generate a translation suggestion model.
Outcome: The proposed model improves on the golden corpus annotated by translators on four translation directions.
Mid-Air Hand Gestures for Post-Editing of Machine Translation (2021.acl-long)

Copied to clipboard

Challenge: In a well-connected world, translation is of everincreasing importance.
Approach: They propose to use mid-air hand gestures in combination with the keyboard for editing in machine translation and post-editing workflows to improve quality.
Outcome: The proposed prototype supports mid-air hand gestures for cursor placement, text selection, deletion, and reordering.
DivEMT: Neural Machine Translation Post-Editing Effort Across Typologically Diverse Languages (2022.emnlp-main)

Copied to clipboard

Challenge: Recent advances in neural language modeling and multilingual training have prompted widespread adoption of machine translation (MT) technologies across an unprecedented range of world languages.
Approach: They propose to use a dataset to assess the impact of two state-of-the-art NMT systems, Google Translate and the multilingual mBART-50 model, on translation productivity.
Outcome: The proposed model is faster than translation from scratch, but the magnitude of productivity gains varies widely across systems and languages.
Automatic Input Rewriting Improves Translation with Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: LLMs can rewrite inputs but in machine translation, they are primarily used to re-write outputs via post-editing.
Approach: They propose to use LLMs to rewrite inputs automatically to improve machine translation (MT) they propose to simplify inputs and use quality estimation to assess translatability.
Outcome: The proposed methods can be improved by using quality estimation to assess translatability.
Evaluating Automatic Subtitling: Correlating Post-editing Effort and Automatic Metrics (2024.lrec-main)

Copied to clipboard

Challenge: Existing metrics for automatic subtitling are not yet fully explored.
Approach: They propose to use machine translation metrics to measure post-editing effort in automatic subtitling to collect data on product-, process- and participant-based data.
Outcome: The proposed metrics correlate with measures of post-editing effort in automatic subtitling.
MMTE: Corpus and Metrics for Evaluating Machine Translation Quality of Metaphorical Language (2024.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation methods focus on fluency and factual reliability, while neglecting figurative quality.
Approach: They propose a set of human evaluation metrics focused on the translation of figurative language and a parallel metaphor corpus generated by post-editing.
Outcome: The proposed evaluation protocol estimates four aspects of MT: Metaphorical Equivalence, Emotion, Authenticity, and Quality.
Correcting Diverse Factual Errors in Abstractive Summarization via Post-Editing and Language Model Infilling (2022.emnlp-main)

Copied to clipboard

Challenge: Abstractive summarization models often generate inconsistent summaries containing factual errors or fabricated content.
Approach: They propose to generate representative examples of non-factual summaries through infilling language models and train a robust fact-correction model to post-edit them to improve factual consistency.
Outcome: The proposed model outperforms previous methods in correcting factual errors on two popular summarization datasets.
Exploring the Potential of Multimodal LLM with Knowledge-Intensive Multimodal ASR (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in multimodal large language models have made significant progress in integrating information across various modalities, yet real-world applications in educational and scientific domains remain challenging.
Approach: They propose a task that focuses on transcribing scientific conference videos by leveraging visual information from slides to enhance the accuracy of technical terminologies.
Outcome: The proposed framework improves transcript quality through post-editing and improves performance over speech-only baselines.
Investigating the Helpfulness of Word-Level Quality Estimation for Post-Editing Machine Translation Output (2021.emnlp-main)

Copied to clipboard

Challenge: Post-editing (PE) machine translation (MT) output can save time and reduce errors.
Approach: They propose to use automatic word-level quality estimation to predict correctness of MT output to flag problematic output.
Outcome: The proposed model is not good enough to support human translations, but is based on a visualization reflecting uncertainty of the model.
PEGRL: Improving Machine Translation by Post-Editing Guided Reinforcement Learning (2026.findings-acl)

Copied to clipboard

Challenge: Reinforcement learning (RL) has shown strong promise for LLM-based machine translation . however, translation-oriented RL remains challenged by high-variance policy gradients induced by Monte Carlo baselines and large trajectory space that favors global exploration over fine-grained local optimization.
Approach: They propose a two-stage RL framework that uses post-editing as an auxiliary task to stabilize training and guide overall optimization.
Outcome: The proposed framework supports global exploration and fine-grained optimization while supporting global exploration.
Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreement (2025.emnlp-main)

Copied to clipboard

Challenge: Modern WQE techniques rely on expensive inference with large language models or ad-hoc training with large amounts of human-labeled data.
Approach: They propose to use word-level quality estimation to identify translation errors from the inner workings of translation models to quantify the impact of human label variation on metric performance.
Outcome: The proposed methods identify translation errors from the inner workings of translation models using human labels.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations