Challenge: Various types of omissions have been described in simultaneous interpretation to improve interpretation quality or train interpreters.
Approach: They analyze missing translations in simultaneous interpretations using a large-scale bilingual speech corpus.
Outcome: The authors found that a high proportion of adverbs were missed in the translations . the authors suggest that omissions can be improved to improve interpretation quality .

Similar Papers

Simultaneous Interpretation Corpus Construction by Large Language Models in Distant Language Pair (2024.emnlp-main)

Copied to clipboard

Challenge: Existing siMT corpora are limited due to high costs and limited annotator capabilities.
Approach: They propose a method to convert ST corpora into interpretation-style corpors by fine-tuning models with Large Language Models.
Outcome: The proposed method reduces latency while achieving better quality compared to other models.
The Role of Mixed-Language Documents for Multilingual Large Language Model Pretraining (2026.acl-long)

Copied to clipboard

Challenge: Existing research suggests that multilingual large language models can achieve impressive cross-lingual understanding despite largely monolingual pretraining.
Approach: They compare a monolingual-only corpus with a standard web corpus that removes all multilingual documents and then retrain the models from scratch under controlled conditions.
Outcome: The results show that removing bilingual data causes translation performance to drop 56% in BLEU, whereas code-switching contributes minimally.
Is It Good Data for Multilingual Instruction Tuning or Just Bad Multilingual Evaluation for Large Language Models? (2024.emnlp-main)

Copied to clipboard

Challenge: Existing practices of fine-tuning and evaluating multilingual large language models may not align with this objective due to a heavy reliance on translation.
Approach: They propose to use translated or native instruction data to fine-tune multilingual large language models.
Outcome: The proposed model can be fine tuned and evaluated in multilingual large language models . the results show that native or translated data can be used to compare model performance .
Lost in Interpretation: Predicting Untranslated Terminology in Simultaneous Interpretation (N19-1)

Copied to clipboard

Challenge: Experimental results on a newly-annotated version of the NAIST Simultaneous Translation Corpus indicate the promise of our proposed method.
Approach: They propose a task of predicting which terminology simultaneous interpreters will leave untranslated using supervised sequence taggers.
Outcome: The proposed method predicts which terminology interpreters leave untranslated . it is based on an annotated version of the NAIST Simultaneous Translation Corpus .
Language Lives in Sparse Dimensions: Toward Interpretable and Efficient Multilingual Control for Large Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Prior studies show that large language models map multilingual content into English-aligned representations at intermediate layers before projecting them back into target-language token spaces in the later layers.
Approach: They propose a method to identify and manipulate dimensions that are sparse and sparsity-based . they propose to use as few as 50 sentences of either parallel or monolingual data to manipulate these dimensions .
Outcome: Experiments on a multilingual generation control task show the interpretability of these dimensions.
Fine-Grained Analysis of Cross-Linguistic Syntactic Divergences (2020.acl-main)

Copied to clipboard

Challenge: Existing work on quantifying the prevalence of syntactic divergences across languages has not been done.
Approach: They propose a framework for extracting divergence patterns for any language pair from a parallel corpus building on Universal Dependencies.
Outcome: The proposed framework provides a detailed picture of cross-language divergences, generalizes previous approaches, and lends itself to full automation.
Autocorrect in the Process of Translation — Multi-task Learning Improves Dialogue Machine Translation (2021.naacl-industry)

Copied to clipboard

Challenge: Existing neural machine translation models are not able to translate dialogues in real life scenarios.
Approach: They propose a joint learning method to identify omission and typos and utilize context to translate dialogue utterances.
Outcome: The proposed method improves translation quality by 3.2 BLEU over baselines and recovers omitted pronouns by 47.16%.
NAIST-SIC-Aligned: An Aligned English-Japanese Simultaneous Interpretation Corpus (2024.lrec-main)

Copied to clipboard

Challenge: Simultaneous interpretation data is a task where an utterance is translated in real-time.
Approach: They propose to use an automatically-aligned parallel English-Japanese SI dataset to make it suitable for model training.
Outcome: The proposed model improves translation quality and latency over baselines.
Simultaneous Translation (2020.emnlp-tutorials)

Copied to clipboard

Challenge: Simultaneous translation is a problem that has long been considered one of the hardest problems in AI . this tutorial will provide a deep understanding of the history and the recent advances in simultaneous translation.
Approach: This tutorial will examine the design and evaluation of policies for simultaneous translation . it will provide an overview of the history and recent advances in simultaneous translation.
Outcome: This tutorial will examine the design and evaluation of policies for simultaneous translation .
Translation Errors Significantly Impact Low-Resource Languages in Cross-Lingual Learning (2024.eacl-short)

Copied to clipboard

Challenge: XNLI benchmarks use parallel versions of English evaluation sets in multiple target languages . a recent study found that translation errors exist in some low-resource languages resulting in incorrect estimates of cross-lingual transfer .
Approach: They propose to measure the gap in performance between zero-shot evaluations on human-translated and machine-transcribed target text across multiple target languages.
Outcome: The proposed benchmarks show that translation errors exist for Hindi and Urdu . the results corroborate previous studies that found translation errors in Hindi and urdu despite translation errors.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations