Statistical Analysis of Missing Translation in Simultaneous Interpretation Using A Large-scale Bilingual Speech Corpus (L18-1)
Copied to clipboard
| Challenge: | Various types of omissions have been described in simultaneous interpretation to improve interpretation quality or train interpreters. |
| Approach: | They analyze missing translations in simultaneous interpretations using a large-scale bilingual speech corpus. |
| Outcome: | The authors found that a high proportion of adverbs were missed in the translations . the authors suggest that omissions can be improved to improve interpretation quality . |
Similar Papers
Simultaneous Interpretation Corpus Construction by Large Language Models in Distant Language Pair (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing siMT corpora are limited due to high costs and limited annotator capabilities. |
| Approach: | They propose a method to convert ST corpora into interpretation-style corpors by fine-tuning models with Large Language Models. |
| Outcome: | The proposed method reduces latency while achieving better quality compared to other models. |
The Role of Mixed-Language Documents for Multilingual Large Language Model Pretraining (2026.acl-long)
Copied to clipboard
Jiandong Shao, Raphael Tang, Crystina Zhang, Karin Sevegnani, Pontus Stenetorp, Jianfei Yang, Yao Lu
| Challenge: | Existing research suggests that multilingual large language models can achieve impressive cross-lingual understanding despite largely monolingual pretraining. |
| Approach: | They compare a monolingual-only corpus with a standard web corpus that removes all multilingual documents and then retrain the models from scratch under controlled conditions. |
| Outcome: | The results show that removing bilingual data causes translation performance to drop 56% in BLEU, whereas code-switching contributes minimally. |
Is It Good Data for Multilingual Instruction Tuning or Just Bad Multilingual Evaluation for Large Language Models? (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing practices of fine-tuning and evaluating multilingual large language models may not align with this objective due to a heavy reliance on translation. |
| Approach: | They propose to use translated or native instruction data to fine-tune multilingual large language models. |
| Outcome: | The proposed model can be fine tuned and evaluated in multilingual large language models . the results show that native or translated data can be used to compare model performance . |
Lost in Interpretation: Predicting Untranslated Terminology in Simultaneous Interpretation (N19-1)
Copied to clipboard
| Challenge: | Experimental results on a newly-annotated version of the NAIST Simultaneous Translation Corpus indicate the promise of our proposed method. |
| Approach: | They propose a task of predicting which terminology simultaneous interpreters will leave untranslated using supervised sequence taggers. |
| Outcome: | The proposed method predicts which terminology interpreters leave untranslated . it is based on an annotated version of the NAIST Simultaneous Translation Corpus . |
Language Lives in Sparse Dimensions: Toward Interpretable and Efficient Multilingual Control for Large Language Models (2026.eacl-long)
Copied to clipboard
| Challenge: | Prior studies show that large language models map multilingual content into English-aligned representations at intermediate layers before projecting them back into target-language token spaces in the later layers. |
| Approach: | They propose a method to identify and manipulate dimensions that are sparse and sparsity-based . they propose to use as few as 50 sentences of either parallel or monolingual data to manipulate these dimensions . |
| Outcome: | Experiments on a multilingual generation control task show the interpretability of these dimensions. |
Fine-Grained Analysis of Cross-Linguistic Syntactic Divergences (2020.acl-main)
Copied to clipboard
Dmitry Nikolaev, Ofir Arviv, Taelin Karidi, Neta Kenneth, Veronika Mitnik, Lilja Maria Saeboe, Omri Abend
| Challenge: | Existing work on quantifying the prevalence of syntactic divergences across languages has not been done. |
| Approach: | They propose a framework for extracting divergence patterns for any language pair from a parallel corpus building on Universal Dependencies. |
| Outcome: | The proposed framework provides a detailed picture of cross-language divergences, generalizes previous approaches, and lends itself to full automation. |
Autocorrect in the Process of Translation — Multi-task Learning Improves Dialogue Machine Translation (2021.naacl-industry)
Copied to clipboard
| Challenge: | Existing neural machine translation models are not able to translate dialogues in real life scenarios. |
| Approach: | They propose a joint learning method to identify omission and typos and utilize context to translate dialogue utterances. |
| Outcome: | The proposed method improves translation quality by 3.2 BLEU over baselines and recovers omitted pronouns by 47.16%. |
NAIST-SIC-Aligned: An Aligned English-Japanese Simultaneous Interpretation Corpus (2024.lrec-main)
Copied to clipboard
| Challenge: | Simultaneous interpretation data is a task where an utterance is translated in real-time. |
| Approach: | They propose to use an automatically-aligned parallel English-Japanese SI dataset to make it suitable for model training. |
| Outcome: | The proposed model improves translation quality and latency over baselines. |
Simultaneous Translation (2020.emnlp-tutorials)
Copied to clipboard
| Challenge: | Simultaneous translation is a problem that has long been considered one of the hardest problems in AI . this tutorial will provide a deep understanding of the history and the recent advances in simultaneous translation. |
| Approach: | This tutorial will examine the design and evaluation of policies for simultaneous translation . it will provide an overview of the history and recent advances in simultaneous translation. |
| Outcome: | This tutorial will examine the design and evaluation of policies for simultaneous translation . |
Translation Errors Significantly Impact Low-Resource Languages in Cross-Lingual Learning (2024.eacl-short)
Copied to clipboard
| Challenge: | XNLI benchmarks use parallel versions of English evaluation sets in multiple target languages . a recent study found that translation errors exist in some low-resource languages resulting in incorrect estimates of cross-lingual transfer . |
| Approach: | They propose to measure the gap in performance between zero-shot evaluations on human-translated and machine-transcribed target text across multiple target languages. |
| Outcome: | The proposed benchmarks show that translation errors exist for Hindi and Urdu . the results corroborate previous studies that found translation errors in Hindi and urdu despite translation errors. |