Papers by Marco Turchi

24 papers
Does Simultaneous Speech Translation need Simultaneous Models? (2022.findings-emnlp)

Copied to clipboard

Challenge: Simultaneous speech translation (SimulST) systems strive for high output quality but also low latency.
Approach: They propose to train SimulST offline without additional training or adaptation . they also show offline training achieves similar or better quality compared to offline training .
Outcome: The proposed model can serve both offline and simultaneous applications without additional training or adaptation.
Cascade versus Direct Speech Translation: Do the Differences Still Make a Difference? (2021.acl-long)

Copied to clipboard

Challenge: a gap between direct approaches to speech translation (ST) and traditional cascade solutions has gradually decreased . a recent study found that the subtle differences observed in their behavior are not sufficient for humans neither to distinguish them nor to prefer one over the other.
Approach: They compare state-of-the-art systems representative of the two paradigms . they find subtle differences observed in their behavior are not sufficient .
Outcome: The proposed system is compared with state-of-the-art systems representative of the two paradigms.
Tutorial: End-to-End Speech Translation (2021.eacl-tutorials)

Copied to clipboard

Challenge: Speech translation is the translation of speech in one language typically to text in another, traditionally accomplished through a combination of automatic speech recognition and machine translation.
Approach: This tutorial introduces the techniques used in cutting-edge research on speech translation.
Outcome: The proposed models achieve state-of-the-art performance with end-to-end speech translation for both high- and low-resource languages.
MuST-Cinema: a Speech-to-Subtitles corpus (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods for subtitling are laborious and costly, says aaron sanchez . he says the current methods are laboriously complex and require manual work .
Approach: They propose to use TED subtitles to build a multilingual speech translation corpus . they propose to annotate existing subtitling corpora with subtitle breaks .
Outcome: The proposed model can be used to segment sentences into subtitles and reduces human work . the proposed model reduces the time and cost of human subtitling tasks .
CLAD-ST: Contrastive Learning with Adversarial Data for Robust Speech Translation (2023.emnlp-main)

Copied to clipboard

Challenge: Cascaded approach is the most popular choice for speech translation, but lacks robustness when dealing with noisy inputs.
Approach: They propose a cascaded approach that uses an automatic speech recognition model and a machine translation model to translate speech in one language to text in another language.
Outcome: The proposed approach achieves significant gains of up to 3 BLEU scores in English-German and English-French speech translation without hurting the translation quality on clean text.
Gender in Danger? Evaluating Speech Translation Technology on the MuST-SHE Corpus (2020.acl-main)

Copied to clipboard

Challenge: a growing number of studies have examined the issue of gender bias in speech translation . a gender bias is a systemic problem that reproduces gender stereotypes discriminating women.
Approach: They present the first thorough investigation of gender bias in speech translation . they compare audio technologies for English-Italian/French translations .
Outcome: The proposed method compares different technologies on two languages, English and French.
Evaluating the IWSLT2023 Speech Translation Tasks: Human Annotations, Automatic Metrics, and Segmentation (2024.lrec-main)

Copied to clipboard

Challenge: a meta-analysis of human evaluation for speech translation has not been conducted . noisy data and segmentation mismatches are challenges for automatic metrics .
Approach: They propose an evaluation strategy based on automatic resegmentation and direct assessment with segment context.
Outcome: The proposed evaluation strategy is robust and scores well-correlated with other types of human judgements.
Dodging the Data Bottleneck: Automatic Subtitling with Automatically Segmented ST Corpora (2022.aacl-short)

Copied to clipboard

Challenge: Existing models for subtitling require parallel data paired with audio inputs and textual translations.
Approach: They propose to convert existing ST corpora into SubST resources without human intervention by exploiting audio and text in a multimodal fashion.
Outcome: The proposed model achieves high segmentation quality in zero-shot conditions with manual and automatic segmentation.
Direct Speech Translation for Automatic Subtitling (2023.tacl-1)

Copied to clipboard

Challenge: Existing models for automatic subtitling generate subtitles in the target language along with their timestamps.
Approach: They propose a direct speech translation model that generates subtitles in the target language along with their timestamps with a single model.
Outcome: The proposed model outperforms a cascade system on 7 language pairs and on new benchmarks.
Speechformer: Reducing Information Loss in Direct Speech Translation (2021.emnlp-main)

Copied to clipboard

Challenge: Current approaches to speech-to-text translation (ST) use a pipeline of two sub-components - an automatic speech recognition (ASR) and a machine translation (MT) model.
Approach: They propose an architecture that avoids initial lossy compression and aggregates information only at a higher level according to more informed linguistic criteria.
Outcome: The proposed architecture achieves gains of up to 0.8 BLEU on the standard MuST-C corpus and up to 4.0 BLUE in a low resource scenario.
How to Split: the Effect of Word Segmentation on Gender Bias in Speech Translation (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods for subword splitting penalize the representation of feminine linguistic markings.
Approach: They propose a method that preserves subword splitting while leveraging character-based segmentation to properly translate gender.
Outcome: The proposed approach preserves BPE overall translation quality while leveraging the higher ability of character-based segmentation to properly translate gender.
ESCAPE: a Large-scale Synthetic Corpus for Automatic Post-Editing (L18-1)

Copied to clipboard

Challenge: eSCAPE is the largest freely-available Synthetic Corpus for Automatic Post-Editing released so far.
Approach: a team of researchers develops a Synthetic Corpus for Automatic Post-Editing . eSCAPE is the largest freely-available Synthetic corpus for automatic post-editing released so far . the results prove that the models always improve MT quality with statistically significant gains .
Outcome: eSCAPE is the largest freely-available Synthetic Corpus for Automatic Post-Editing released so far.
Gender Bias in Machine Translation (2021.tacl-1)

Copied to clipboard

Challenge: Interest in understanding, assessing, and mitigating gender bias in machine translation (MT) still lacks cohesion.
Approach: They propose to review current conceptualizations of gender bias in machine translation (MT) they summarize previous studies and propose ways to mitigate bias.
Outcome: This paper summarizes the current conceptualizations and proposes strategies to mitigate biases in machine translation (MT) .
Cross-lingual Evaluation of Multilingual Text Generation (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for multilingual text generation are limited by language and data leakage.
Approach: They propose an annotation-free cross-lingual evaluation protocol for multilingual text generation . they first generate English references from the translated non-English inputs into English .
Outcome: The proposed protocol shows a high correlation to the reference-based ROUGE metric in four languages on news text summarization.
The Two Shades of Dubbing in Neural Machine Translation (2020.coling-main)

Copied to clipboard

Challenge: Dubbing has two shades; synchronisation constraints are applied only when the actor’s mouth is visible on screen, while the translation is unconstrained for off-screen dubbing.
Approach: They annotate an existing dubbing corpus for this dichotomy and find that on-screen dubbing is more difficult for MT than off-screen.
Outcome: The results show that on-screen dubbing is more difficult for MT than off-screen translation, and that synchronisation constraints dramatically decrease translation quality for off- screen dubbing.
Under the Morphosyntactic Lens: A Multifaceted Evaluation of Gender Bias in Speech Translation (2022.acl-long)

Copied to clipboard

Challenge: grammatical gender languages are characterized by morphosyntactic chains of gender agreement marked on a variety of lexical items and parts-of-speech (POS).
Approach: They propose to enrich the natural, gender-sensitive MuST-SHE corpus with two new linguistic annotation layers to explore gender bias.
Outcome: The proposed models shed light on gender bias and its detection at several levels of granularity.
MuST-C: a Multilingual Speech Translation Corpus (N19-1)

Copied to clipboard

Challenge: Current research on spoken language translation (SLT) has to confront the scarcity of sizeable and publicly available training corpora.
Approach: They propose a multilingual speech translation corpus that will facilitate the training of end-to-end systems for SLT from English into 8 languages.
Outcome: The proposed multilingual speech translation corpus will facilitate the training of end-to-end systems for spoken language translation from English into 8 languages.
CTC-based Compression for Direct Speech Translation (2021.eacl-main)

Copied to clipboard

Challenge: Existing studies have shown that a dynamic phone-informed compression of the input audio is beneficial for speech translation (ST).
Approach: They propose a method which performs a phone-informed compression of the input audio in direct ST models by exploiting the Connectionist Temporal Classification (CTC) they demonstrate that their method brings a 1.3-1.5 BLEU improvement over a strong baseline on two language pairs (English-Italian and English-German)
Outcome: The proposed method brings a 1.3-1.5 BLEU improvement over a strong baseline on two language pairs (English-Italian and English-German) it reduces memory footprint by more than 10%, and is faster than previous approaches.
Gradient-based Gradual Pruning for Language-Specific Multilingual Neural Machine Translation (2023.emnlp-main)

Copied to clipboard

Challenge: Multilingual neural machine translation suffers from performance degradation in high-resource languages compared to bilingual counterparts.
Approach: They propose a gradient-based gradual pruning technique for multilingual neural machine translation that allows for partial parameter sharing across language pairs to alleviate interference.
Outcome: The proposed approach yields a notable performance gain on IWSLT and WMT datasets.
Is “moby dick” a Whale or a Bird? Named Entities and Terminology in Speech Translation (2021.emnlp-main)

Copied to clipboard

Challenge: Among rare words, named entities and domain-specific terms are crucial . previous studies have neglected these important words due to limited options .
Approach: They propose a benchmark to evaluate automatic translation systems for rare words . named entities and domain-specific terms are crucial for their translation .
Outcome: The proposed benchmark is based on European Parliament speeches annotated with NEs and terminology.
Attention as a Guide for Simultaneous Speech Translation (2023.acl-long)

Copied to clipboard

Challenge: EDAtt uses attention patterns to determine when to emit partial translations . results show that it yields better results compared to existing SimulST policies .
Approach: They propose an adaptive policy that exploits attention patterns between audio source and target textual translation to guide an offline-trained ST model during simultaneous inference.
Outcome: The proposed policy yields better results on en->de, compared to the current state of the art.
Machine Translation for Machines: the Sentiment Classification Use Case (D19-1)

Copied to clipboard

Challenge: Traditionally, machine translation (MT) pursues a "human-oriented" objective: generating fluent output for a downstream task.
Approach: They propose a neural machine translation approach that uses weak feedback to generate translations that are best suited for a downstream task.
Outcome: The proposed approach outperforms general-purpose models and reinforcement learning methods on German and Italian tweets.
Select, Prompt, Filter: Distilling Large Language Models for Summarizing Conversations (2023.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) can be expensive to train, deploy, and use for specific natural language generation tasks.
Approach: They propose a method to distill ChatGPT and fine-tune smaller LMs for summarizing forum conversations using a semantic similarity metric.
Outcome: The proposed method leads to significant improvements of up to 6.6 ROUGE-2 score by leveraging sufficient in-domain pseudo-labeled data over standard KD approach given the same size of training data.
Breeding Gender-aware Direct Speech Translation Systems (2020.coling-main)

Copied to clipboard

Challenge: In automatic speech translation, traditional cascade approaches involving separate transcription and translation steps are giving ground to more robust direct solutions.
Approach: They compare different approaches to inform direct ST models about the speaker’s gender and test their ability to handle gender translation from English into Italian and French.
Outcome: The proposed models outperform strong but gender-unaware direct ST models in the translation of English into Italian and French.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations