Papers by Matteo Negri
Does Simultaneous Speech Translation need Simultaneous Models? (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Simultaneous speech translation (SimulST) systems strive for high output quality but also low latency. |
| Approach: | They propose to train SimulST offline without additional training or adaptation . they also show offline training achieves similar or better quality compared to offline training . |
| Outcome: | The proposed model can serve both offline and simultaneous applications without additional training or adaptation. |
Cascade versus Direct Speech Translation: Do the Differences Still Make a Difference? (2021.acl-long)
Copied to clipboard
Luisa Bentivogli, Mauro Cettolo, Marco Gaido, Alina Karakanta, Alberto Martinelli, Matteo Negri, Marco Turchi
| Challenge: | a gap between direct approaches to speech translation (ST) and traditional cascade solutions has gradually decreased . a recent study found that the subtle differences observed in their behavior are not sufficient for humans neither to distinguish them nor to prefer one over the other. |
| Approach: | They compare state-of-the-art systems representative of the two paradigms . they find subtle differences observed in their behavior are not sufficient . |
| Outcome: | The proposed system is compared with state-of-the-art systems representative of the two paradigms. |
Tutorial: End-to-End Speech Translation (2021.eacl-tutorials)
Copied to clipboard
| Challenge: | Speech translation is the translation of speech in one language typically to text in another, traditionally accomplished through a combination of automatic speech recognition and machine translation. |
| Approach: | This tutorial introduces the techniques used in cutting-edge research on speech translation. |
| Outcome: | The proposed models achieve state-of-the-art performance with end-to-end speech translation for both high- and low-resource languages. |
Different Speech Translation Models Encode and Translate Speaker Gender Differently (2025.acl-short)
Copied to clipboard
| Challenge: | Recent studies on interpreting the hidden states of speech models have shown their ability to capture speaker-specific features, including gender. |
| Approach: | They propose to use probing methods to assess gender encoding across ST models. |
| Outcome: | The proposed models capture speaker-specific features, including gender, while older models do not . low gender encoding capabilities result in systems’ tendency toward a masculine default, a translation bias that is more pronounced in newer architectures. |
MuST-Cinema: a Speech-to-Subtitles corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing methods for subtitling are laborious and costly, says aaron sanchez . he says the current methods are laboriously complex and require manual work . |
| Approach: | They propose to use TED subtitles to build a multilingual speech translation corpus . they propose to annotate existing subtitling corpora with subtitle breaks . |
| Outcome: | The proposed model can be used to segment sentences into subtitles and reduces human work . the proposed model reduces the time and cost of human subtitling tasks . |
Gender in Danger? Evaluating Speech Translation Technology on the MuST-SHE Corpus (2020.acl-main)
Copied to clipboard
| Challenge: | a growing number of studies have examined the issue of gender bias in speech translation . a gender bias is a systemic problem that reproduces gender stereotypes discriminating women. |
| Approach: | They present the first thorough investigation of gender bias in speech translation . they compare audio technologies for English-Italian/French translations . |
| Outcome: | The proposed method compares different technologies on two languages, English and French. |
Evaluating the IWSLT2023 Speech Translation Tasks: Human Annotations, Automatic Metrics, and Segmentation (2024.lrec-main)
Copied to clipboard
Matthias Sperber, Ondřej Bojar, Barry Haddow, Dávid Javorský, Xutai Ma, Matteo Negri, Jan Niehues, Peter Polák, Elizabeth Salesky, Katsuhito Sudoh, Marco Turchi
| Challenge: | a meta-analysis of human evaluation for speech translation has not been conducted . noisy data and segmentation mismatches are challenges for automatic metrics . |
| Approach: | They propose an evaluation strategy based on automatic resegmentation and direct assessment with segment context. |
| Outcome: | The proposed evaluation strategy is robust and scores well-correlated with other types of human judgements. |
Hi Guys or Hi Folks? Benchmarking Gender-Neutral Machine Translation with the GeNTE Corpus (2023.emnlp-main)
Copied to clipboard
| Challenge: | Societal gender asymmetries and inequalities are perpetuated through language . MT often defaults to masculine representations by making undue binary gender assumptions . |
| Approach: | They propose a benchmark and automated evaluation methods to assess gender-neutral translation from English to Italian. |
| Outcome: | The proposed method is based on a survey on gender-neutral translation. |
Dodging the Data Bottleneck: Automatic Subtitling with Automatically Segmented ST Corpora (2022.aacl-short)
Copied to clipboard
| Challenge: | Existing models for subtitling require parallel data paired with audio inputs and textual translations. |
| Approach: | They propose to convert existing ST corpora into SubST resources without human intervention by exploiting audio and text in a multimodal fashion. |
| Outcome: | The proposed model achieves high segmentation quality in zero-shot conditions with manual and automatic segmentation. |
What the Harm? Quantifying the Tangible Impact of Gender Bias in Machine Translation with a Human-centered Study (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing bias measurements do not reflect the gender disparities found in machine translation. |
| Approach: | They conduct a human-centered study to examine if and to what extent bias in machine translation brings harms with tangible costs, such as quality of service gaps between women and men. |
| Outcome: | The findings advocate for human-centered approaches that can inform the societal impact of bias. |
Direct Speech Translation for Automatic Subtitling (2023.tacl-1)
Copied to clipboard
| Challenge: | Existing models for automatic subtitling generate subtitles in the target language along with their timestamps. |
| Approach: | They propose a direct speech translation model that generates subtitles in the target language along with their timestamps with a single model. |
| Outcome: | The proposed model outperforms a cascade system on 7 language pairs and on new benchmarks. |
Speechformer: Reducing Information Loss in Direct Speech Translation (2021.emnlp-main)
Copied to clipboard
| Challenge: | Current approaches to speech-to-text translation (ST) use a pipeline of two sub-components - an automatic speech recognition (ASR) and a machine translation (MT) model. |
| Approach: | They propose an architecture that avoids initial lossy compression and aggregates information only at a higher level according to more informed linguistic criteria. |
| Outcome: | The proposed architecture achieves gains of up to 0.8 BLEU on the standard MuST-C corpus and up to 4.0 BLUE in a low resource scenario. |
How to Split: the Effect of Word Segmentation on Gender Bias in Speech Translation (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for subword splitting penalize the representation of feminine linguistic markings. |
| Approach: | They propose a method that preserves subword splitting while leveraging character-based segmentation to properly translate gender. |
| Outcome: | The proposed approach preserves BPE overall translation quality while leveraging the higher ability of character-based segmentation to properly translate gender. |
StreamAtt: Direct Streaming Speech-to-Text Translation with Attention-based Audio History Selection (2024.acl-long)
Copied to clipboard
| Challenge: | Existing studies on streaming translation focus on SimulST only focusing on StreamST . StreamAtt is the first Stream ST policy and proposes StreamLAAL . |
| Approach: | They propose StreamAtt, the first StreamST policy, and StreamLAAL, the second Stream ST latency metric. |
| Outcome: | Experiments in 8 languages show that StreamAtt is more efficient than SimulST . StreamLAAL is the first StreamST latency metric comparable with existing metrics for Simul ST. |
SBAAM! Eliminating Transcript Dependency in Automatic Subtitling (2024.acl-long)
Copied to clipboard
| Challenge: | Subtitling is a crucial task for enhancing the accessibility of audiovisual content and relying on automatic transcripts for the three subtasks is uncharted territory. |
| Approach: | They propose a model capable of producing automatic subtitles, completely eliminating any dependence on intermediate transcripts also for timestamp prediction. |
| Outcome: | Experimental results show that the proposed model eliminates the need for intermediate transcripts for timestamp prediction across multiple language pairs and diverse conditions. |
ESCAPE: a Large-scale Synthetic Corpus for Automatic Post-Editing (L18-1)
Copied to clipboard
| Challenge: | eSCAPE is the largest freely-available Synthetic Corpus for Automatic Post-Editing released so far. |
| Approach: | a team of researchers develops a Synthetic Corpus for Automatic Post-Editing . eSCAPE is the largest freely-available Synthetic corpus for automatic post-editing released so far . the results prove that the models always improve MT quality with statistically significant gains . |
| Outcome: | eSCAPE is the largest freely-available Synthetic Corpus for Automatic Post-Editing released so far. |
Integrating Language Models into Direct Speech Translation: An Inference-Time Solution to Control Gender Inflection (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing solutions to control speaker-related gender inflections in ST involve dedicated model retraining on gender-labeled data. |
| Approach: | They propose to use a gender-based inference-time solution to control speaker-related gender inflections in ST by replacing the implicitly learned internal language model with gender-specific external LMs. |
| Outcome: | The proposed approach outperforms the base models and the best training-time mitigation strategy by up to 31.0 and 1.6 points in gender accuracy, respectively, for feminine forms. |
Speech Foundation Models and Crowdsourcing for Efficient, High-Quality Data Collection (2025.coling-main)
Copied to clipboard
| Challenge: | Existing methods for crowdsourcing data collection require a human workforce, which is hard to sustain. |
| Approach: | They propose to use Speech Foundation Models to automate validation processes . they find that SFMs can reduce reliance on human validation . |
| Outcome: | The proposed model reduces the reliance on human validation without degrading the quality of the final data. |
Gender Bias in Machine Translation (2021.tacl-1)
Copied to clipboard
| Challenge: | Interest in understanding, assessing, and mitigating gender bias in machine translation (MT) still lacks cohesion. |
| Approach: | They propose to review current conceptualizations of gender bias in machine translation (MT) they summarize previous studies and propose ways to mitigate bias. |
| Outcome: | This paper summarizes the current conceptualizations and proposes strategies to mitigate biases in machine translation (MT) . |
Evaluating Automatic Subtitling: Correlating Post-editing Effort and Automatic Metrics (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing metrics for automatic subtitling are not yet fully explored. |
| Approach: | They propose to use machine translation metrics to measure post-editing effort in automatic subtitling to collect data on product-, process- and participant-based data. |
| Outcome: | The proposed metrics correlate with measures of post-editing effort in automatic subtitling. |
The Two Shades of Dubbing in Neural Machine Translation (2020.coling-main)
Copied to clipboard
| Challenge: | Dubbing has two shades; synchronisation constraints are applied only when the actor’s mouth is visible on screen, while the translation is unconstrained for off-screen dubbing. |
| Approach: | They annotate an existing dubbing corpus for this dichotomy and find that on-screen dubbing is more difficult for MT than off-screen. |
| Outcome: | The results show that on-screen dubbing is more difficult for MT than off-screen translation, and that synchronisation constraints dramatically decrease translation quality for off- screen dubbing. |
When Good and Reproducible Results are a Giant with Feet of Clay: The Importance of Software Quality in NLP (2024.acl-long)
Copied to clipboard
| Challenge: | despite its crucial role in research experiments, code correctness is often presumed on the perceived quality of results. |
| Approach: | They propose to promote code-quality checklists to promote coding best practices . they propose to fix bugs in conformer implementations to mitigate this risk . |
| Outcome: | The proposed checklists aim to promote coding best practices and improve software quality within the NLP community. |
Under the Morphosyntactic Lens: A Multifaceted Evaluation of Gender Bias in Speech Translation (2022.acl-long)
Copied to clipboard
| Challenge: | grammatical gender languages are characterized by morphosyntactic chains of gender agreement marked on a variety of lexical items and parts-of-speech (POS). |
| Approach: | They propose to enrich the natural, gender-sensitive MuST-SHE corpus with two new linguistic annotation layers to explore gender bias. |
| Outcome: | The proposed models shed light on gender bias and its detection at several levels of granularity. |
MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languages (2024.emnlp-main)
Copied to clipboard
Marco Gaido, Sara Papi, Luisa Bentivogli, Alessio Brutti, Mauro Cettolo, Roberto Gretter, Marco Matassoni, Mohamed Nabih, Matteo Negri
| Challenge: | Existing speech FMs fall short of full compliance with open-source principles . existing models do not have model weights, code, and training data publicly available . |
| Approach: | They propose to use a CC-BY license to create open-source speech FMs for EU languages . they collect suitable training data by surveying automatic speech recognition datasets . |
| Outcome: | The proposed model can be used in the 24 official languages of the European Union. |
MuST-C: a Multilingual Speech Translation Corpus (N19-1)
Copied to clipboard
| Challenge: | Current research on spoken language translation (SLT) has to confront the scarcity of sizeable and publicly available training corpora. |
| Approach: | They propose a multilingual speech translation corpus that will facilitate the training of end-to-end systems for SLT from English into 8 languages. |
| Outcome: | The proposed multilingual speech translation corpus will facilitate the training of end-to-end systems for spoken language translation from English into 8 languages. |
CTC-based Compression for Direct Speech Translation (2021.eacl-main)
Copied to clipboard
| Challenge: | Existing studies have shown that a dynamic phone-informed compression of the input audio is beneficial for speech translation (ST). |
| Approach: | They propose a method which performs a phone-informed compression of the input audio in direct ST models by exploiting the Connectionist Temporal Classification (CTC) they demonstrate that their method brings a 1.3-1.5 BLEU improvement over a strong baseline on two language pairs (English-Italian and English-German) |
| Outcome: | The proposed method brings a 1.3-1.5 BLEU improvement over a strong baseline on two language pairs (English-Italian and English-German) it reduces memory footprint by more than 10%, and is faster than previous approaches. |
Translation in the Hands of Many: Centering Lay Users in Machine Translation Interactions (2025.emnlp-main)
Copied to clipboard
| Challenge: | Multilingual demands and accessibility have made MT a global tool . however, the understanding of MT consumed by such a diverse group of users remains limited. |
| Approach: | They first trace the evolution of MT user profiles, focusing on non-experts and how their engagement with technology may shift with the rise of LLMs. |
| Outcome: | The proposed approach will help to align MT with user needs and improve the quality of the language. |
Mind the Inclusivity Gap: Multilingual Gender-Neutral Translation Evaluation with mGeNTE (2025.emnlp-main)
Copied to clipboard
Beatrice Savoldi, Giuseppe Attanasio, Eleonora Cupin, Eleni Gkovedarou, Janiça Hackenbuchner, Anne Lauscher, Matteo Negri, Andrea Piergentili, Manjinder Thind, Luisa Bentivogli
| Challenge: | Genderneutral translation (GNT) is a linguistic strategy towards fairer communication across languages. |
| Approach: | They propose to use a multilingual evaluation resource to evaluate inclusive translation with state-of-the-art instruction-following language models (LMs) |
| Outcome: | The proposed model can recognize when neutrality is appropriate, but cannot consistently produce neutral translations, limiting their usability. |
Is “moby dick” a Whale or a Bird? Named Entities and Terminology in Speech Translation (2021.emnlp-main)
Copied to clipboard
| Challenge: | Among rare words, named entities and domain-specific terms are crucial . previous studies have neglected these important words due to limited options . |
| Approach: | They propose a benchmark to evaluate automatic translation systems for rare words . named entities and domain-specific terms are crucial for their translation . |
| Outcome: | The proposed benchmark is based on European Parliament speeches annotated with NEs and terminology. |
Attention as a Guide for Simultaneous Speech Translation (2023.acl-long)
Copied to clipboard
| Challenge: | EDAtt uses attention patterns to determine when to emit partial translations . results show that it yields better results compared to existing SimulST policies . |
| Approach: | They propose an adaptive policy that exploits attention patterns between audio source and target textual translation to guide an offline-trained ST model during simultaneous inference. |
| Outcome: | The proposed policy yields better results on en->de, compared to the current state of the art. |
A Prompt Response to the Demand for Automatic Gender-Neutral Translation (2024.eacl-short)
Copied to clipboard
| Challenge: | Advancements in machine translation (MT) are hindered by the lack of dedicated parallel data, which are necessary to adapt MT systems to satisfy neutral constraints. |
| Approach: | They propose to use GPT-4 to generate GNTs that avoid bias and undue binary assumptions by comparing MT with the popular GPT-3 model. |
| Outcome: | The proposed model outperforms the existing model and provides valuable insights into the potential and challenges associated with prompting for neutrality. |
How Do Hyenas Deal with Human Speech? Speech Recognition and Translation with ConfHyena (2024.lrec-main)
Copied to clipboard
| Challenge: | Currently, attention-based models face computational hurdles in processing long sequences due to its quadratic complexity. |
| Approach: | They propose a conformer whose encoder self-attentions are replaced with Hyena for speech processing . they propose 'confhyena' model that reduces training time by 27% at minimal cost . |
| Outcome: | The proposed model reduces training time by 27% at the cost of minimal quality degradation. |
Machine Translation for Machines: the Sentiment Classification Use Case (D19-1)
Copied to clipboard
| Challenge: | Traditionally, machine translation (MT) pursues a "human-oriented" objective: generating fluent output for a downstream task. |
| Approach: | They propose a neural machine translation approach that uses weak feedback to generate translations that are best suited for a downstream task. |
| Outcome: | The proposed approach outperforms general-purpose models and reinforcement learning methods on German and Italian tweets. |
Breeding Gender-aware Direct Speech Translation Systems (2020.coling-main)
Copied to clipboard
| Challenge: | In automatic speech translation, traditional cascade approaches involving separate transcription and translation steps are giving ground to more robust direct solutions. |
| Approach: | They compare different approaches to inform direct ST models about the speaker’s gender and test their ability to handle gender translation from English into Italian and French. |
| Outcome: | The proposed models outperform strong but gender-unaware direct ST models in the translation of English into Italian and French. |