Papers by Matteo Negri

34 papers
Does Simultaneous Speech Translation need Simultaneous Models? (2022.findings-emnlp)

Copied to clipboard

Challenge: Simultaneous speech translation (SimulST) systems strive for high output quality but also low latency.
Approach: They propose to train SimulST offline without additional training or adaptation . they also show offline training achieves similar or better quality compared to offline training .
Outcome: The proposed model can serve both offline and simultaneous applications without additional training or adaptation.
Cascade versus Direct Speech Translation: Do the Differences Still Make a Difference? (2021.acl-long)

Copied to clipboard

Challenge: a gap between direct approaches to speech translation (ST) and traditional cascade solutions has gradually decreased . a recent study found that the subtle differences observed in their behavior are not sufficient for humans neither to distinguish them nor to prefer one over the other.
Approach: They compare state-of-the-art systems representative of the two paradigms . they find subtle differences observed in their behavior are not sufficient .
Outcome: The proposed system is compared with state-of-the-art systems representative of the two paradigms.
Tutorial: End-to-End Speech Translation (2021.eacl-tutorials)

Copied to clipboard

Challenge: Speech translation is the translation of speech in one language typically to text in another, traditionally accomplished through a combination of automatic speech recognition and machine translation.
Approach: This tutorial introduces the techniques used in cutting-edge research on speech translation.
Outcome: The proposed models achieve state-of-the-art performance with end-to-end speech translation for both high- and low-resource languages.
Different Speech Translation Models Encode and Translate Speaker Gender Differently (2025.acl-short)

Copied to clipboard

Challenge: Recent studies on interpreting the hidden states of speech models have shown their ability to capture speaker-specific features, including gender.
Approach: They propose to use probing methods to assess gender encoding across ST models.
Outcome: The proposed models capture speaker-specific features, including gender, while older models do not . low gender encoding capabilities result in systems’ tendency toward a masculine default, a translation bias that is more pronounced in newer architectures.
MuST-Cinema: a Speech-to-Subtitles corpus (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods for subtitling are laborious and costly, says aaron sanchez . he says the current methods are laboriously complex and require manual work .
Approach: They propose to use TED subtitles to build a multilingual speech translation corpus . they propose to annotate existing subtitling corpora with subtitle breaks .
Outcome: The proposed model can be used to segment sentences into subtitles and reduces human work . the proposed model reduces the time and cost of human subtitling tasks .
Gender in Danger? Evaluating Speech Translation Technology on the MuST-SHE Corpus (2020.acl-main)

Copied to clipboard

Challenge: a growing number of studies have examined the issue of gender bias in speech translation . a gender bias is a systemic problem that reproduces gender stereotypes discriminating women.
Approach: They present the first thorough investigation of gender bias in speech translation . they compare audio technologies for English-Italian/French translations .
Outcome: The proposed method compares different technologies on two languages, English and French.
Evaluating the IWSLT2023 Speech Translation Tasks: Human Annotations, Automatic Metrics, and Segmentation (2024.lrec-main)

Copied to clipboard

Challenge: a meta-analysis of human evaluation for speech translation has not been conducted . noisy data and segmentation mismatches are challenges for automatic metrics .
Approach: They propose an evaluation strategy based on automatic resegmentation and direct assessment with segment context.
Outcome: The proposed evaluation strategy is robust and scores well-correlated with other types of human judgements.
Hi Guys or Hi Folks? Benchmarking Gender-Neutral Machine Translation with the GeNTE Corpus (2023.emnlp-main)

Copied to clipboard

Challenge: Societal gender asymmetries and inequalities are perpetuated through language . MT often defaults to masculine representations by making undue binary gender assumptions .
Approach: They propose a benchmark and automated evaluation methods to assess gender-neutral translation from English to Italian.
Outcome: The proposed method is based on a survey on gender-neutral translation.
Dodging the Data Bottleneck: Automatic Subtitling with Automatically Segmented ST Corpora (2022.aacl-short)

Copied to clipboard

Challenge: Existing models for subtitling require parallel data paired with audio inputs and textual translations.
Approach: They propose to convert existing ST corpora into SubST resources without human intervention by exploiting audio and text in a multimodal fashion.
Outcome: The proposed model achieves high segmentation quality in zero-shot conditions with manual and automatic segmentation.
What the Harm? Quantifying the Tangible Impact of Gender Bias in Machine Translation with a Human-centered Study (2024.emnlp-main)

Copied to clipboard

Challenge: Existing bias measurements do not reflect the gender disparities found in machine translation.
Approach: They conduct a human-centered study to examine if and to what extent bias in machine translation brings harms with tangible costs, such as quality of service gaps between women and men.
Outcome: The findings advocate for human-centered approaches that can inform the societal impact of bias.
Direct Speech Translation for Automatic Subtitling (2023.tacl-1)

Copied to clipboard

Challenge: Existing models for automatic subtitling generate subtitles in the target language along with their timestamps.
Approach: They propose a direct speech translation model that generates subtitles in the target language along with their timestamps with a single model.
Outcome: The proposed model outperforms a cascade system on 7 language pairs and on new benchmarks.
Speechformer: Reducing Information Loss in Direct Speech Translation (2021.emnlp-main)

Copied to clipboard

Challenge: Current approaches to speech-to-text translation (ST) use a pipeline of two sub-components - an automatic speech recognition (ASR) and a machine translation (MT) model.
Approach: They propose an architecture that avoids initial lossy compression and aggregates information only at a higher level according to more informed linguistic criteria.
Outcome: The proposed architecture achieves gains of up to 0.8 BLEU on the standard MuST-C corpus and up to 4.0 BLUE in a low resource scenario.
How to Split: the Effect of Word Segmentation on Gender Bias in Speech Translation (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods for subword splitting penalize the representation of feminine linguistic markings.
Approach: They propose a method that preserves subword splitting while leveraging character-based segmentation to properly translate gender.
Outcome: The proposed approach preserves BPE overall translation quality while leveraging the higher ability of character-based segmentation to properly translate gender.
StreamAtt: Direct Streaming Speech-to-Text Translation with Attention-based Audio History Selection (2024.acl-long)

Copied to clipboard

Challenge: Existing studies on streaming translation focus on SimulST only focusing on StreamST . StreamAtt is the first Stream ST policy and proposes StreamLAAL .
Approach: They propose StreamAtt, the first StreamST policy, and StreamLAAL, the second Stream ST latency metric.
Outcome: Experiments in 8 languages show that StreamAtt is more efficient than SimulST . StreamLAAL is the first StreamST latency metric comparable with existing metrics for Simul ST.
SBAAM! Eliminating Transcript Dependency in Automatic Subtitling (2024.acl-long)

Copied to clipboard

Challenge: Subtitling is a crucial task for enhancing the accessibility of audiovisual content and relying on automatic transcripts for the three subtasks is uncharted territory.
Approach: They propose a model capable of producing automatic subtitles, completely eliminating any dependence on intermediate transcripts also for timestamp prediction.
Outcome: Experimental results show that the proposed model eliminates the need for intermediate transcripts for timestamp prediction across multiple language pairs and diverse conditions.
ESCAPE: a Large-scale Synthetic Corpus for Automatic Post-Editing (L18-1)

Copied to clipboard

Challenge: eSCAPE is the largest freely-available Synthetic Corpus for Automatic Post-Editing released so far.
Approach: a team of researchers develops a Synthetic Corpus for Automatic Post-Editing . eSCAPE is the largest freely-available Synthetic corpus for automatic post-editing released so far . the results prove that the models always improve MT quality with statistically significant gains .
Outcome: eSCAPE is the largest freely-available Synthetic Corpus for Automatic Post-Editing released so far.
Integrating Language Models into Direct Speech Translation: An Inference-Time Solution to Control Gender Inflection (2023.emnlp-main)

Copied to clipboard

Challenge: Existing solutions to control speaker-related gender inflections in ST involve dedicated model retraining on gender-labeled data.
Approach: They propose to use a gender-based inference-time solution to control speaker-related gender inflections in ST by replacing the implicitly learned internal language model with gender-specific external LMs.
Outcome: The proposed approach outperforms the base models and the best training-time mitigation strategy by up to 31.0 and 1.6 points in gender accuracy, respectively, for feminine forms.
Speech Foundation Models and Crowdsourcing for Efficient, High-Quality Data Collection (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for crowdsourcing data collection require a human workforce, which is hard to sustain.
Approach: They propose to use Speech Foundation Models to automate validation processes . they find that SFMs can reduce reliance on human validation .
Outcome: The proposed model reduces the reliance on human validation without degrading the quality of the final data.
Gender Bias in Machine Translation (2021.tacl-1)

Copied to clipboard

Challenge: Interest in understanding, assessing, and mitigating gender bias in machine translation (MT) still lacks cohesion.
Approach: They propose to review current conceptualizations of gender bias in machine translation (MT) they summarize previous studies and propose ways to mitigate bias.
Outcome: This paper summarizes the current conceptualizations and proposes strategies to mitigate biases in machine translation (MT) .
Evaluating Automatic Subtitling: Correlating Post-editing Effort and Automatic Metrics (2024.lrec-main)

Copied to clipboard

Challenge: Existing metrics for automatic subtitling are not yet fully explored.
Approach: They propose to use machine translation metrics to measure post-editing effort in automatic subtitling to collect data on product-, process- and participant-based data.
Outcome: The proposed metrics correlate with measures of post-editing effort in automatic subtitling.
The Two Shades of Dubbing in Neural Machine Translation (2020.coling-main)

Copied to clipboard

Challenge: Dubbing has two shades; synchronisation constraints are applied only when the actor’s mouth is visible on screen, while the translation is unconstrained for off-screen dubbing.
Approach: They annotate an existing dubbing corpus for this dichotomy and find that on-screen dubbing is more difficult for MT than off-screen.
Outcome: The results show that on-screen dubbing is more difficult for MT than off-screen translation, and that synchronisation constraints dramatically decrease translation quality for off- screen dubbing.
When Good and Reproducible Results are a Giant with Feet of Clay: The Importance of Software Quality in NLP (2024.acl-long)

Copied to clipboard

Challenge: despite its crucial role in research experiments, code correctness is often presumed on the perceived quality of results.
Approach: They propose to promote code-quality checklists to promote coding best practices . they propose to fix bugs in conformer implementations to mitigate this risk .
Outcome: The proposed checklists aim to promote coding best practices and improve software quality within the NLP community.
Under the Morphosyntactic Lens: A Multifaceted Evaluation of Gender Bias in Speech Translation (2022.acl-long)

Copied to clipboard

Challenge: grammatical gender languages are characterized by morphosyntactic chains of gender agreement marked on a variety of lexical items and parts-of-speech (POS).
Approach: They propose to enrich the natural, gender-sensitive MuST-SHE corpus with two new linguistic annotation layers to explore gender bias.
Outcome: The proposed models shed light on gender bias and its detection at several levels of granularity.
MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languages (2024.emnlp-main)

Copied to clipboard

Challenge: Existing speech FMs fall short of full compliance with open-source principles . existing models do not have model weights, code, and training data publicly available .
Approach: They propose to use a CC-BY license to create open-source speech FMs for EU languages . they collect suitable training data by surveying automatic speech recognition datasets .
Outcome: The proposed model can be used in the 24 official languages of the European Union.
MuST-C: a Multilingual Speech Translation Corpus (N19-1)

Copied to clipboard

Challenge: Current research on spoken language translation (SLT) has to confront the scarcity of sizeable and publicly available training corpora.
Approach: They propose a multilingual speech translation corpus that will facilitate the training of end-to-end systems for SLT from English into 8 languages.
Outcome: The proposed multilingual speech translation corpus will facilitate the training of end-to-end systems for spoken language translation from English into 8 languages.
CTC-based Compression for Direct Speech Translation (2021.eacl-main)

Copied to clipboard

Challenge: Existing studies have shown that a dynamic phone-informed compression of the input audio is beneficial for speech translation (ST).
Approach: They propose a method which performs a phone-informed compression of the input audio in direct ST models by exploiting the Connectionist Temporal Classification (CTC) they demonstrate that their method brings a 1.3-1.5 BLEU improvement over a strong baseline on two language pairs (English-Italian and English-German)
Outcome: The proposed method brings a 1.3-1.5 BLEU improvement over a strong baseline on two language pairs (English-Italian and English-German) it reduces memory footprint by more than 10%, and is faster than previous approaches.
Translation in the Hands of Many: Centering Lay Users in Machine Translation Interactions (2025.emnlp-main)

Copied to clipboard

Challenge: Multilingual demands and accessibility have made MT a global tool . however, the understanding of MT consumed by such a diverse group of users remains limited.
Approach: They first trace the evolution of MT user profiles, focusing on non-experts and how their engagement with technology may shift with the rise of LLMs.
Outcome: The proposed approach will help to align MT with user needs and improve the quality of the language.
Mind the Inclusivity Gap: Multilingual Gender-Neutral Translation Evaluation with mGeNTE (2025.emnlp-main)

Copied to clipboard

Challenge: Genderneutral translation (GNT) is a linguistic strategy towards fairer communication across languages.
Approach: They propose to use a multilingual evaluation resource to evaluate inclusive translation with state-of-the-art instruction-following language models (LMs)
Outcome: The proposed model can recognize when neutrality is appropriate, but cannot consistently produce neutral translations, limiting their usability.
Is “moby dick” a Whale or a Bird? Named Entities and Terminology in Speech Translation (2021.emnlp-main)

Copied to clipboard

Challenge: Among rare words, named entities and domain-specific terms are crucial . previous studies have neglected these important words due to limited options .
Approach: They propose a benchmark to evaluate automatic translation systems for rare words . named entities and domain-specific terms are crucial for their translation .
Outcome: The proposed benchmark is based on European Parliament speeches annotated with NEs and terminology.
Attention as a Guide for Simultaneous Speech Translation (2023.acl-long)

Copied to clipboard

Challenge: EDAtt uses attention patterns to determine when to emit partial translations . results show that it yields better results compared to existing SimulST policies .
Approach: They propose an adaptive policy that exploits attention patterns between audio source and target textual translation to guide an offline-trained ST model during simultaneous inference.
Outcome: The proposed policy yields better results on en->de, compared to the current state of the art.
A Prompt Response to the Demand for Automatic Gender-Neutral Translation (2024.eacl-short)

Copied to clipboard

Challenge: Advancements in machine translation (MT) are hindered by the lack of dedicated parallel data, which are necessary to adapt MT systems to satisfy neutral constraints.
Approach: They propose to use GPT-4 to generate GNTs that avoid bias and undue binary assumptions by comparing MT with the popular GPT-3 model.
Outcome: The proposed model outperforms the existing model and provides valuable insights into the potential and challenges associated with prompting for neutrality.
How Do Hyenas Deal with Human Speech? Speech Recognition and Translation with ConfHyena (2024.lrec-main)

Copied to clipboard

Challenge: Currently, attention-based models face computational hurdles in processing long sequences due to its quadratic complexity.
Approach: They propose a conformer whose encoder self-attentions are replaced with Hyena for speech processing . they propose 'confhyena' model that reduces training time by 27% at minimal cost .
Outcome: The proposed model reduces training time by 27% at the cost of minimal quality degradation.
Machine Translation for Machines: the Sentiment Classification Use Case (D19-1)

Copied to clipboard

Challenge: Traditionally, machine translation (MT) pursues a "human-oriented" objective: generating fluent output for a downstream task.
Approach: They propose a neural machine translation approach that uses weak feedback to generate translations that are best suited for a downstream task.
Outcome: The proposed approach outperforms general-purpose models and reinforcement learning methods on German and Italian tweets.
Breeding Gender-aware Direct Speech Translation Systems (2020.coling-main)

Copied to clipboard

Challenge: In automatic speech translation, traditional cascade approaches involving separate transcription and translation steps are giving ground to more robust direct solutions.
Approach: They compare different approaches to inform direct ST models about the speaker’s gender and test their ability to handle gender translation from English into Italian and French.
Outcome: The proposed models outperform strong but gender-unaware direct ST models in the translation of English into Italian and French.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations