Papers by Emmanuelle Esperança-Rodier

5 papers
Online Versus Offline NMT Quality: An In-depth Analysis on English-German and German-English (2020.coling-main)

Copied to clipboard

Challenge: Existing studies compare offline and online neural machine translation architectures . we examine the impact of online decoding constraints on the translation quality .
Approach: They evaluate offline and online neural machine translation architectures using human evaluations on English-German and German-English language pairs.
Outcome: The proposed models are particularly sensitive to latency constraints and are well-suited for offline translation tasks.
A Multimodal French Corpus of Aligned Speech, Text, and Pictogram Sequences for Speech-to-Pictogram Machine Translation (2024.lrec-main)

Copied to clipboard

Challenge: Existing algorithms for the automatic translation of spoken language into pictogram units are lacking for language impairments.
Approach: They propose to use a French dataset that contains 230 hours of speech resources to create a rule-based pictogram grammar with a restricted vocabulary and a discussion of strategic decisions.
Outcome: The proposed model is validated through multiple post-editing phases by expert annotators and is freely available under a non-commercial licence.
DOLFIN - Document-Level Financial Test-Set for Machine Translation (2025.findings-naacl)

Copied to clipboard

Challenge: Existing document-level machine translation test-sets cover general domain but fall short on specialised domains, such as legal and financial.
Approach: They propose to use a document-level machine translation test-set to replace perfectly aligned sentences by presenting data in units of sections rather than sentences.
Outcome: The proposed dataset is built from specialised financial documents and it shows that it can discriminate between context-sensitive and context-agnostic models and shows the weaknesses when models fail to accurately translate financial texts.
Limitations of Human Identification of Automatically Generated Text (2024.lrec-main)

Copied to clipboard

Challenge: Neural text generation tools such as ChatGPT are gaining popularity . human annotations are considered gold standard labels for multiple tasks .
Approach: They propose a new corpus in French and English for recognising automatically generated texts . they propose 'incontext' setup which makes explicit the interaction between two parties .
Outcome: The proposed model generates fluent text, which requires much closer reading than the current model.
Jargon: A Suite of Language Models and Evaluation Tasks for French Specialized Domains (2024.lrec-main)

Copied to clipboard

Challenge: Pretrained language models are the de facto backbone of most state-of-the-art NLP systems.
Approach: They propose a family of domain-specific pretrained PLMs for French focusing on three important domains: transcribed speech, medicine, and law.
Outcome: The proposed models perform better on transcribed speech, medicine, and law domains than state-of-the-art models on a diverse set of tasks and datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations