Papers by Benjamin Marie

7 papers
Disfluency Generation for More Robust Dialogue Systems (2023.findings-acl)

Copied to clipboard

Challenge: Disfluencies in user utterances can trigger a chain of errors impacting all the modules of a dialogue system.
Approach: They propose to augment existing dialogue datasets with disfluent utterances by paraphrasing them into disfluente ones.
Outcome: The proposed method improves dialogue state tracking and response generation by combining disfluent utterances with disfluency utteraces.
Unsupervised Extraction of Partial Translations for Neural Machine Translation (N19-1)

Copied to clipboard

Challenge: Neural machine translation systems usually require a large quantity of bilingual parallel data for training.
Approach: They propose an algorithm for extracting from monolingual data what they call partial translations . partial translation is a pair of source and target sentences that contain sequences of tokens that are translations of each other.
Outcome: The proposed algorithm extracts from monolingual data what we call partial translations . it takes only source and target monolingual datasets as input .
Synthesizing Parallel Data of User-Generated Texts with Zero-Shot Neural Machine Translation (2020.tacl-1)

Copied to clipboard

Challenge: Neural machine translation systems are usually trained on clean parallel data, but the quality of translations is poor when translating noisy texts.
Approach: They synthesize parallel data of UGT and exploit monolingual data to generate translations . they propose to use monolingual parallel data to train or adapt NMT systems .
Outcome: The proposed approach improves the translation quality of noisy texts while making them more robust.
Supervised and Unsupervised Machine Translation for Myanmar-English and Khmer-English (D19-52)

Copied to clipboard

Challenge: Using cleaned and normalized noisy monolingual data, supervised neural and statistical machine translation systems performed among the best for the four translation directions.
Approach: They present supervised and unsupervised machine translation systems for the WAT2019 Myanmar-English and Khmer-English translation tasks.
Outcome: The proposed systems performed among the best for the four translation directions.
Scientific Credibility of Machine Translation Research: A Meta-Evaluation of 769 Papers (2021.acl-long)

Copied to clipboard

Challenge: a meta-evaluation of machine translation (MT) has been conducted in 769 research papers . a recent study shows that evaluation practices have changed over the past decade .
Approach: They propose a meta-evaluation method for machine translation that uses BLEU scores to evaluate MT performance.
Outcome: The proposed meta-evaluation of machine translation shows that evaluation practices have changed over the past decade . the authors suggest that the evaluation process should be streamlined and standardized to ensure the validity of the evaluation method .
Unsupervised Joint Training of Bilingual Word Embeddings (P19-1)

Copied to clipboard

Challenge: Existing methods for unsupervised bilingual word embeddings are limited by the dissimilarity between the word embedded spaces.
Approach: They propose a method that trains unsupervised bilingual word embeddings jointly on parallel data generated through unsupervised machine translation.
Outcome: The proposed method outperforms unsupervised mapped bilingual word embeddings in cross-lingual NLP tasks.
Tagged Back-translation Revisited: Why Does It Really Work? (2020.acl-main)

Copied to clipboard

Challenge: In this paper, we show that neural machine translation systems trained on large back-translated data overfit some of the characteristics of machine-transcribed texts.
Approach: They propose to add a tag to back-translations to help distinguish back-translated data from original parallel training data.
Outcome: The proposed tag helps the system distinguish back-translated data from original parallel training data and is as effective as a tag in high-resource training.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations