Papers by Eleftheria Briakou

14 papers
Can Synthetic Translations Improve Bitext Quality? (2022.acl-long)

Copied to clipboard

Challenge: Synthetic translations have been used for a wide range of NLP tasks, but it remains unclear how they differ from naturally occurring data.
Approach: They propose to use a semantic equivalence classifier to improve bitext quality without additional bilingual supervision to replace the originals.
Outcome: The proposed samples improve bitext quality without additional bilingual supervision and are validated intrinsically and extrinsically through bilingual induction and MT tasks.
Understanding and Detecting Hallucinations in Neural Machine Translation via Model Introspection (2023.tacl-1)

Copied to clipboard

Challenge: Neural sequence generation models produce outputs that are unrelated to the source text, and are potentially harmful, yet it remains unclear in what conditions they arise and how to mitigate their impact.
Approach: They first identify internal model symptoms of hallucinations by analyzing the relative token contributions to the generation in contrastive hallucinous vs. non-hallucinated outputs generated via source perturbations.
Outcome: The proposed detector outperforms both baseline models and strong classifiers on English-Chinese and German-English translation test beds.
Evaluating the Evaluation Metrics for Style Transfer: A Case Study in Multilingual Formality Transfer (2021.emnlp-main)

Copied to clipboard

Challenge: a lack of standardized and reliable methods for automatic evaluation hinders ST . prior work has employed as many as nine different automatic systems to rate formality alone .
Approach: They evaluate automatic metrics on the oft-researched task of formality style transfer . they outline best practices for automatic evaluation in (formality) style transfer and identify models that correlate well with human judgments.
Outcome: The proposed models correlate well with human judgments and are robust across languages.
Cross-Topic Distributional Semantic Representations Via Unsupervised Mappings (N19-1)

Copied to clipboard

Challenge: Existing distributional semantic models cannot capture the distinct meanings of polysemous words, resulting in conflated word representations of diverse contextual semantics.
Approach: They propose a distributional semantic model that learns multiple representations of a word based on different topics.
Outcome: The proposed model outperforms single-prototype models on NLP downstream tasks.
AfriMTE and AfriCOMET: Enhancing COMET to Embrace Under-resourced African Languages (2024.naacl-long)

Copied to clipboard

Challenge: Recent advances in machine translation (MT) have focused on scaling multilingual machine translation models and evaluation data to hundreds of languages, including multiple under-resourced languages.
Approach: They propose to use n-gram matching metrics to measure progress in multilingual machine translation to 13 typologically diverse African languages to create high-quality human evaluation data with simplified MQM guidelines.
Outcome: The proposed metrics have a higher correlation with human judgments than n-gram matching metrics such as BLEU and METEOR.
Olá, Bonjour, Salve! XFORMAL: A Benchmark for Multilingual Formality Style Transfer (2021.naacl-main)

Copied to clipboard

Challenge: XFORMAL benchmarks formal reformulations of informal text in Brazilian Portuguese, French, and Italian . most work on style transfer within English, while covering different languages has received disproportional interest.
Approach: They create a benchmark of multiple formal reformulations of informal text in Brazil, Brazil, and Italy.
Outcome: XFORMAL benchmarks formal reformulations of informal text in Brazilian Portuguese, French, and Italian . results show that state-of-the-art approaches perform close to simple baselines .
SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages? (2025.emnlp-main)

Copied to clipboard

Challenge: Existing metrics for machine translation quality for under-resourced African languages suffer from limited language coverage and poor performance in low-resource settings.
Approach: They propose a large-scale human-annotated machine translation evaluation dataset . they use a reference-based and reference-free evaluation model to compare MT quality .
Outcome: The proposed models outperform AfriCOMET and the strongest LLM on low-resource languages.
Searching for Needles in a Haystack: On the Role of Incidental Bilingualism in PaLM’s Translation Capability (2023.acl-long)

Copied to clipboard

Challenge: Large multilingual language models exhibit impressive zero- or few-shot machine translation capabilities, despite never having been explicitly and intentionally exposed to translation data.
Approach: They propose a mixed-method approach to measure and understand incidental bilingualism at scale using the Pathways Language Model.
Outcome: The proposed model is exposed to over 30 million translation pairs across at least 44 languages.
BitextEdit: Automatic Bitext Editing for Improved Low-Resource Machine Translation (2022.findings-naacl)

Copied to clipboard

Challenge: Existing methods to improve Neural Machine Translation (NMT) for lowresource languages are often trained on heuristically aligned or automatically mined data.
Approach: They propose to filter out imperfect translations that yield unreliable training signals for Neural Machine Translation (NMT) instead, they propose to refine mined bitexts by automatic editing .
Outcome: The proposed method improves the quality of mined bitexts for low-resource languages by up to 8 BLEU points.
Beyond Noise: Mitigating the Impact of Fine-grained Semantic Divergences on Neural Machine Translation (2021.acl-long)

Copied to clipboard

Challenge: Prior work treats all types of mismatches between source and target as noise . Consequently, it remains unclear how noisy parallel training samples impact NMT training.
Approach: They propose a divergent-aware NMT framework that uses factors to help NMT recover from the degradation caused by naturally occurring divergences.
Outcome: The proposed framework improves translation quality and model calibration on EN-FR tasks.
Explaining with Contrastive Phrasal Highlighting: A Case Study in Assisting Humans to Detect Translation Differences (2023.emnlp-main)

Copied to clipboard

Challenge: a common strategy to explain NLP predictions is to highlight salient tokens in their inputs.
Approach: They propose a technique to generate contrastive phrasal highlights that explain the predictions of a semantic divergence model via phrase alignment guided erasure.
Outcome: The proposed techniques match human rationales of cross-lingual semantic differences better than popular post-hoc saliency techniques and help people detect fine-grained meaning differences in human translations and critical machine translation errors.
What Else Do I Need to Know? The Effect of Background Information on Users’ Reliance on QA Systems (2023.emnlp-main)

Copied to clipboard

Challenge: Existing NLP systems can only access the retrieved context to determine the answer, resulting in a knowledge gap between the information that is required to answer the question and the information available to assess the model’s correctness.
Approach: They ask whether adding relevant background helps mitigate users’ over-reliance on predictions.
Outcome: The proposed approach reduces over-reliance on model predictions even in the absence of sufficient information to assess their correctness.
WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects (2025.findings-acl)

Copied to clipboard

Challenge: In order to evaluate large language models (LLMs), it is important to collect benchmark datasets in order to assess their multilingual performance.
Approach: They extend the WMT24 dataset to cover 55 languages by collecting new human-written references and post-edits for 46 new languages/dialects.
Outcome: The proposed dataset covers 55 languages and provides best-performing MT systems in all 55 languages.
Detecting Fine-Grained Cross-Lingual Semantic Divergences without Supervision by Learning to Rank (2020.emnlp-main)

Copied to clipboard

Challenge: Detecting fine-grained differences in content conveyed in different languages is expensive and hard to scale.
Approach: They propose a training strategy for multilingual BERT models by learning to rank divergent examples of varying granularity.
Outcome: The proposed model improves the prediction and annotation of fine-grained semantic divergences.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations