Papers by David Dale

12 papers
Improving Language and Modality Transfer in Translation by Character-level Modeling (2025.acl-long)

Copied to clipboard

Challenge: Current translation systems cover only 5% of the world's languages . expanding to the long-tail of low-resource languages requires data-efficient methods that rely on cross-lingual and cross-modal knowledge transfer.
Approach: They propose a character-based approach to improve adaptability to new languages and modalities by using a teacher-student approach and parallel translation data to obtain a SONAR character-level encoder.
Outcome: The proposed model outperforms subword-based models in speech-to-text translation on the FLEURS benchmark on 33 languages and achieves state-of-the-art generalizability to unseen languages.
SpeechAlign: A Framework for Speech Translation Alignment Evaluation (2024.lrec-main)

Copied to clipboard

Challenge: Speech-to-Speech and Speech- to-Text translation are currently dynamic areas of research.
Approach: They propose a framework to evaluate source-target alignment in speech models . they introduce a speech gold alignment dataset and introduce two new metrics .
Outcome: The proposed framework evaluates source-target alignment quality within speech models.
BOUQuET : dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation (2025.emnlp-main)

Copied to clipboard

Challenge: BOUQUET is a multi-way, multicentric and multi-register/domain dataset and benchmark . the dataset is handcrafted in 8 non-English languages .
Approach: They propose to use BOUQuET to collect a multi-way, multicentric and multi-register/domain dataset and benchmark in 8 non-English languages.
Outcome: The proposed dataset is available at https://huggingface.co/datasets/facebook/bouquet.
BLASER 2.0: a metric for evaluation and quality estimation of massively multilingual speech and text translation (2024.findings-emnlp)

Copied to clipboard

Challenge: Automatic evaluation of machine translation (MT) is difficult because of the number of possible ways to express a thought in a language.
Approach: They propose to use BLASER 2.0 to evaluate machine translation quality . they propose to apply the reference-based model to a sentence-based version .
Outcome: The proposed model is applicable to detecting translation hallucinations and filtering training datasets to obtain more reliable translation models.
Detecting and Mitigating Hallucinations in Machine Translation: Model Internal Workings Alone Do Well, Sentence Similarity Even Better (2023.acl-long)

Copied to clipboard

Challenge: a recent study shows that without artificially encouraging models to hallucinate, existing methods fall short . hallucinations are cases when the model generates output that is partially or fully unrelated to the source sentence.
Approach: They propose a method that evaluates the percentage of the source contribution to a generated translation.
Outcome: The proposed method improves detection accuracy for the most severe hallucinations by a factor of 2.
Text Detoxification using Large Pre-trained Neural Models (2021.emnlp-main)

Copied to clipboard

Challenge: Existing studies on text detoxification cast this task as style transfer . text detox requires better preservation of the original meaning, authors argue .
Approach: They propose two unsupervised methods for eliminating toxicity in text . they use a paraphraser guided by style-trained language models to keep the text content .
Outcome: The proposed methods yield new SOTA results.
Less Mature is More Adaptable for Sentence-level Language Modeling (2025.acl-long)

Copied to clipboard

Challenge: Existing studies fine-tune encoders or contrastive learning approaches to learn sentence representations.
Approach: They propose to use sentence-level models to study how sentence representations influence downstream task performance.
Outcome: The proposed models outperform token-level models in terms of time and data efficiency.
LCFO: Long Context and Long Form Output Dataset and Benchmarking (2025.findings-acl)

Copied to clipboard

Challenge: Using long text outputs to evaluate progress in summarization and summary expansion tasks is challenging.
Approach: They propose a framework for assessing gradual summarization and summary expansion capabilities across diverse domains.
Outcome: The proposed framework provides alignments between specific QA pairs and corresponding summaries in 7 domains.
MuTox: Universal MUltilingual Audio-based TOXicity Dataset and Zero-shot Detector (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies on text-based toxicity detection for other languages are limited, especially for languages other than English.
Approach: They propose a multilingual audio-based toxicity classifier which covers 14 different linguistic families and a dataset of 20,000 audio utterances for English and Spanish.
Outcome: The new classifier improves F1-Score by an average of 100% when compared to existing wordlist-based classifiers.
ParaDetox: Detoxification with Parallel Data (2022.acl-long)

Copied to clipboard

Challenge: a pipeline for the collection of parallel data for the detoxification task is available.
Approach: They propose a pipeline for the collection of parallel data for the detoxification task . they collect non-toxic paraphrases for over 10,000 English toxic sentences .
Outcome: The proposed pipeline outperforms state-of-the-art models on both automatic and manual evaluations.
A large-scale computational study of content preservation measures for text style transfer and paraphrase generation (2022.acl-srw)

Copied to clipboard

Challenge: Text style transfer and paraphrases generation are growing areas of NLP . many researchers still use BLEU-like measures to evaluate content preservation .
Approach: They compare 57 different measures based on different principles on 19 annotated datasets . they find that measures relying on cross-encoder models outperform alternative approaches .
Outcome: The proposed methods outperform traditional methods on 19 datasets.
HalOmi: A Manually Annotated Benchmark for Multilingual Hallucination and Omission Detection in Machine Translation (2023.emnlp-main)

Copied to clipboard

Challenge: Previously available quality assessments do not distinguish between hallucinations and omissions.
Approach: They propose to annotate hallucinations and omissions in machine translation using a single language pair.
Outcome: The proposed dataset covers 18 translation directions with varying resource levels and scripts.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations