Papers by Ahmed El-Kishky

11 papers
Facebook AI’s WAT19 Myanmar-English Translation Task Submission (D19-52)

Copied to clipboard

Challenge: Using back-translation, we can improve generalization by using noisy channel re-ranking and ensembling.
Approach: They propose to use BPE-based transformer models to leverage monolingual data to improve generalization and use noisy channel re-ranking and ensembling to improve results.
Outcome: The proposed system improves on the baseline system trained exclusively on the provided small parallel dataset, and the human evaluation and BLEU score are higher.
CCAligned: A Massive Collection of Cross-Lingual Web-Document Pairs (2020.emnlp-main)

Copied to clipboard

Challenge: Cross-lingual document alignment aims to identify pairs of documents in two distinct languages that are of comparable content or translations of each other.
Approach: They exploit the signals embedded in URLs to label web documents at scale with an average precision of 94.5% across different language pairs.
Outcome: The proposed method can label documents at 94.5% across languages with high precision . the proposed method is useful for low-resource languages with limited resources .
Classification-based Quality Estimation: Small and Efficient Models for Real-world Applications (2021.emnlp-main)

Copied to clipboard

Challenge: Sentence-level Quality estimation (QE) is traditionally a regression task . but large multilingual contextualized language models are expensive and infeasible for real-world applications.
Approach: They evaluate several model compression techniques for QE and find they are inefficient . they argue that a full model parameterization is required to achieve SoTA results .
Outcome: The proposed models are poorly expressive in a regression task, the authors argue . they show that reframing QE as a classification problem and evaluating models would improve their performance in real-world applications.
An Exploratory Study on Multilingual Quality Estimation (2020.aacl-main)

Copied to clipboard

Challenge: Existing approaches to predict the quality of machine translation use language-specific models, but they lack labelled data for each language pair.
Approach: They propose to use scores from translation models to estimate quality of machine translations by predicting the quality of a translation at test time.
Outcome: The proposed models outperform single-language models in less balanced quality label distributions and low-resource settings.
Massively Multilingual Document Alignment with Cross-lingual Sentence-Mover’s Distance (2020.aacl-main)

Copied to clipboard

Challenge: Document alignment aims to identify pairs of documents in two distinct languages that are of comparable content or translations of each other.
Approach: They propose an unsupervised scoring function that leverages cross-lingual sentence embeddings to compute the semantic distance between documents in different languages.
Outcome: The proposed scoring function outperforms baseline methods on high-resource language pairs, 15% on mid-resourced language pairs and 22% on low-resourcing language pairs.
Adapting High-resource NMT Models to Translate Low-resource Related Languages without Parallel Data (2021.acl-long)

Copied to clipboard

Challenge: linguistic overlap between low-resource languages and high-resourced languages is a major obstacle for training high-quality machine translation systems.
Approach: They exploit linguistic overlap to facilitate translation to and from low-resource languages . they use monolingual data and parallel data in related high-resourced languages based on their method .
Outcome: The proposed method significantly improves translation into low-resource language compared to baselines on 7 languages from three different language families.
Putting words into the system’s mouth: A targeted attack on neural machine translation using monolingual data poisoning (2021.findings-acl)

Copied to clipboard

Challenge: Neural machine translation systems are known to be vulnerable to adversarial test inputs, however, they are also vulnerable to training attacks.
Approach: They propose a poisoning attack in which a malicious adversary inserts a small poisoned sample of monolingual text into a training set of a system trained using back-translation.
Outcome: The proposed attack is based on two methods that can be used to craft poisoned examples.
Simple Temporal Adaptation to Changing Label Sets: Hashtag Prediction via Dense KNN (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods to adapt to temporal change of user-generated social media data are stale without retraining.
Approach: They propose a non-parametric dense retrieval technique to adapt to temporal change . they use a Twitter dataset to study temporal distribution shift in tweet-hashtag prediction .
Outcome: The proposed method improves over the best static parametric baseline on a year-long Twitter dataset while avoiding costly re-training.
XLEnt: Mining a Large Cross-lingual Entity Dataset with Lexical-Semantic-Phonetic Word Alignment (2021.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to generate named entity lexica for lower-resource languages are under performing.
Approach: They propose a technique to automatically mine cross-lingual named-entity lexica from mined web data.
Outcome: The proposed technique outperforms baselines at extracting cross-lingual entity pairs and mines 164 million entity pairs from 120 different languages aligned with English.
Quality Estimation without Human-labeled Data (2021.eacl-main)

Copied to clipboard

Challenge: Quality estimation aims to measure the quality of translated content without access to a reference translation.
Approach: They propose a method that uses synthetic training data to train supervised quality estimation models.
Outcome: The proposed model outperforms models trained on human-annotated data for sentence and word-level prediction.
As Easy as 1, 2, 3: Behavioural Testing of NMT Systems for Numerical Translation (2021.findings-acl)

Copied to clipboard

Challenge: Mistranslated numbers can cause financial loss or medical misinformation.
Approach: They propose a method to assess the robustness of neural machine translation systems to numerical text via behavioural testing.
Outcome: The proposed method systematically assesses four fundamental capabilities of neural machine translation systems in translation numbers by virtue of a variety of test cases.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations