Papers by Luke Gessler

7 papers
From Priest to Doctor: Domain Adaptation for Low-Resource Neural Machine Translation (2025.coling-main)

Copied to clipboard

Challenge: Existing data for low-resource languages are limited; the languages that could most benefit from domain adaptation (DA) are the ones left behind.
Approach: They propose a realistic setting in which they aim to translate between a high-resource and a low-resourced language with limited parallel data, a bilingual dictionary, and c) a monolingual target-domain corpus in the high-rsource language.
Outcome: The proposed methods are compared with a human evaluation of DALI and show that the most effective is the simplest.
Xposition: An Online Multilingual Database of Adpositional Semantics (2022.lrec-1)

Copied to clipboard

Challenge: Xposition is an online platform for documenting adpositional semantics across languages . SNACS provides a unified metalanguage for characterizing the major classes of meanings expressed with appositions .
Approach: They propose to use Xposition to document adpositional semantics across languages . Xpos houses annotation guidelines, structured lexicographic documentation, annotated corpora .
Outcome: The proposed platform houses annotation guidelines, structured lexicographic documentation, and annotated corpora.
Supervised Grapheme-to-Phoneme Conversion of Orthographic Schwas in Hindi and Punjabi (2020.acl-main)

Copied to clipboard

Challenge: Existing methods to predict schwa deletion in Hindi are based on prosodic or phonetic analysis.
Approach: They propose to use Hindi grapheme-to-phoneme (G2P) conversion to predict whether a schwa represented in the orthography is pronounced or unpronounced (deleted).
Outcome: The proposed model outperforms existing models on a newly-compiled pronunciation lexicon extracted from various online dictionaries.
PrOnto: Language Model Evaluations for 859 Languages (2024.lrec-main)

Copied to clipboard

Challenge: Evaluation datasets are scarce for most languages other than English due to high cost of annotation . authors present method for evaluating pretrained language models using evaluation datasets .
Approach: They propose a method which enables any language with a New Testament translation to receive evaluation datasets suitable for pretrained language models.
Outcome: The proposed method can be used in any language with a New Testament translation without manual annotation.
TAMS: Translation-Assisted Morphological Segmentation (2024.acl-long)

Copied to clipboard

Challenge: Canonical morphological segmentation is a key task in endangered language documentation . training data for canonical segmentation can be difficult, making it difficult to train high quality models.
Approach: They propose a model that leverages translation data to speed up canonical segmentation . they propose to use translation data as an additional signal to leverage the data .
Outcome: The proposed model outperforms baseline models in a super-low resource setting but yields mixed results on training splits with more data.
Understanding the Gap: an Analysis of Research Collaborations in NLP and Language Documentation (2025.findings-acl)

Copied to clipboard

Challenge: despite 20 years of NLP work, practical use of this work remains vanishingly scarce.
Approach: They propose to use interviews and surveys to examine the lack of NLP adoption in LD . they find that linguists and language communities have little or no use of Nlp in their work .
Outcome: a new study shows that linguists and language researchers are not using NLP in LD . the findings highlight the importance of misaligned professional incentives and LD software .
AMALGUM – A Free, Balanced, Multilayer English Web Corpus (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of 4M tokens is available online with a large number of high-quality annotation layers.
Approach: They propose to use a genre-balanced English web corpus with multiple annotation layers . they harness knowledge from multiple annotation layer to achieve a "better than NLP" benchmark .
Outcome: The proposed corpus is genre-balanced and features high-quality automatic annotation layers.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations