Papers by Luke Gessler
From Priest to Doctor: Domain Adaptation for Low-Resource Neural Machine Translation (2025.coling-main)
Copied to clipboard
| Challenge: | Existing data for low-resource languages are limited; the languages that could most benefit from domain adaptation (DA) are the ones left behind. |
| Approach: | They propose a realistic setting in which they aim to translate between a high-resource and a low-resourced language with limited parallel data, a bilingual dictionary, and c) a monolingual target-domain corpus in the high-rsource language. |
| Outcome: | The proposed methods are compared with a human evaluation of DALI and show that the most effective is the simplest. |
Xposition: An Online Multilingual Database of Adpositional Semantics (2022.lrec-1)
Copied to clipboard
| Challenge: | Xposition is an online platform for documenting adpositional semantics across languages . SNACS provides a unified metalanguage for characterizing the major classes of meanings expressed with appositions . |
| Approach: | They propose to use Xposition to document adpositional semantics across languages . Xpos houses annotation guidelines, structured lexicographic documentation, annotated corpora . |
| Outcome: | The proposed platform houses annotation guidelines, structured lexicographic documentation, and annotated corpora. |
Supervised Grapheme-to-Phoneme Conversion of Orthographic Schwas in Hindi and Punjabi (2020.acl-main)
Copied to clipboard
| Challenge: | Existing methods to predict schwa deletion in Hindi are based on prosodic or phonetic analysis. |
| Approach: | They propose to use Hindi grapheme-to-phoneme (G2P) conversion to predict whether a schwa represented in the orthography is pronounced or unpronounced (deleted). |
| Outcome: | The proposed model outperforms existing models on a newly-compiled pronunciation lexicon extracted from various online dictionaries. |
PrOnto: Language Model Evaluations for 859 Languages (2024.lrec-main)
Copied to clipboard
| Challenge: | Evaluation datasets are scarce for most languages other than English due to high cost of annotation . authors present method for evaluating pretrained language models using evaluation datasets . |
| Approach: | They propose a method which enables any language with a New Testament translation to receive evaluation datasets suitable for pretrained language models. |
| Outcome: | The proposed method can be used in any language with a New Testament translation without manual annotation. |
TAMS: Translation-Assisted Morphological Segmentation (2024.acl-long)
Copied to clipboard
| Challenge: | Canonical morphological segmentation is a key task in endangered language documentation . training data for canonical segmentation can be difficult, making it difficult to train high quality models. |
| Approach: | They propose a model that leverages translation data to speed up canonical segmentation . they propose to use translation data as an additional signal to leverage the data . |
| Outcome: | The proposed model outperforms baseline models in a super-low resource setting but yields mixed results on training splits with more data. |
Understanding the Gap: an Analysis of Research Collaborations in NLP and Language Documentation (2025.findings-acl)
Copied to clipboard
| Challenge: | despite 20 years of NLP work, practical use of this work remains vanishingly scarce. |
| Approach: | They propose to use interviews and surveys to examine the lack of NLP adoption in LD . they find that linguists and language communities have little or no use of Nlp in their work . |
| Outcome: | a new study shows that linguists and language researchers are not using NLP in LD . the findings highlight the importance of misaligned professional incentives and LD software . |
AMALGUM – A Free, Balanced, Multilayer English Web Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of 4M tokens is available online with a large number of high-quality annotation layers. |
| Approach: | They propose to use a genre-balanced English web corpus with multiple annotation layers . they harness knowledge from multiple annotation layer to achieve a "better than NLP" benchmark . |
| Outcome: | The proposed corpus is genre-balanced and features high-quality automatic annotation layers. |