Papers by Emily Prud’hommeaux

13 papers
A dataset for identifying actionable feedback in collaborative software development (P18-2)

Copied to clipboard

Challenge: a dataset of code reviews for the Google Chromium project analyzed linguistic features of code review feedback that elicited responsive actions from coworkers.
Approach: They analyze code reviews for Google Chromium and extract linguistic features that elicit responsive responses from coworkers.
Outcome: The proposed dataset shows that using NLP can be useful in code reviews . it also shows that it can be used to improve code reviews in a collaborative environment .
ASR for Documenting Acutely Under-Resourced Indigenous Languages (L18-1)

Copied to clipboard

Challenge: Automatic speech recognition (ASR) has not been widely explored as a tool for documenting endangered languages.
Approach: They propose to use automatic speech recognition (ASR) to bootstrap new data to improve the acoustic model.
Outcome: The proposed system improves the model for a polysynthetic language with few audio and text resources.
Not always about you: Prioritizing community needs when developing endangered language technology (2022.acl-long)

Copied to clipboard

Challenge: low-resource languages lack the quantity of data needed to train statistical and machine learning tools and models.
Approach: They propose to use language technology to support endangered languages' revitalization . they propose to work with indigenous speakers to develop technology for such training .
Outcome: The authors discuss the challenges that researchers and indigenous speech community members face when working together to develop language technology to support endangered languages.
Predicting pragmatic discourse features in the language of adults with autism spectrum disorder (2021.acl-srw)

Copied to clipboard

Challenge: Existing tools to quantify atypicality in discourse and pragmatics are difficult to precisely identify and quantify.
Approach: They present a corpus of transcribed natural conversations produced in an experimental setting and annotate them for three pragmatic features on a three-point scale.
Outcome: The proposed model yields higher accuracy than previous approaches for deriving these features, with F1 exceeding 0.82 for all three pragmatic features.
Evaluating the Performance of Transformer-based Language Models for Neuroatypical Language (2022.coling-1)

Copied to clipboard

Challenge: Difficulties with social aspects of language are among the hallmarks of autism spectrum disorder (ASD).
Approach: They propose a transformer-based framework for identifying linguistic features associated with social aspects of communication using a corpus of conversations between adults with and without ASD and neurotypical conversational partners.
Outcome: The proposed framework yields strong accuracy overall, but performance is significantly worse for the language of participants with ASD, suggesting they use a more diverse set of strategies for some social linguistic functions.
An (unhelpful) guide to selecting the best ASR architecture for your under-resourced language (2023.acl-short)

Copied to clipboard

Challenge: English ASR now has word error rates comparable to that of human transcriptionists, but only for the handful of the world's 7000 languages with abundant training resources.
Approach: They propose to use four of the most popular ASR toolkits to train ASR models for eleven languages with limited ASR training resources: eleven widely spoken languages of Africa, Asia, and South America, one endangered language of Central America, and three critically endangered languages of North America.
Outcome: The proposed architecture outperforms four of the most popular ASR toolkits for eleven languages with limited training resources.
That doesn’t sound right: Evaluating speech transcription quality in field linguistics corpora (2025.acl-short)

Copied to clipboard

Challenge: Automated speech recognition (ASR) is a popular tool for documenting languages, but field linguists do not have the data to train robust models.
Approach: They propose to use fieldwork data to identify speech transcriptions that may be unsuitable for training ASR models.
Outcome: The proposed measures can be used to identify transcriptions with characteristics common in field data but could be detrimental to ASR training.
Data-driven Model Generalizability in Crosslinguistic Low-resource Morphological Segmentation (2022.tacl-1)

Copied to clipboard

Challenge: morphological segmentation is a common method of evaluation for multilingual tasks . authors often examine models with one data set that is representative of all possible data .
Approach: They compare three broad classes of models with different parameterizations using morphological segmentation as the test case.
Outcome: The results show that the extent of model generalization depends on the characteristics of the data set, and does not necessarily rely heavily on the data sets size.
Are modern neural ASR architectures robust for polysynthetic languages? (2024.findings-emnlp)

Copied to clipboard

Challenge: Traditional morphological typology recognizes a range of morphology in the world's languages.
Approach: They investigate the performance of modern automatic speech recognition architectures on morphologically complex languages.
Outcome: The proposed architectures perform better on morphologically complex languages, the authors show . they show that they are less robust in managing high OOV rates for morphology complex languages .
UniMorph 4.0: Universal Morphology (2022.lrec-1)

Copied to clipboard

Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieraś, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Abbott Lane, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóǧa, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan North Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo M. Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Prud’hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya D. McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova
Challenge: The Universal Morphology project provides broad-coverage instantiated morphological inflection tables for hundreds of diverse languages.
Approach: They propose a language-independent feature schema for rich morphological annotation and a type-level resource of annotated data in diverse languages realizing that schema.
Outcome: The proposed schema has added 66 new languages, including 24 endangered languages.
Investigating data partitioning strategies for crosslinguistic low-resource ASR evaluation (2023.eacl-main)

Copied to clipboard

Challenge: Automatic speech recognition data sets include a single pre-defined test set consisting of one or more speakers whose speech never appears in the training set.
Approach: They propose to use hold-speaker(s)-out partitioning to partition data for five languages . utterance duration and intensity are more predictive factors of variability .
Outcome: The proposed method can produce results that do not reflect model performance on unseen data or speakers.
How Important is a Language Model for Low-resource ASR? (2024.findings-acl)

Copied to clipboard

Challenge: Using an n-gram language model in ASR may seem obvious, but its absence in most implementations suggests otherwise.
Approach: They examine whether using an n-gram language model in ASR can improve accuracy in low-resource languages.
Outcome: The proposed model is absent in most implementations, but it does improve accuracy in English and Mandarin.
What data should I include in my POS tagging training set? (2025.findings-emnlp)

Copied to clipboard

Challenge: POS tagging is a crucial task for descriptive linguistics and language documentation . POS tags are not available in all languages, but are used for training sets for understudied languages .
Approach: They compare POS tagging with in-context learning, active learning, and random sampling . they find that POS can deliver reasonable results for communities with limited resources .
Outcome: The proposed training set for Indigenous and endangered languages performs better than random sampling.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations