Papers by Ekaterina Vylomova

10 papers
Simpson’s Paradox and the Accuracy-Fluency Tradeoff in Translation (2024.acl-short)

Copied to clipboard

Challenge: Existing studies suggest that accuracy and fluency should trade off against each other, and that capturing every detail of the source is difficult for human raters to distinguish.
Approach: They propose to evaluate the relationship between accuracy and fluency at the segment level and to use probabilities to estimate probabilities.
Outcome: The proposed model relies on human judgments of accuracy and fluency collected in prior work on translation quality estimation.
Predicting Human Translation Difficulty Using Automatic Word Alignment (2023.findings-acl)

Copied to clipboard

Challenge: Translation difficulty is a problem when translators are required to resolve translation ambiguity from multiple possible translations.
Approach: They use word alignments computed over large scale bilingual corpora to develop predictors of lexical translation difficulty.
Outcome: The proposed method improves on a previous embedding-based approach and can contribute to a deeper understanding of cross-lingual differences and of causes of translation difficulty.
A Multidimensional Framework for Evaluating Lexical Semantic Change with Social Science Applications (2024.acl-long)

Copied to clipboard

Challenge: Historical linguists have identified multiple forms of lexical semantic change.
Approach: They propose a framework for integrating and evaluating lexical semantic changes in historical linguists and a unified computational methodology for evaluating them concurrently.
Outcome: The proposed framework enables lexical semantic change to be mapped economically and systematically and has applications in computational social science.
Where the Cat Sat: A Multilingual Framework for Spatial Language Understanding (2026.acl-long)

Copied to clipboard

Challenge: Existing work exhibits biases toward English and prepositional marking . Existing models are limited in understanding spatial relations across typologically diverse languages .
Approach: They propose a multilingual framework and benchmark for spatial language understanding . they decompose spatial relations into surface elements and semantic components . their results suggest surface parsing does not entail spatial understanding - they argue .
Outcome: The proposed framework and benchmark decomposes spatial relations into surface elements and semantic components.
UniMorph 2.0: Universal Morphology (L18-1)

Copied to clipboard

Challenge: The Universal Morphology project is a collaborative effort to improve how NLP handles complex morphology across the world's languages.
Approach: They propose to use a universal tagset to annotate morphological data using a schema that includes a lemma and a bundle of morphology features.
Outcome: The project releases annotated morphological data using a universal tagset, the UniMorph schema.
Tulun: Transparent and Adaptable Low-resource Machine Translation (2025.acl-demo)

Copied to clipboard

Challenge: a low-resource language that is the lingua franca in Timor-Leste lacks available corpora in the health domain.
Approach: They propose a solution that combines neural MT with large language model-based post-editing guided by existing glossaries and translation memories.
Outcome: The proposed system outperforms both standalone MT and LLM approaches across six low-resource languages on the FLORES dataset.
LSC-Eval: A General Framework to Evaluate Methods for Assessing Dimensions of Lexical Semantic Change Using LLM-Generated Synthetic Data (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for measuring Lexical Semantic Change are lacking historical benchmarks.
Approach: They propose a three-stage general-purpose evaluation framework that simulates theory-driven LSC using In-Context Learning and a lexical database.
Outcome: The proposed framework evaluates the sensitivity of computational methods to synthetic change and their suitability for detecting change in specific dimensions and domains.
UniMorph 4.0: Universal Morphology (2022.lrec-1)

Copied to clipboard

Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieraś, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Abbott Lane, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóǧa, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan North Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo M. Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Prud’hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya D. McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova
Challenge: The Universal Morphology project provides broad-coverage instantiated morphological inflection tables for hundreds of diverse languages.
Approach: They propose a language-independent feature schema for rich morphological annotation and a type-level resource of annotated data in diverse languages realizing that schema.
Outcome: The proposed schema has added 66 new languages, including 24 endangered languages.
Contextualization of Morphological Inflection (N19-1)

Copied to clipboard

Challenge: In this paper, we isolate the task of predicting a fully inflected sentence from its partially lemmatized version.
Approach: They propose a task that requires morphological features to be inferred from sentential context . they propose morphology-based models that explicitly reconstruct morphologic features before predicting inflected forms .
Outcome: The proposed model is able to predict inflected sentences without relying on morphological annotations.
UniMorph 3.0: Universal Morphology (2020.lrec-1)

Copied to clipboard

Challenge: Explicit modeling of morphology has demonstrable benefits for language modeling, speech recognition, word embedding and keyword search.
Approach: They propose a language-independent feature schema for rich morphological annotation and a type-level resource for annotated data in diverse languages.
Outcome: The proposed schema has been improved to make it more complete and correct, and adds 66 new languages and parts of speech for 12 languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations