Papers by Heather Lent
Text Embedding Inversion Security for Multilingual Language Models (2024.acl-long)
Copied to clipboard
| Challenge: | storing sensitive information as embeddings is susceptible to security breaches, as text can be reconstructed from embeddables . study explores multilingual inversion attacks using a masking defense . |
| Approach: | They propose a simple masking defense that can be used to decode embedded text . they define the problem of black-box multilingual and crosslingual inversion attacks . |
| Outcome: | The proposed defense is effective for both monolingual and multilingual models. |
Compositional Generalization in Multilingual Semantic Parsing over Wikidata (2022.tacl-1)
Copied to clipboard
| Challenge: | Semantic parsers are mostly designed for and evaluated on English resources, such as CFQ. |
| Approach: | They propose a method for creating a multilingual, parallel question-query dataset . they analyze compositional generalization of parsers in Hebrew, Kannada, Chinese, and English . |
| Outcome: | The proposed method analyzes compositional generalization of parsers in Hebrew, Kannada, Chinese, and English. |
Eidos, INDRA, & Delphi: From Free Text to Executable Causal Models (N19-4)
Copied to clipboard
Rebecca Sharp, Adarsh Pyarelal, Benjamin Gyori, Keith Alcock, Egoitz Laparra, Marco A. Valenzuela-Escárcega, Ajay Nagesh, Vikas Yadav, John Bachman, Zheng Tang, Heather Lent, Fan Luo, Mithun Paul, Steven Bethard, Kobus Barnard, Clayton Morrison, Mihai Surdeanu
| Challenge: | a paper proposes a method for building probabilistic models of complex phenomena such as food insecurity . currently, these models are hand-built for each new situation and require months to construct . |
| Approach: | They propose an approach that builds executable probabilistic models from raw, free text. |
| Outcome: | The proposed approach builds executable probabilistic models from raw, free text. |
Rewarding Coreference Resolvers for Being Consistent with World Knowledge (D19-1)
Copied to clipboard
Rahul Aralikatte, Heather Lent, Ana Valeria Gonzalez, Daniel Hershcovich, Chen Qiu, Anders Sandholm, Michael Ringaard, Anders Søgaard
| Challenge: | Unresolved coreference is a bottleneck for relation extraction systems . a state-of-the-art system may be able to infer the relation using distributional information about the phrase the Sunshine State, but is likely to have limited evidence for the decision that it is coreferential with Florida rather than with Skynyrd. |
| Approach: | They propose to forward coreference input to relation extraction system and reward them for producing triples that are found in knowledge bases. |
| Outcome: | The proposed approach improves over the state-of-the-art by forwarding their input to a relation extraction system and rewarding resolvers for producing triples that are found in knowledge bases. |
What a Creole Wants, What a Creole Needs (2022.lrec-1)
Copied to clipboard
| Challenge: | Recent efforts to improve the quality of high-resource languages focus on translating existing datasets into other languages, but this approach ignores that different language communities have different needs. |
| Approach: | They examine how things needed from language technology can change dramatically from one language to another. |
| Outcome: | The proposed method ignores that different language communities have different needs. |
Challenges and Strategies in Cross-Cultural NLP (2022.acl-long)
Copied to clipboard
Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, Anders Søgaard
| Challenge: | Various efforts have been made to accommodate linguistic diversity and serve speakers of many different languages. |
| Approach: | They propose a framework to examine cultural differences in NLP to better serve users . they argue that cultural knowledge, preferences and values can affect NLP practices . |
| Outcome: | The proposed framework examines how cultural knowledge, preferences and values can affect NLP practices. |
How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP (2026.acl-long)
Copied to clipboard
Kushal Tatariya, Artur Kulmizev, Wessel Poelman, Esther Ploeger, Marcel Bollmann, Johannes Bjerva, Jiaming Luo, Heather Lent, Miryam de Lhoneux
| Challenge: | Wikipedia’s perceived high quality and broad language coverage have established it as a fundamental resource in NLP. |
| Approach: | They propose a data filtering procedure which removes a large percentage of Wikipedia's data and a 4-level quality ranking of the site. |
| Outcome: | The results show that the proposed filtering procedure outperforms the raw Wikipedia models in three language modelling scenarios. |
Limited-Resource Adapters Are Regularizers, Not Linguists (2025.acl-short)
Copied to clipboard
Marcell Fekete, Nathaniel Romney Robinson, Ernests Lavrinovics, Djeride Jean-Baptiste, Raj Dabre, Johannes Bjerva, Heather Lent
| Challenge: | Existing studies show that cross-lingual transfer from high-resource languages is promising for low-resourced machine translation. |
| Approach: | They propose to use adapter souping and cross-attention fine-tuning to leverage language transfer for Creoles, an under-served group of low-resource languages. |
| Outcome: | The proposed method improves performance over baselines but not meaningfully with adapters. |