Papers by Hila Gonen
That was the last straw, we need more: Are Translation Systems Sensitive to Disambiguating Context? (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing models for translation of ambiguous text use context to disambiguate meaning . current models for MTs consistently translate English idioms literally, whereas LMs are context-aware . |
| Approach: | They use a dataset of 512 pairs of English sentences to study semantic ambiguities . they use literal and figurative idioms to disambiguate intended meaning . |
| Outcome: | The results show that current models translate English idioms literally, even when the context suggests a figurative interpretation. |
Analyzing the Mono- and Cross-Lingual Pretraining Dynamics of Multilingual Language Models (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies on multilingual models have focused on their cross-lingual transfer behavior . a recent study examined multilingual model learning from the multilingual pretraining signal . |
| Approach: | They analyze checkpoints during multilingual pretraining to identify when models acquire in-language and cross-lingual abilities. |
| Outcome: | The proposed model achieves high in-language performance early on, with lower-level linguistic skills acquired before more complex ones. |
LEXPLAIN: Improving Model Explanations via Lexicon Supervision (2023.starsem-1)
Copied to clipboard
| Challenge: | Existing methods that extract features from input text to explain a classifier's prediction are limiting to models that are faithful to their predictions. |
| Approach: | They propose a framework for guiding model explanations by supervising them explicitly using task-related lexicons to direct supervise model explanation. |
| Outcome: | The proposed method improves model explanations without sacrificing performance on sentiment analysis and toxicity detection tasks while demoting spurious correlations with African American English dialects. |
McPhraSy: Multi-Context Phrase Similarity and Clustering (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for estimating phrase similarity use the phrase context only during training, instead relying on the phrase itself. |
| Approach: | They propose a novel algorithm that leverages multiple contexts during inference to estimate the similarity of phrases based on multiple context. |
| Outcome: | The proposed method outperforms existing models on two phrase similarity datasets by 13.3% and a new task that relies on phrase similarities in the product reviews domain. |
MULTIGUARD: An Efficient Approach for AI Safety Moderation Across Languages and Modalities (2025.emnlp-main)
Copied to clipboard
Sahil Verma, Keegan Hines, Jeff Bilmes, Charlotte Siska, Luke Zettlemoyer, Hila Gonen, Chandan Singh
| Challenge: | Existing approaches to detect harmful queries to large language models are fallible and vulnerable to attacks that exploit mismatched generalization of model capabilities. |
| Approach: | They propose an approach to detect harmful queries to large language models (LLMs) OMNIGUARD identifies internal representations of an LLM/MLLM that are aligned across languages or modalities and builds a language-agnostic or modality-adic classifier for detecting harmful prompts. |
| Outcome: | OMNIGUARD improves harmful prompt classification accuracy by 11.57% over the strongest baseline in a multilingual setting, by 20.44% for image-based prompts, and sets a new SOTA for audio-based ones. |
Universal NER: A Gold-Standard Multilingual Named Entity Recognition Benchmark (2024.naacl-long)
Copied to clipboard
Stephen Mayhew, Terra Blevins, Shuheng Liu, Marek Suppa, Hila Gonen, Joseph Marvin Imperial, Börje Karlsson, Peiqin Lin, Nikola Ljubešić, Lester James Miranda, Barbara Plank, Arij Riabi, Yuval Pinter
| Challenge: | In named entity recognition, the majority of annotation efforts are centered on English, and cross-lingual transfer performance remains brittle. |
| Approach: | They propose to develop gold-standard named entity recognition benchmarks in many languages using a cross-lingual consistent schema. |
| Outcome: | The proposed benchmarks will be released to the public in 2022 . they will provide baselines on in-language and cross-lingual learning settings. |
Prompting Language Models for Linguistic Structure (2023.acl-long)
Copied to clipboard
| Challenge: | Existing prompting methods can test this hypothesis on autoregressive PLMs. |
| Approach: | They propose a structured prompting approach for linguistic structured prediction tasks that performs zero- and few-shot sequence tagging with autoregressive PLMs. |
| Outcome: | The proposed approach shows that the model can perform few-shot sequence tagging on part-of-speech taging, named entity recognition, and sentence chunking tasks. |
Automatically Identifying Gender Issues in Machine Translation using Perturbations (2020.findings-emnlp)
Copied to clipboard
| Challenge: | a novel approach to machine translation has addressed outstanding challenges, including the modeling and treatment of gendered language. |
| Approach: | They propose a method to mine examples from real world data to explore challenges for deployed systems. |
| Outcome: | The proposed method exposes where model representations are gendered and the unintended consequences of genderes in downstream applications. |
Lipstick on a Pig: Debiasing Methods Cover up Systematic Gender Biases in Word Embeddings But do not Remove Them (N19-1)
Copied to clipboard
| Challenge: | Existing methods to remove gender bias from word embeddings are insufficient, we argue . existing methods for gender-neutral modeling are ineffective, we conclude . |
| Approach: | They propose methods to reduce gender bias in word embeddings by debiasing them using text corpora. |
| Outcome: | The proposed methods show that they can reduce gender bias in word embeddings . the proposed methods are insufficient and should not be trusted, the authors argue . |
BUFFET: Benchmarking Large Language Models for Few-shot Cross-lingual Transfer (2024.naacl-long)
Copied to clipboard
Akari Asai, Sneha Kudugunta, Xinyan Yu, Terra Blevins, Hila Gonen, Machel Reid, Yulia Tsvetkov, Sebastian Ruder, Hannaneh Hajishirzi
| Challenge: | Recent advances in few-shot generalization in natural language processing focus on English. |
| Approach: | They propose a benchmark that unifies 15 diverse tasks across 54 languages in a sequence-to-sequence format and provides a fixed set of few-shot examples and instructions. |
| Outcome: | The proposed framework unifies 15 diverse tasks across 54 languages in a sequence-to-sequence format and provides a fixed set of few-shot examples and instructions. |
XLM-V: Overcoming the Vocabulary Bottleneck in Multilingual Masked Language Models (2023.emnlp-main)
Copied to clipboard
Davis Liang, Hila Gonen, Yuning Mao, Rui Hou, Naman Goyal, Marjan Ghazvininejad, Luke Zettlemoyer, Madian Khabsa
| Challenge: | Large multilingual models rely on a single vocabulary shared across 100+ languages . this vocabulary bottleneck limits the representational capabilities of multilingual model XLM-R . |
| Approach: | They propose a new approach for scaling to large multilingual vocabularies by de-emphasizing token sharing between languages with little lexical overlap and assigning vocabulary capacity to achieve sufficient coverage for each individual language. |
| Outcome: | The proposed model outperforms XLM-R on all language tasks and is particularly effective on low-resource tasks. |
Language Modeling for Code-Switching: Evaluation, Integration of Monolingual Data, and Discriminative Training (D19-1)
Copied to clipboard
| Challenge: | Code-switching (CS) is a linguistic phenomenon defined as "the alternation of two languages within a single discourse, sentence or constituent." |
| Approach: | They propose an ASR-motivated evaluation setup which is decoupled from an ASL system and the choice of vocabulary . they propose a discriminative training approach which works better than generative language modeling . |
| Outcome: | The proposed evaluation setup is better than generative language modeling, the authors show . the proposed setup is decoupled from an ASR system and the choice of vocabulary . |
Does Liking Yellow Imply Driving a School Bus? Semantic Leakage in Language Models (2025.naacl-long)
Copied to clipboard
| Challenge: | Despite their wide adoption, the biases and unintended behaviors of language models remain poorly understood. |
| Approach: | They propose an evaluation setting to detect semantic leakage by humans and automatically . they also curate a diverse test suite for diagnosing this behavior in 13 flagship models . |
| Outcome: | The proposed evaluation setting detects semantic leakage by humans and automatically, and measures it in 13 flagship models. |
It’s All in the Name: Mitigating Gender Bias with Name-Based Counterfactual Data Substitution (D19-1)
Copied to clipboard
| Challenge: | Existing attempts to mitigate gender bias rely on operationalisation of gender bias as a projection over a linear subspace. |
| Approach: | They propose to operationalise gender bias as a linear subspace and augmented a corpus to remove bias by swapping all inherently-gendered words in the copy. |
| Outcome: | The proposed approach outperforms projection-based methods at the task of drawing non-biased gender analogies by an average of 19% across both corpora. |
Identifying Helpful Sentences in Product Reviews (2021.naacl-main)
Copied to clipboard
| Challenge: | a key advantage of online shopping is the ability to read what other customers are saying about products of interest. |
| Approach: | They propose a task to extract a representative helpful sentence from reviews . they collect a dataset in english and use crowd-sourcing to test their model . |
| Outcome: | The proposed model outperforms baselines in a crowd-sourced model of representative helpful sentences from product reviews. |
Null It Out: Guarding Protected Attributes by Iterative Nullspace Projection (2020.acl-main)
Copied to clipboard
| Challenge: | Word embeddings, pre-trained language models, and deep learning methods are becoming effective for text classification. |
| Approach: | They propose a method for removing information from neural representations using null-space projection. |
| Outcome: | The proposed method mitigates bias in word embeddings and increases fairness in multi-class classification. |
Breaking the Curse of Multilinguality with Cross-lingual Expert Language Models (2024.emnlp-main)
Copied to clipboard
Terra Blevins, Tomasz Limisiewicz, Suchin Gururangan, Margaret Li, Hila Gonen, Noah Smith, Luke Zettlemoyer
| Challenge: | Multilingual language models often underperform monolingual ones due to inter-language competition for model parameters. |
| Approach: | They propose Cross-lingual Expert Language Models (X-ELM) which mitigates inter-language competition by independently training language models on subsets of the multilingual corpus. |
| Outcome: | The proposed model outperforms jointly trained multilingual models across all 16 considered languages and transfer the gains to downstream tasks. |
Do All Languages Cost the Same? Tokenization in the Era of Commercial Language Models (2023.emnlp-main)
Copied to clipboard
Orevaoghene Ahia, Sachin Kumar, Hila Gonen, Jungo Kasai, David Mortensen, Noah Smith, Yulia Tsvetkov
| Challenge: | Language models have evolved from being research prototypes to commercialized products offered as web APIs. |
| Approach: | They conduct a systematic analysis of the cost and utility of OpenAI’s language model API on multilingual benchmarks in 22 typologically diverse languages. |
| Outcome: | The proposed language model API performs poorly on multiple languages and speakers of a large number of languages are overcharged while obtaining poorer results. |
Pick a Fight or Bite your Tongue: Investigation of Gender Differences in Idiomatic Language Usage (2020.coling-main)
Copied to clipboard
| Challenge: | Existing studies on gender-linked language have established foundations regarding cross-gender differences in lexical, emotional, and topical preferences, along with their sociological underpinnings. |
| Approach: | They compile a corpus of spontaneous linguistic productions annotated with speakers’ gender and perform an empirical study of gender differences in the usage of figurative language between male and female authors. |
| Outcome: | The results show that gender-specific idiomatic choices reflect gender-based lexical and semantic preferences in general language, men's and women's idioms express higher emotion than their literal language, and contextual analysis of idiomatic expressions reveals considerable differences, reflecting subtle divergences in usage environments, shaped by cross-gender communication styles and semantic biases. |
Voices Unheard: NLP Resources and Models for Yorùbá Regional Dialects (2024.emnlp-main)
Copied to clipboard
Orevaoghene Ahia, Anuoluwapo Aremu, Diana Abagyan, Hila Gonen, David Adelani, Daud Abolade, Noah Smith, Yulia Tsvetkov
| Challenge: | Recent efforts to develop NLP tools for low-resource languages focus on their standard dialects. |
| Approach: | They propose a high-quality parallel text and speech corpus for Yoruba . they use native speakers to collect data from four regional yoruba dialects . |
| Outcome: | The proposed dataset shows that dialect-adaptive finetuning can narrow performance disparities . the dataset will be released publicly under an open license . |
Simple, Interpretable and Stable Method for Detecting Words with Usage Change across Corpora (2020.acl-main)
Copied to clipboard
| Challenge: | comparing two corpus texts and searching for words that differ in their usage between them is a common problem in digital humanities and computational social science. |
| Approach: | They propose an alternative approach that does not use vector space alignment, and instead considers the neighbors of each word. |
| Outcome: | The proposed method is interpretable and stable in 9 different setups and is highly reliable. |
Toward Human Readable Prompt Tuning: Kubrick’s The Shining is a good movie, and a good prompt too? (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models can perform downstream tasks in a zero-shot fashion, given natural language prompts that specify the desired behavior. |
| Approach: | They propose a human readable prompt tuning method that incorporates a fluency constraint to find a distribution of effective and fluent prompts. |
| Outcome: | The proposed method outperforms baselines by 7.0% across three tasks. |
Demystifying Prompts in Language Models via Perplexity Estimation (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Language models can be prompted to perform a wide variety of tasks with zero- and few-shot learning. |
| Approach: | They propose a method to automatically extend a small seed set of manually written prompts by paraphrasing with GPT3 and backtranslation. |
| Outcome: | The proposed method extends a small seed set of manually written prompts by paraphrasing with GPT3 and backtranslation. |
Dementia Through Different Eyes: Explainable Modeling of Human and LLM Perceptions for Early Awareness (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a new study examines how dementia is perceived by non-experts . human perception of dementia is inconsistent and relies on a narrow set of cues compared to LLMs based on broader clinical patterns . |
| Approach: | They propose a method that uses LLMs to extract high-level, expert-guided features . human perception of dementia is inconsistent and relies on a narrow set of cues, they say . |
| Outcome: | The proposed method analyzes picture descriptions to assess whether they were produced by non-experts or by nonexperts. |