Papers by Emily Chang
How many words does it take to understand a low-resource language? (2025.naacl-srw)
Copied to clipboard
| Challenge: | We evaluated the documentation needed to create a sentence embedding space using widely spoken languages. |
| Approach: | They propose to use widely spoken languages as a proxy for low-resource languages to evaluate the documentation needed to create a sentence embedding space. |
| Outcome: | The proposed language model can be used to improve the performance of sentences embedded in low-resource languages. |
On Measures of Biases and Harms in NLP (2022.findings-aacl)
Copied to clipboard
Sunipa Dev, Emily Sheng, Jieyu Zhao, Aubrie Amstutz, Jiao Sun, Yu Hou, Mattie Sanseverino, Jiin Kim, Akihiro Nishi, Nanyun Peng, Kai-Wei Chang
| Challenge: | Recent studies show that natural language processing (NLP) technologies propagate societal biases about demographic groups associated with attributes such as gender, race, and nationality. |
| Approach: | They propose a framework for harms and questions to help practitioners understand biases . they propose measurable measures to detect and mitigate biased groups . |
| Outcome: | The proposed framework provides a framework for harms and questions for practitioners to answer to guide the development of bias measures. |
“Nice Try, Kiddo”: Investigating Ad Hominems in Dialogue Responses (2021.naacl-main)
Copied to clipboard
| Challenge: | Ad hominem attacks target a person's character instead of the position the person is maintaining. |
| Approach: | They propose to use salient n-gram similarity as a soft constraint to reduce the amount of ad hominems generated in Twitter conversations. |
| Outcome: | The proposed method reduces the amount of ad hominems generated in human and dialogue system responses to English Twitter posts by using salient n-gram similarity as a soft constraint. |
DialUp! Modeling the Language Continuum by Adapting Models to Dialects and Dialects to Models (2025.acl-long)
Copied to clipboard
Niyati Bafna, Emily Chang, Nathaniel Romney Robinson, David R. Mortensen, Kenton Murray, David Yarowsky, Hale Sirin
| Challenge: | Recent advances in MT quality and language coverage have shown that language varieties with low baseline performance are more likely to benefit from these approaches. |
| Approach: | They propose a training-time technique for adapting a pretrained model to dialectal data and an inference-time intervention adapting dialectal datasets to the model expertise. |
| Outcome: | The proposed model shows significant performance gains for several dialects from four language families, and modest gains for two other language families. |
Towards Controllable Biases in Language Generation (2020.findings-emnlp)
Copied to clipboard
| Challenge: | a new method to induce societal biases in natural language generation is being developed . a method to equalize the amount of biased text across demographics is effective . |
| Approach: | They propose a method to induce societal biases in natural language generation by using demographic inequalities. |
| Outcome: | The proposed method is effective at equalizing biases across demographics while generating less negatively biased text overall. |
ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks for large language models (LLMs) are restricted to high- or mid-resource languages, and evaluate performance on higher-order tasks in reasoning and generation. |
| Approach: | They propose a multilingual benchmarking tool to evaluate lexical comprehension and generation abilities of large language models. |
| Outcome: | The proposed benchmarks cover 2700+ languages and surpasses existing benchmarks in terms of language coverage. |
The Woman Worked as a Babysitter: On Biases in Language Generation (D19-1)
Copied to clipboard
| Challenge: | a systematic study of biases in natural language generation (NLG) is presented . a study of language models in NLG is conducted by examining language models. |
| Approach: | They propose a systematic study of biases in natural language generation by analyzing text generated from prompts that contain mentions of different demographic groups. |
| Outcome: | The proposed method reveals biases in natural language generation (NLG) by analyzing text generated from demographic prompts. |
Societal Biases in Language Generation: Progress and Challenges (2021.acl-long)
Copied to clipboard
| Challenge: | Language generation techniques can produce undesirable societal biases that can negatively impact marginalized populations. |
| Approach: | They propose to examine how decoding techniques contribute to biases in language generation . they also conduct experiments to quantify the effects of these techniques . |
| Outcome: | The proposed methods can reduce biases and improve user experience, the authors argue . they also show that the proposed techniques can reduce societal biase . |