Papers by Marzena Karpinska
CaLMQA: Exploring culturally specific long-form question answering across 23 languages (2025.acl-long)
Copied to clipboard
| Challenge: | Despite rising global usage of large language models, their ability to generate *long-form* answers to *culturally specific* questions remains unexplored in many languages. |
| Approach: | They perform the first study of textual multilingual long-form QA by creating a dataset of culturally specific questions across 23 different languages. |
| Outcome: | The results show that the best models make critical surface-level errors for many languages and their understanding of diverse cultures. |
Program Chairs’ Report on Peer Review at ACL 2023 (2023.acl-long)
Copied to clipboard
| Challenge: | ACL'23 makes its peer review report public and an official part of the conference proceedings. |
| Approach: | They present an analysis of the factors affecting peer review and identify the most problematic issues that the authors complained about. |
| Outcome: | The authors identified the most problematic issues and provided suggestions for the future chairs. |
AI use in American newspapers is widespread, uneven, and rarely disclosed (2026.acl-long)
Copied to clipboard
Jenna Russell, Marzena Karpinska, Destiny Akinode, James Zhou, Katherine Thai, Bradley Emi, Max Spero, Mohit Iyyer
| Challenge: | a large-scale dataset of 186K articles from 1.5K newspapers published in the summer of 2025 is audited. |
| Approach: | They audit 186K articles from 1.5K newspapers published in summer of 2025 . they use Pangram, a state-of-the-art AI detector, to detect whether articles are partially or fully AI-generated . |
| Outcome: | The findings highlight the need for greater transparency and updated editorial standards regarding the use of AI in journalism to maintain public trust. |
Exploring Document-Level Literary Machine Translation with Parallel Paragraphs from World Literature (2022.emnlp-main)
Copied to clipboard
Katherine Thai, Marzena Karpinska, Kalpesh Krishna, Bill Ray, Moira Inghilleri, John Wieting, Mohit Iyyer
| Challenge: | Literary translation is a culturally significant task, but it is bottlenecked by the small number of qualified literary translators . a dataset of non-English language novels is used to study literary MT . |
| Approach: | They use a dataset of non-English language novels aligned to human and automatic English translations to study literary MT. |
| Outcome: | The proposed model prefers human translations over machine translations at a rate of 84% . state-of-the-art MT metrics do not correlate with preferences, the study finds . |
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text (2025.acl-long)
Copied to clipboard
| Challenge: | Qualitative analysis of experts’ free-form explanations shows that while they rely heavily on specific lexical clues (‘AI vocabulary’), they also pick up on more complex phenomena within the text (e.g., formality, originality, clarity). |
| Approach: | They hire annotators to read 300 non-fiction English articles, label them as either human-written or AI-generated, and provide paragraph-length explanations for their decisions. |
| Outcome: | The annotators who frequently use LLMs for writing tasks outperform commercial and open-source detectors even without evasion tactics like paraphrasing and humanization. |
ezCoref: Towards Unifying Annotation Guidelines for Coreference Resolution (2023.findings-eacl)
Copied to clipboard
Ankita Gupta, Marzena Karpinska, Wenlong Zhao, Kalpesh Krishna, Jack Merullo, Luke Yeh, Mohit Iyyer, Brendan O’Connor
| Challenge: | Existing datasets vary in definition of coreferences and are curated for linguistic experts. |
| Approach: | They propose to use ezCoref to create a crowdsourcing-friendly coreference annotation methodology that teaches annotators only cases that are treated similarly across existing datasets. |
| Outcome: | The proposed method reannotates 240 passages from seven existing english coreference datasets while teaching annotators only cases that are treated similarly across them. |
Aurora-M: Open Source Continual Pre-training for Multilingual Language and Code (2025.coling-industry)
Copied to clipboard
Taishi Nakamura, Mayank Mishra, Simone Tedeschi, Yekun Chai, Jason T. Stillerman, Felix Friedrich, Prateek Yadav, Tanmay Laud, Vu Minh Chien, Terry Yue Zhuo, Diganta Misra, Ben Bogin, Xuan-Son Vu, Marzena Karpinska, Arnav Varma Dantuluri, Wojciech Kusa, Tommaso Furlanello, Rio Yokota, Niklas Muennighoff, Suhas Pai, Tosin Adewumi, Veronika Laippala, Xiaozhe Yao, Adalberto Barbosa Junior, Aleksandr Drozd, Jordan Clive, Kshitij Gupta, Liangyu Chen, Qi Sun, Ken Tsui, Nour Moustafa-Fahmy, Nicolo Monti, Tai Dang, Ziyang Luo, Tien-Tung Bui, Roberto Navigli, Virendra Mehta, Matthew Blumberg, Victor May, Hiep Nguyen, Sampo Pyysalo
| Challenge: | Pretrained language models are integral part of AI applications, but their high computational cost limits accessibility. |
| Approach: | They evaluate Aurora-M, a 15B parameter multilingual open-source model trained on English, Finnish, Hindi, Japanese, Vietnamese, and code. |
| Outcome: | The proposed model outperforms existing models on English, Finnish, Hindi, Japanese, Vietnamese, and code. |
OWL: Probing Cross-Lingual Recall of Memorized Texts via World Literature (2025.emnlp-main)
Copied to clipboard
Alisha Srivastava, Emir Kaan Korukluoglu, Minh Nhat Le, Duyen Tran, Chau Minh Pham, Marzena Karpinska, Mohit Iyyer
| Challenge: | Large language models (LLMs) are known to memorize and recall English text from their pretraining data, but the extent to which this ability generalizes to non-English languages or transfers across languages remains unclear. |
| Approach: | They propose a dataset of 31.5K aligned excerpts from 20 books in ten languages, including English originals, official translations and new translations in six low-resource languages. |
| Outcome: | The proposed model can recall English content in translations, but perturbations reduce performance, causing the model to fail. |
Does quantization affect models’ performance on long-context tasks? (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language models support context windows exceeding 128K tokens, but this comes with significant memory requirements and high inference latency. |
| Approach: | They present the first systematic evaluation of quantized LLMs on tasks with long inputs and long-form outputs. |
| Outcome: | The proposed method preserves accuracy, while 4-bit methods lead to substantial losses . the results highlight the importance of a careful evaluation before deploying quantized LLMs . |
One Thousand and One Pairs: A “novel” challenge for long-context language models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing long-context evaluation methods measure surface-level retrieval capabilities, but do not assess performance on the more challenging task of synthesizing distant and underlying information. |
| Approach: | They propose a dataset of 1,001 minimally different pairs of true and false claims about 67 recently-published English fictional books. |
| Outcome: | The proposed model performs better on pairs that require only sentence-level retrieval vs. global reasoning . the proposed model also performs worse on speculative fiction books with extensive world-building . |
Revisiting Statistical Laws of Semantic Shift in Romance Cognates (2022.coling-1)
Copied to clipboard
| Challenge: | Despite their shared etymology, some cognate pairs have experienced semantic shift. |
| Approach: | They examine the relationship between lexical semantic shift and six intra-linguistic variables, such as frequency and polysemy, and examine the effect of morphologically complex etyma on semantic shift. |
| Outcome: | The results show that frequency and polysemy have positive effects on semantic shift and that morphologically complex etyma are more resistant to it. |
DEMETR: Diagnosing Evaluation Metrics for Translation (2022.emnlp-main)
Copied to clipboard
| Challenge: | BLEU scores are based on string overlap, but they are opaque in comparison to newer learned metrics. |
| Approach: | They propose a dataset to evaluate MT evaluation metrics based on linguistic perturbations in English . they find learned metrics perform substantially better than string-based metrics . |
| Outcome: | The proposed dataset shows that learned metrics perform better than string-based metrics . the dataset contains 31K English examples that cover 35 different linguistic phenomena . |
NarrativeTime: Dense Temporal Annotation on a Timeline (2024.lrec-main)
Copied to clipboard
| Challenge: | e.g. TimeBank contains 1-5% of all possible tlinks, and this information is underspecified in the text. |
| Approach: | They propose a timeline-based framework that achieves full coverage of all possible TLINKs. |
| Outcome: | The proposed framework achieves full coverage of all possible TLINKs in a text. |
An Interdisciplinary Approach to Human-Centered Machine Translation (2025.emnlp-main)
Copied to clipboard
Marine Carpuat, Omri Asscher, Kalika Bali, Luisa Bentivogli, Fred Blain, Lynne Bowker, Monojit Choudhury, Hal Daumé Iii, Kevin Duh, Ge Gao, Alvin C Grissom II, Marzena Karpinska, Elaine C Khoong, William D. Lewis, Andre Martins, Mary Nurminen, Douglas W. Oard, Maja Popovic, Michel Simard, François Yvon
| Challenge: | Despite progress in MT, a gap persists between how the technology is developed and how it is used in real-world contexts. |
| Approach: | They propose a human-centered approach to machine translation (MT) they argue that MT should be evaluated with diverse goals and contexts of use . |
| Outcome: | The proposed approach emphasizes alignment of evaluation and design with diverse communicative goals and contexts of use. |
The Perils of Using Mechanical Turk to Evaluate Open-Ended Text Generation (2021.emnlp-main)
Copied to clipboard
| Challenge: | Recent research has focused on open-ended text generation tasks because they are difficult to evaluate automatically. |
| Approach: | They conduct a survey of 45 open-ended text generation papers to determine whether models are reproducible . they then run story evaluation experiments with AMT workers and English teachers . |
| Outcome: | The results show that AMT workers and English teachers perform better when shown model-generated output alongside human-generated references. |