Papers by Marzena Karpinska

15 papers
CaLMQA: Exploring culturally specific long-form question answering across 23 languages (2025.acl-long)

Copied to clipboard

Challenge: Despite rising global usage of large language models, their ability to generate *long-form* answers to *culturally specific* questions remains unexplored in many languages.
Approach: They perform the first study of textual multilingual long-form QA by creating a dataset of culturally specific questions across 23 different languages.
Outcome: The results show that the best models make critical surface-level errors for many languages and their understanding of diverse cultures.
Program Chairs’ Report on Peer Review at ACL 2023 (2023.acl-long)

Copied to clipboard

Challenge: ACL'23 makes its peer review report public and an official part of the conference proceedings.
Approach: They present an analysis of the factors affecting peer review and identify the most problematic issues that the authors complained about.
Outcome: The authors identified the most problematic issues and provided suggestions for the future chairs.
AI use in American newspapers is widespread, uneven, and rarely disclosed (2026.acl-long)

Copied to clipboard

Challenge: a large-scale dataset of 186K articles from 1.5K newspapers published in the summer of 2025 is audited.
Approach: They audit 186K articles from 1.5K newspapers published in summer of 2025 . they use Pangram, a state-of-the-art AI detector, to detect whether articles are partially or fully AI-generated .
Outcome: The findings highlight the need for greater transparency and updated editorial standards regarding the use of AI in journalism to maintain public trust.
Exploring Document-Level Literary Machine Translation with Parallel Paragraphs from World Literature (2022.emnlp-main)

Copied to clipboard

Challenge: Literary translation is a culturally significant task, but it is bottlenecked by the small number of qualified literary translators . a dataset of non-English language novels is used to study literary MT .
Approach: They use a dataset of non-English language novels aligned to human and automatic English translations to study literary MT.
Outcome: The proposed model prefers human translations over machine translations at a rate of 84% . state-of-the-art MT metrics do not correlate with preferences, the study finds .
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text (2025.acl-long)

Copied to clipboard

Challenge: Qualitative analysis of experts’ free-form explanations shows that while they rely heavily on specific lexical clues (‘AI vocabulary’), they also pick up on more complex phenomena within the text (e.g., formality, originality, clarity).
Approach: They hire annotators to read 300 non-fiction English articles, label them as either human-written or AI-generated, and provide paragraph-length explanations for their decisions.
Outcome: The annotators who frequently use LLMs for writing tasks outperform commercial and open-source detectors even without evasion tactics like paraphrasing and humanization.
ezCoref: Towards Unifying Annotation Guidelines for Coreference Resolution (2023.findings-eacl)

Copied to clipboard

Challenge: Existing datasets vary in definition of coreferences and are curated for linguistic experts.
Approach: They propose to use ezCoref to create a crowdsourcing-friendly coreference annotation methodology that teaches annotators only cases that are treated similarly across existing datasets.
Outcome: The proposed method reannotates 240 passages from seven existing english coreference datasets while teaching annotators only cases that are treated similarly across them.
OWL: Probing Cross-Lingual Recall of Memorized Texts via World Literature (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are known to memorize and recall English text from their pretraining data, but the extent to which this ability generalizes to non-English languages or transfers across languages remains unclear.
Approach: They propose a dataset of 31.5K aligned excerpts from 20 books in ten languages, including English originals, official translations and new translations in six low-resource languages.
Outcome: The proposed model can recall English content in translations, but perturbations reduce performance, causing the model to fail.
Does quantization affect models’ performance on long-context tasks? (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models support context windows exceeding 128K tokens, but this comes with significant memory requirements and high inference latency.
Approach: They present the first systematic evaluation of quantized LLMs on tasks with long inputs and long-form outputs.
Outcome: The proposed method preserves accuracy, while 4-bit methods lead to substantial losses . the results highlight the importance of a careful evaluation before deploying quantized LLMs .
One Thousand and One Pairs: A “novel” challenge for long-context language models (2024.emnlp-main)

Copied to clipboard

Challenge: Existing long-context evaluation methods measure surface-level retrieval capabilities, but do not assess performance on the more challenging task of synthesizing distant and underlying information.
Approach: They propose a dataset of 1,001 minimally different pairs of true and false claims about 67 recently-published English fictional books.
Outcome: The proposed model performs better on pairs that require only sentence-level retrieval vs. global reasoning . the proposed model also performs worse on speculative fiction books with extensive world-building .
Revisiting Statistical Laws of Semantic Shift in Romance Cognates (2022.coling-1)

Copied to clipboard

Challenge: Despite their shared etymology, some cognate pairs have experienced semantic shift.
Approach: They examine the relationship between lexical semantic shift and six intra-linguistic variables, such as frequency and polysemy, and examine the effect of morphologically complex etyma on semantic shift.
Outcome: The results show that frequency and polysemy have positive effects on semantic shift and that morphologically complex etyma are more resistant to it.
DEMETR: Diagnosing Evaluation Metrics for Translation (2022.emnlp-main)

Copied to clipboard

Challenge: BLEU scores are based on string overlap, but they are opaque in comparison to newer learned metrics.
Approach: They propose a dataset to evaluate MT evaluation metrics based on linguistic perturbations in English . they find learned metrics perform substantially better than string-based metrics .
Outcome: The proposed dataset shows that learned metrics perform better than string-based metrics . the dataset contains 31K English examples that cover 35 different linguistic phenomena .
NarrativeTime: Dense Temporal Annotation on a Timeline (2024.lrec-main)

Copied to clipboard

Challenge: e.g. TimeBank contains 1-5% of all possible tlinks, and this information is underspecified in the text.
Approach: They propose a timeline-based framework that achieves full coverage of all possible TLINKs.
Outcome: The proposed framework achieves full coverage of all possible TLINKs in a text.
An Interdisciplinary Approach to Human-Centered Machine Translation (2025.emnlp-main)

Copied to clipboard

Challenge: Despite progress in MT, a gap persists between how the technology is developed and how it is used in real-world contexts.
Approach: They propose a human-centered approach to machine translation (MT) they argue that MT should be evaluated with diverse goals and contexts of use .
Outcome: The proposed approach emphasizes alignment of evaluation and design with diverse communicative goals and contexts of use.
The Perils of Using Mechanical Turk to Evaluate Open-Ended Text Generation (2021.emnlp-main)

Copied to clipboard

Challenge: Recent research has focused on open-ended text generation tasks because they are difficult to evaluate automatically.
Approach: They conduct a survey of 45 open-ended text generation papers to determine whether models are reproducible . they then run story evaluation experiments with AMT workers and English teachers .
Outcome: The results show that AMT workers and English teachers perform better when shown model-generated output alongside human-generated references.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations