Papers by Mark Dredze
Task Matters: Knowledge Requirements Shape LLM Responses to Context–Memory Conflict (2026.findings-acl)
Copied to clipboard
| Challenge: | Prior work has shown that large language models favor parametric knowledge under conflict, but this setting assumes that tasks should always rely on the provided passage. |
| Approach: | They propose a model-agnostic diagnostic framework that holds underlying knowledge constant while injecting controlled conflicts across tasks with varying knowledge requirements. |
| Outcome: | Evaluating representative open-source LLMs, the proposed framework holds underlying knowledge constant while injecting controlled conflicts across tasks with varying knowledge requirements. |
Domain Generalizable AI Guardrails with Augmented Policy Training (2026.acl-long)
Copied to clipboard
| Challenge: | Current guardrails overfit the training policies, preventing adaptation to new domains and policies. |
| Approach: | They propose a training recipe that uses a suite of policy perturbation strategies to reduce overfitting and increase generalization to guardrails. |
| Outcome: | The proposed training recipe reduces overfitting and increases generalization on unseen policies and achieves comparable or better performance than existing 8B guardrails on unsen policies. |
Evaluating the Evaluators: Are readability metrics good measures of readability? (2025.emnlp-main)
Copied to clipboard
| Challenge: | Plain language summarization (PLS) aims to distill complex documents into accessible summaries for non-expert audiences. |
| Approach: | They conduct a thorough survey of literature on plain language summarization (PLS) and find that traditional readability metrics are not compared to human judgments. |
| Outcome: | The proposed language models better capture deeper measures of readability, with the best-performing model achieving a Pearson correlation of 0.56 with human judgments. |
Do LLMs Plan Like Human Writers? Comparing Journalist Coverage of Press Releases with LLMs (2024.emnlp-main)
Copied to clipboard
| Challenge: | Journalists engage in multiple steps in news writing that depend on human creativity, such as exploring different “angles” and selecting sources. |
| Approach: | They propose to use large language models to help journalists plan their news coverage . they find that LLMs recommend more creative angles and more informational sources . |
| Outcome: | The proposed models align better with humans when recommending angles, compared with informational sources. |
Multi-Task Transfer Matters During Instruction-Tuning (2024.findings-acl)
Copied to clipboard
| Challenge: | Instruction-tuning improves a model’s ability to learn in-context, but the mechanisms that drive in-constext learning are poorly understood. |
| Approach: | They propose to train a model on hundreds of tasks to improve its ability to learn in-context. |
| Outcome: | The proposed methods improve model transfer and in-context generalization, suggesting catastrophic forgetting may impact in-constext learning. |
LLMs are Better Than You Think: Label-Guided In-Context Learning for Named Entity Recognition (2025.emnlp-main)
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) tasks are performed using only a few demonstrations. |
| Approach: | They propose a method that leverages training labels through token-level statistics to improve ICL performance. |
| Outcome: | The proposed method outperforms existing methods on five NER datasets and is robust in low-resource settings. |
Do Models of Mental Health Based on Social Media Data Generalize? (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing literature on the validity of proxy-based methods for annotating mental health status in social media has raised new concerns regarding their use in clinical applications. |
| Approach: | They explore the generalization ability of machine learning classifiers trained to detect depression in individuals across multiple social media platforms. |
| Outcome: | The proposed methods show that they can be used to train and analyze large datasets and that they are robust to large dataset sizes. |
MedScore: Generalizable Factuality Evaluation of Open-ended Long-form Medical Answers by Domain-adapted Claim Decomposition and Verification (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing factuality evaluation pipelines are poor matches for medical domains . existing methods are limited to objective, entity-centric, formulaic texts . |
| Approach: | They propose a pipeline to decompose medical answers into condition-aware valid facts . they use a decomposition-then-verify approach to evaluate generated text . |
| Outcome: | The proposed method extracts up to three times as many valid facts as existing methods . the resulting factuality score substantially varies by decomposition method, corpus, and used backbone LLM . |
RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models (2025.naacl-long)
Copied to clipboard
| Challenge: | Efforts to ensure the safety of large language models (LLMs) include safety fine-tuning, evaluation, and red teaming. |
| Approach: | They conduct a comparative analysis of RAG and non-RAG frameworks with eleven LLMs to examine how RAG can make models less safe and change their safety profile. |
| Outcome: | The proposed methods are less effective than those used for non-RAG settings. |
Joint End-to-end Semantic Proto-role Labeling (2023.acl-short)
Copied to clipboard
| Challenge: | Existing systems for semantic proto-role labeling assign binary properties to arguments based on agent-like or patient-like properties. |
| Approach: | They propose to use a deep transformer model to model the performance of semantic proto-role labeling . they propose to include an error analysis to understand correlations between system stages . |
| Outcome: | The proposed system is robust in the presence of predicted arguments, the authors show . the proposed system also reduces annotation errors, the researchers conclude . |
A Closer Look at Claim Decomposition (2024.starsem-1)
Copied to clipboard
| Challenge: | Recent work uses claim decomposition to determine how well supported a claim is for applications in factual precision of generated text, entailment of human generated text and claim verification. |
| Approach: | They propose an LLM-based approach to generating decompositions inspired by Bertrand Russell’s theory of logical atomism and neo-Davidsonian semantics and demonstrate its improved decomposing quality over previous methods. |
| Outcome: | The proposed method improves on the FActScore and a Bertrand Russell-inspired approach to generating decompositions inspired by neo-Davidsonian semantics and improves decomposability quality. |
Probing Multimodal Large Language Models on Cognitive Biases in Chinese Short-Video Misinformation (2026.findings-acl)
Copied to clipboard
| Challenge: | Short-video platforms have become major channels for misinformation, but their robustness against misinformation entangled with cognitive biases remains under-explored. |
| Approach: | They propose a framework for evaluation of short-video platforms that use visual cues and social cue. |
| Outcome: | The proposed framework evaluates MLLMs across five modality settings. |
Gender and Racial Fairness in Depression Research using Social Media (2021.eacl-main)
Copied to clipboard
| Challenge: | Existing studies show that social media behavior can indicate mental health of an individual . previous studies have raised concerns about possible biases in models produced from such data, but no study has investigated how these biase recur with demographic groups. |
| Approach: | They analyze the fairness of depression classifiers trained on Twitter data with respect to gender and racial/ethnic demographic groups. |
| Outcome: | The proposed model performs better for gender and racial/ethnic groups than other models and provides recommendations on how to avoid biases in future research. |
Do Explicit Alignments Robustly Improve Multilingual Encoders? (2020.emnlp-main)
Copied to clipboard
| Challenge: | Explicit alignment objectives based on bitexts like Europarl and MultiUN have been shown to improve cross-lingual representations. |
| Approach: | They propose a new contrastive alignment objective that can better utilize bitexts . they propose to use a random sample of 1 million pair subset of OPUS data . |
| Outcome: | The proposed objective outperforms existing alignment objectives on a random 1 million pair subset of the OPUS dataset. |
Bernice: A Multilingual Pre-trained Encoder for Twitter (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing language models for Twitter are monolingual, adapted from other domains, or trained on limited amount of in-domain data. |
| Approach: | They propose a multilingual RoBERTa language model that is trained from scratch on 2.5 billion tweets with a custom tweet-focused tokenizer. |
| Outcome: | The proposed model outperforms or matches models trained on monolingual and multilingual tweets on a variety of benchmarks and is more efficient compute- and data-wise to train completely on in-domain data with a specialized domain-specific tokenizer. |
Evaluating Biases in Context-Dependent Sexual and Reproductive Health Questions (2024.findings-emnlp)
Copied to clipboard
| Challenge: | With the rise in accessibility of chat-based large language models, the public increasingly uses them as question-answering systems for personalized answers. |
| Approach: | They curate a dataset of sexual and reproductive healthcare questions dependent on age, sex, and location attributes and compare their outputs with and without demographic context to determine answer alignment . |
| Outcome: | The results show that young adult female users are favored in the model answers to underspecified questions in the healthcare domain. |
Beto, Bentz, Becas: The Surprising Cross-Lingual Effectiveness of BERT (D19-1)
Copied to clipboard
| Challenge: | Pretrained contextual representation models have pushed forward the state-of-the-art on many NLP tasks. |
| Approach: | They propose to use a model that is pretrained on 104 languages for cross-lingual transfer. |
| Outcome: | The proposed model performs well on 5 NLP tasks covering 39 languages from various language families. |
Geo-Seq2seq: Twitter User Geolocation on Noisy Data through Sequence to Sequence Learning (2023.findings-acl)
Copied to clipboard
| Challenge: | a new method for Twitter user geolocation rewrites noisy, multilingual location strings into structured English location names. |
| Approach: | They propose a sequence-to-sequence (seq2sequ) model that rewrites noisy location strings into structured English location names. |
| Outcome: | The proposed model can generalize well to unseen temporal data, but performance does vary by language and country. |
Gender Bias in Decision-Making with Large Language Models: A Study of Relationship Conflicts (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models acquire beliefs about gender from training data and can therefore generate text with stereotypical gender attitudes. |
| Approach: | They use a decision-making lens to examine gender equity within large language models . they explore relationships through typical and gender-neutral names . |
| Outcome: | The proposed model generation and classification models exhibit stereotypical gender biases . the proposed model generates gender-neutral names, with and without safety enhancements, and egalitarian versus traditional scenarios across topics. |
Schema-Driven Information Extraction from Heterogeneous Tables (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing work on information extraction from tables has focused on developing custom pipelines for each table collection. |
| Approach: | They propose a task that transforms tabular data into structured records following a human-authored schema. |
| Outcome: | The proposed task achieves F1 scores ranging from 74.2 to 96.1 while maintaining cost efficiency. |
Making FETCH! Happen: Finding Emergent Dog Whistles Through Common Habitats (2025.acl-long)
Copied to clipboard
| Challenge: | Dog whistles are coded expressions with dual meanings that slip by content moderation filters . a new study finds that state-of-the-art systems fail to identify novel dog whistles . |
| Approach: | They propose a task to find novel dog whistles in massive social media corpora . they use a strong baseline system that combines vector databases and Large Language Models to identify new dog whistle. |
| Outcome: | The proposed system fails to identify dog whistles across three social media cases . it combines vector databases and Large Language Models to efficiently and effectively identify new dog whistle expressions. |
DnDScore: Decontextualization and Decomposition for Factuality Verification in Long-Form Text Generation (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent measures of factual precision use a decompose-then-verify framework . decontextualization is the process of augmenting subclaims with necessary context . |
| Approach: | They evaluate different decomposition, decontextualization and verification strategies . they introduce a deconstructualization aware verification method that validates subclaims in context . |
| Outcome: | The proposed method decomposes claims and independently verifyes them . it introduces a decontextualization aware verification method that validates subclaims in context . |
Characterization of Stigmatizing Language in Medical Records (2023.acl-short)
Copied to clipboard
Keith Harrigian, Ayah Zirikly, Brant Chee, Alya Ahmad, Anne Links, Somnath Saha, Mary Catherine Beach, Mark Dredze
| Challenge: | Widespread disparities in healthcare outcomes exist between demographic groups in the United States. |
| Approach: | They characterize disparities in medical documentation using domain-informed NLP techniques and highlight important differences between them. |
| Outcome: | The proposed methods highlight important differences between the task and bias-related tasks studied within the NLP community. |
Challenges of Using Text Classifiers for Causal Inference (D18-1)
Copied to clipboard
| Challenge: | a number of scientific analyses focus on low-dimensional structured data, but text classifiers can be used to produce structured variables. |
| Approach: | They propose to use text classifiers to conduct causal analyses on simulated and Yelp data. |
| Outcome: | The proposed method can be used on simulated and Yelp data. |
Academics Can Contribute to Domain-Specialized Language Models (2024.emnlp-main)
Copied to clipboard
Mark Dredze, Genta Winata, Prabhanjan Kambadur, Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, David Rosenberg, Sebastian Gehrmann
| Challenge: | Commercially available models dominate academic leaderboards, focusing on creating and adapting general-purpose models . however, general- purpose models often underperform in specialized domains, and domain-specific models yield superior results. |
| Approach: | They advocate for a renewed focus on developing and evaluating domain- and task-specific models . they advocate for an adapted or adapted model that can be used to improve academic leaderboard standings . |
| Outcome: | The proposed model can do well on professional and linguistic examinations, college-level knowledge questions, and collections of reasoning tasks. |
Do Text-to-Text Multi-Task Learners Suffer from Task Conflict? (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing multi-task learning architectures learn a single model across multiple tasks through a shared encoder followed by task-specific decoders. |
| Approach: | They propose to use a shared encoder and language model decoder to learn a single model across multiple tasks. |
| Outcome: | The proposed architecture does surprisingly well across a range of diverse tasks. |
Sources of Transfer in Multilingual Named Entity Recognition (2020.acl-main)
Copied to clipboard
| Challenge: | naive training of named-entity recognition models using annotated data from multiple languages consistently underperforms monolingual models. |
| Approach: | They propose a polyglot named-entity recognition model where one model is trained using annotated data drawn from multiple languages. |
| Outcome: | The proposed model outperforms models trained on monolingual data despite more training data . the proposed model shares many parameters across languages and fine-tunes them to outperFORM monolingual models. |
Clinical Concept Linking with Contextualized Neural Representations (2020.acl-main)
Copied to clipboard
| Challenge: | Entity linking systems rely on three sources of information: 1) similarity between mention string and entity name; 2) similarity of context of document to entity; 3) broader information about knowledge base; 4) contextual information; 5) semantic information; and 6) semantic information. |
| Approach: | They propose an approach to linking medical concepts to a medical concept ontology that leverages recent work in contextualized neural models. |
| Outcome: | The proposed approach outperforms a baseline approach and provides better initialization for the ranker. |
Fine-tuning Encoders for Improved Monolingual and Zero-shot Polylingual Neural Topic Modeling (2021.naacl-main)
Copied to clipboard
| Challenge: | Topic models can augment or replace bag-of-words inputs with pre-trained transformer-based word prediction models. |
| Approach: | They propose several methods for fine-tuning encoders to improve both monolingual and zero-shot polylingual topic modeling. |
| Outcome: | The proposed methods improve both monolingual and zero-shot polylingual topic modeling. |
MixCE: Training Autoregressive Language Models by Mixing Forward and Reverse Cross-Entropies (2023.acl-long)
Copied to clipboard
| Challenge: | Autoregressive language models are trained by minimizing the cross-entropy of the model distribution Q relative to the data distribution P. However, these systems still struggle in many openended generation settings, where they are asked to produce a long text following a short prompt. |
| Approach: | They propose to combine forward and reverse cross-entropy to train autoregressive language models by minimizing the cross-Entropy of the model distribution Q relative to the data distribution P. |
| Outcome: | The proposed model overgeneralizes and produces non-human-like text without complex decoding strategies. |
Updated Headline Generation: Creating Updated Summaries for Evolving News Stories (2022.acl-long)
Copied to clipboard
| Challenge: | Existing systems that generate headlines for updated articles are not as efficient as static ones. |
| Approach: | They propose a task where a system generates a headline for an updated article, considering both the previous article and headline. |
| Outcome: | The proposed model produces headlines judged by humans to be as factual as gold headlines while making fewer unnecessary edits compared to a standard headline generation model. |
Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to eliminate implicit biases in LLMs do not eradicate underlying behavioral bias. |
| Approach: | They propose a framework that uses logic grid puzzles to probe the influence of social stereotypes on logical reasoning and decision making in LLMs. |
| Outcome: | The proposed framework systematically probes the influence of social stereotypes on logical reasoning and decision making in LLMs. |
Deep Dirichlet Multinomial Regression (N18-1)
Copied to clipboard
| Challenge: | supervised topic models can incorporate arbitrary document-level features to inform topic priors, but their ability to model corpora is limited by the representation and selection of these features. |
| Approach: | They propose a generative topic model that simultaneously learns document feature representations and topics. |
| Outcome: | The proposed model outperforms DMR and LDA on three datasets and human subjects judge it more representative of associated document features. |
Benchmarking Large Language Models on Answering and Explaining Challenging Medical Questions (2025.naacl-long)
Copied to clipboard
| Challenge: | Medical board exams or general clinical questions do not capture the complexity of real clinical cases. |
| Approach: | They construct two datasets that are structured as multiple-choice question-answering tasks accompanied by expert-written explanations. |
| Outcome: | The proposed datasets are harder than previous benchmarks. |
Everything Is All It Takes: A Multipronged Strategy for Zero-Shot Cross-Lingual Information Extraction (2021.emnlp-main)
Copied to clipboard
Mahsa Yarmohammadi, Shijie Wu, Marc Marone, Haoran Xu, Seth Ebner, Guanghui Qin, Yunmo Chen, Jialiang Guo, Craig Harman, Kenton Murray, Aaron Steven White, Mark Dredze, Benjamin Van Durme
| Challenge: | Zero-shot cross-lingual information extraction (IE) is a technique for training data in a source language but not in . |
| Approach: | They explore techniques including data projection and self-training to improve zero-shot cross-lingual information extraction (IE) IE is a construction of an IE model for some target language given existing annotations exclusively in English. |
| Outcome: | The proposed techniques show that they perform better than any single strategy. |
Cross-Lingual Transfer in Zero-Shot Cross-Language Entity Linking (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing work on cross-language entity linking grounds mentions written in multiple languages to a monolingual knowledge base is lacking. |
| Approach: | They propose a task that uses multilingual BERT representations of both the mention and context as input and explore zero-shot language transfer. |
| Outcome: | The proposed model performs well in both monolingual and multilingual settings. |