Papers by Thomas Hartvigsen
Decoding the Rule Book: Extracting Hidden Moderation Criteria from Reddit Communities (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to train classifiers that predict norm violations are often opacity-prone . a new approach to identify and extract these implicit criteria from historical moderation data is proposed . |
| Approach: | They propose to extract implicit criteria from historical moderation data using an interpretable architecture. |
| Outcome: | The proposed model replicates neural moderation models while providing transparent insights into decision-making processes. |
A Survey of Toxicity Mitigation Strategies for Multilingual Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Large language models can reproduce and amplify toxic content, including hate speech, harassment, and bias. |
| Approach: | They propose a comprehensive survey of the many detoxification methods tailored to multilingual LLMs. |
| Outcome: | The proposed methods are based on data filtering, style transfer, expert-based logit steering, retrieval augmentation, and human feedback. |
TAXI: Evaluating Categorical Knowledge Editing for Language Models (2024.findings-acl)
Copied to clipboard
| Challenge: | Knowledge editing aims to inject new facts into language models to improve factuality, but current benchmarks fail to evaluate consistency, which is critical to ensure efficient, accurate, and generalizable edits. |
| Approach: | They manually create a new benchmark dataset specifically created to evaluate consistency in categorical knowledge edits. |
| Outcome: | The results show that the editors achieve marginal, yet non-random consistency, and their consistency far underperforms human baselines. |
Efficient Knowledge Editing via Minimal Precomputation (2025.acl-short)
Copied to clipboard
| Challenge: | Knowledge editing methods like MEMIT require a one-time but significant computational cost. |
| Approach: | They propose to pre-compute 44 million hidden vectors per edited layer . authors show that this precomputation step is unnecessary . |
| Outcome: | The proposed methods can be performed by pre-computing a small portion of 44 million hidden vectors. |
ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection (2022.acl-long)
Copied to clipboard
| Challenge: | Toxic language detection systems often falsely flag text that contains minority group mentions as toxic . this over-reliance on spurious correlations also causes systems to struggle with detecting implicitly toxic language. |
| Approach: | They develop a machine-generated dataset of toxic and benign statements about 13 minority groups that generates subtly toxic and harmless text with a massive pretrained language model. |
| Outcome: | The proposed method can detect toxic and benign statements on a large scale . it can also detect hate speech on 94.5% of the toxic examples . |
ModelCitizens: Representing Community Voices in Online Safety (2025.emnlp-main)
Copied to clipboard
Ashima Suvarna, Christina A Chance, Karolina Naranjo, Hamid Palangi, Sophie Hao, Thomas Hartvigsen, Saadia Gabriel
| Challenge: | Existing toxic language detection models are trained on annotations that collapse diverse perspectives into a single ground truth. |
| Approach: | They propose to augment social media posts with conversational scenarios to reflect the impact of conversational context on toxicity. |
| Outcome: | The proposed model outperforms existing models on social media with conversational scenarios. |
MATHWELL: Generating Educational Math Word Problems Using Teacher Annotations (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing models and data fail to be educationally appropriate, causing teachers to write boilerplate questions and use boilerplate question sets. |
| Approach: | They propose that large language models (LLMs) can generate educational word problems by generating word problems using annotations from experts. |
| Outcome: | The proposed model generates more solvable, accurate, and appropriate word problems than public models while avoiding harmful questions. |
EDUMATH: Generating Standards-aligned Educational Math Word Problems (2026.acl-long)
Copied to clipboard
Bryan R Christ, Penelope Molitz, Beau LeBlond, Zachary Gottesman, Jonathan Kropko, Thomas Hartvigsen
| Challenge: | Math word problems (MWPs) are critical elements of K-12 math education and can be customized to students' interests and ability levels. |
| Approach: | They propose that LLMs can generate MWPs customized to student interests and math education standards by using an open and closed LLM to evaluate over 11,000 MWps and develop a teacher-annotated dataset for standards-aligned educational MWPS generation. |
| Outcome: | The proposed model outperforms existing closed models without training and is more similar to human-written MWPs but prefers customized MWPS with grade school students. |
Low-Bit Quantization Favors Undertrained LLMs (2025.acl-long)
Copied to clipboard
| Challenge: | Larger models or those trained on fewer tokens exhibit less quantization-induced degradation (QiD), while smaller, well-trained models face significant performance losses. |
| Approach: | They propose to use QiD to measure an LLM’s training levels and determine the number of training tokens required for fully training LLMs of various sizes. |
| Outcome: | The proposed scaling laws can predict the quantization performance of different-sized LLMs trained with tokens. |
Human Attention Maps for Text Classification: Do Humans and Neural Networks Focus on the Same Words? (2020.acl-main)
Copied to clipboard
| Challenge: | Attention-based models have been claimed to add interpretability, but little is known about the actual relationships between machine and human attention. |
| Approach: | They conduct the first quantitative assessment of human versus computational attention mechanisms for the text classification task. |
| Outcome: | The proposed models are compared against machine attention maps on a publicly available YELP dataset. |
Evaluating Temporal Consistency in Multi-Turn Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Language models are increasingly deployed in interactive settings where users reason about facts over time . we study temporal scope stability, the ability to preserve, override, or transfer time-scoped factual context across dialogue turns. |
| Approach: | They propose a diagnostic benchmark to isolate temporal scope stability in controlled multi-turn interactions. |
| Outcome: | The proposed model can preserve, override, or transfer time-scoped factual context across dialogue turns. |
TWEET-FID: An Annotated Dataset for Multiple Foodborne Illness Detection Tasks (2022.lrec-1)
Copied to clipboard
| Challenge: | Approximately 1 in 6 Americans (or 48 million people) are sickened by foodborne illness each year. |
| Approach: | They propose to use Twitter's TWEET-FID dataset to create annotated datasets for multiple foodborne illness incident detection tasks. |
| Outcome: | The proposed dataset is the first publicly available annotated dataset for multiple foodborne illness incident detection tasks. |
Inferring Events from Time Series using Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Prior work on reasoning about time series in conjunction with natural language has largely overlooked event descriptions and focused on tasks involving just numeric data like trend analysis or anomaly detection. |
| Approach: | They propose a method for generating tasks that test a model’s ability to reason about events associated with time series data based on sports data and develop a benchmarking method. |
| Outcome: | The proposed method can infer unobserved events from time series data, even when providing minimal context. |
Language Models Still Struggle to Zero-shot Reason about Time Series (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Time series are critical for decision-making in fields like finance and healthcare. |
| Approach: | They propose a framework for time series reasoning that includes formal tasks and a dataset of multi-scale time series paired with text captions across ten domains. |
| Outcome: | The proposed framework combines formal tasks and a dataset of multi-scale time series paired with text captions across ten domains to examine whether language models achieve three forms of reasoning. |
Language Models are Surprisingly Fragile to Drug Names in Biomedical Benchmarks (2024.findings-emnlp)
Copied to clipboard
Jack Gallifant, Shan Chen, Pedro Moreira, Nikolaj Munch, Mingye Gao, Jackson Pond, Leo Anthony Celi, Hugo Aerts, Thomas Hartvigsen, Danielle Bitterman
| Challenge: | Medical knowledge is context-dependent and requires consistent reasoning across various natural language expressions of semantically equivalent phrases. |
| Approach: | They create a robustness dataset to evaluate performance differences on medical benchmarks . they swap brand and generic drug names using physician expert annotations based on medical terminology . |
| Outcome: | The proposed model shows a consistent performance drop of 1-10% on medical benchmarks. |
Lifelong Model Editing with Graph-Based External Memory (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for post-training model editing suffer from overfitting and catastrophic forgetting. |
| Approach: | They propose a framework that leverages hyperbolic geometry and graph neural networks for precise and stable model edits. |
| Outcome: | Experiments on CounterFact, CounterFACT+, and MQuAKE with GPT2-XL and GPT-J show that HYPE significantly enhances edit stability, factual accuracy, and multi-hop reasoning. |
Math Neurosurgery: Isolating Language Models’ Math Reasoning Abilities Using Only Forward Passes (2025.acl-long)
Copied to clipboard
| Challenge: | Math reasoning is an active area of Large Language Model (LLM) research because it is a hallmark of artificial intelligence and has implications in several domains, including math education. |
| Approach: | They propose a method to isolate math-specific parameters in LLMs using only forward passes. |
| Outcome: | The proposed method improves a model's performance on GSM8K and MATH by 4-17% while leaving non-math behavior unaltered. |
Sparse Autoencoder Features for Classifications and Transferability (2025.emnlp-main)
Copied to clipboard
| Challenge: | Sparse Autoencoders (SAEs) provide potential for uncovering structured, human-interpretable representations in Large Language Models (LLMs). |
| Approach: | They analyze SAEs for interpretable feature extraction from Large Language Models in safety-critical classification tasks. |
| Outcome: | The proposed framework outperforms hidden-state and BoW models while demonstrating cross-lingual toxicity detection and visual classification tasks. |
Lifelong Knowledge Editing requires Better Regularization (2025.findings-emnlp)
Copied to clipboard
Akshat Gupta, Phudish Prateepamornkul, Maochuan Lu, Ahmed Alaa, Thomas Hartvigsen, Gopala Anumanchipalli
| Challenge: | Knowledge editing is a promising way to improve factuality in large language models, but recent studies have shown significant model degradation during sequential editing. |
| Approach: | They formalize locate-then-edit methods as a two-step fine-tuning process . they show that model degradation occurs due to over-optimization of internal activations . |
| Outcome: | The proposed methods reduce time and improve factuality by 42-61%. |