Papers by Tim Althoff
Responsible Evaluation of AI for Mental Health (2026.acl-long)
Copied to clipboard
Hiba Arnaout, Anmol Goel, H. Andrew Schwartz, Steffen T. Eberhardt, Dana Atzil-Slonim, Gavin Doherty, Brian Schwartz, Wolfgang Lutz, Tim Althoff, Munmun De Choudhury, Hamidreza Jamalabadi, Raj Sanjay Shah, Flor Miriam Plaza-del-Arco, Dirk Hovy, Maria Liakata, Iryna Gurevych
| Challenge: | Existing approaches to evaluating AI tools in this domain remain fragmented and inconsistent. |
| Approach: | They propose a taxonomy of AI mental health support types that integrates clinical soundness, social context, and equity to provide a structured basis for evaluation. |
| Outcome: | The proposed framework integrates clinical soundness, social context, and equity, providing a structured basis for evaluation. |
Cognitive Reframing of Negative Thoughts through Human-Language Model Interaction (2023.acl-long)
Copied to clipboard
Ashish Sharma, Kevin Rushton, Inna Lin, David Wadden, Khendra Lucas, Adam Miner, Theresa Nguyen, Tim Althoff
| Challenge: | Psychotherapy can help people overcome negative thoughts by replacing them with a more hopeful "reframed thought" but clinician shortages and mental health stigma often limit access to therapy. |
| Approach: | They propose a framework of seven linguistic attributes that can be used to reframe a thought . they use a retrieval-enhanced in-context learning model to generate reframed thoughts . |
| Outcome: | The proposed model is based on a human-centered study of 600 situations, thoughts and reframes on 2,000 mental health websites. |
Substance over Style: Evaluating Proactive Conversational Coaching Agents (2025.acl-long)
Copied to clipboard
Vidya Srinivas, Xuhai Xu, Xin Liu, Kumar Ayush, Isaac Galatzer-Levy, Shwetak Patel, Daniel McDuff, Tim Althoff
| Challenge: | Recent NLP research has focused on single-turn tasks with well-defined objectives or evaluation criteria. |
| Approach: | They describe five multi-turn coaching agents that exhibit distinct conversational styles and evaluate them through a user study. |
| Outcome: | The authors compare user feedback with third-person evaluations from health experts and an LM to find that stylistic components in absence of core functionality are viewed negatively. |
BLADE: Benchmarking Language Model Agents for Data-Driven Science (2024.findings-emnlp)
Copied to clipboard
Ken Gu, Ruoxi Shang, Ruien Jiang, Keying Kuang, Richard-John Lin, Donghe Lyu, Yue Mao, Youran Pan, Teng Wu, Jiaqian Yu, Yikun Zhang, Tianmai Zhang, Lanyi Zhu, Mike Merrill, Jeffrey Heer, Tim Althoff
| Challenge: | Language model-based agents can be used to conduct and support data-driven science, but evaluating them on open-ended tasks is challenging due to multiple valid approaches, partially correct steps, and different ways to express the same decisions. |
| Approach: | They propose a benchmark to automatically evaluate agents’ multifaceted approaches to open-ended research questions. |
| Outcome: | BLADE evaluates agents’ multifaceted approaches to open-ended research questions using data from 12 datasets and research questions drawn from existing scientific literature. |
A Computational Approach to Understanding Empathy Expressed in Text-Based Mental Health Support (2020.emnlp-main)
Copied to clipboard
| Challenge: | Empathy measurement has predominantly occurred in synchronous, face-to-face settings, and may not translate to asynchronous, text-based contexts. |
| Approach: | They propose a computational approach to understanding how empathy is expressed in online mental health platforms. |
| Outcome: | The proposed model can identify empathic conversations and extract rationales from them. |
IMBUE: Improving Interpersonal Effectiveness through Simulation and Just-in-time Feedback with Human-Language Model Interaction (2024.acl-long)
Copied to clipboard
| Challenge: | Various communication frameworks assist individuals in conducting difficult conversations by providing a set of skills to apply. |
| Approach: | They propose a language-based simulation system that provides just-in-time feedback to support the practice and learning of interpersonal effectiveness skills. |
| Outcome: | The proposed training system improves self-efficacy and reduces negative emotions by 27% compared to the GPT-4 training system. |
Gendered Mental Health Stigma in Masked Language Models (2022.emnlp-main)
Copied to clipboard
Inna Lin, Lucille Njoo, Anjalie Field, Ashish Sharma, Katharina Reinecke, Tim Althoff, Yulia Tsvetkov
| Challenge: | Mental health stigma prevents many individuals from receiving appropriate care, and social psychology studies have shown that mental health tends to be overlooked in men. |
| Approach: | They propose to use clinical psychology literature to curate prompts, then evaluate models’ propensity to generate gendered words. |
| Outcome: | The proposed framework captures stigma about gender in mental health and is more likely to predict female subjects than male in sentences about mental health conditions (32% vs. 19%), and this disparity is exacerbated for sentences that indicate treatment-seeking behavior. |
Inferring Events from Time Series using Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Prior work on reasoning about time series in conjunction with natural language has largely overlooked event descriptions and focused on tasks involving just numeric data like trend analysis or anomaly detection. |
| Approach: | They propose a method for generating tasks that test a model’s ability to reason about events associated with time series data based on sports data and develop a benchmarking method. |
| Outcome: | The proposed method can infer unobserved events from time series data, even when providing minimal context. |
Language Models Still Struggle to Zero-shot Reason about Time Series (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Time series are critical for decision-making in fields like finance and healthcare. |
| Approach: | They propose a framework for time series reasoning that includes formal tasks and a dataset of multi-scale time series paired with text captions across ten domains. |
| Outcome: | The proposed framework combines formal tasks and a dataset of multi-scale time series paired with text captions across ten domains to examine whether language models achieve three forms of reasoning. |
What Are the Odds? Language Models Are Capable of Probabilistic Reasoning (2024.emnlp-main)
Copied to clipboard
Akshay Paruchuri, Jake Garrison, Shun Liao, John Hernandez, Jacob Sunshine, Tim Althoff, Xin Liu, Daniel McDuff
| Challenge: | Language models (LMs) are capable of remarkably complex linguistic tasks, but numerical reasoning is an area in which they struggle. |
| Approach: | They evaluate the probabilistic reasoning capabilities of language models using idealized and real-world statistical distributions. |
| Outcome: | The proposed model can make inferences about distributions, even if assumptions are incorrect or misspecified. |