Papers by H. Schwartz
Evaluation of LLMs-based Hidden States as Author Representations for Psychological Human-Centered NLP Tasks (2025.findings-naacl)
Copied to clipboard
| Challenge: | Many human-centered NLP tasks focus on assessing human-attributes of a user based on their language. |
| Approach: | They evaluate different ways of representing documents and users using different LM and HuLM architectures to predict task outcomes as dynamically changing states and averaged trait-like user-level attributes. |
| Outcome: | The proposed representations predict valence, arousal, empathy, and distress as well as trait-like user-level attributes. |
WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher Learning (2025.acl-long)
Copied to clipboard
Rajath Rao, Adithya V Ganesan, Oscar Kjell, Jonah Luby, Akshay Raghavan, Scott M. Feltman, Whitney Ringwald, Ryan L. Boyd, Benjamin J. Luft, Camilo J. Ruggero, Neville Ryant, Roman Kotov, H. Schwartz
| Challenge: | Current speech encoding pipelines rely on an additional text-based LM to get robust representations of human communication, even though speech-to-text models often have a LM within. |
| Approach: | They propose to align Whisper's latent space with semantic representations from a text autoencoder and lexically derived embeddings of basic psychological dimensions: emotion and personality. |
| Outcome: | The proposed approach surpasses current speech encoders over self-supervised affective tasks and downstream psychological tasks, achieving an error reduction of 73.4% and 83.8%, respectively. |
SOCIALITE-LLAMA: An Instruction-Tuned Model for Social Scientific Tasks (2024.eacl-short)
Copied to clipboard
Gourab Dey, Adithya V Ganesan, Yash Kumar Lal, Manal Shah, Shreyashee Sinha, Matthew Matero, Salvatore Giorgi, Vivek Kulkarni, H. Schwartz
| Challenge: | Social science NLP tasks require large data to capture semantics and implicit pragmatics. |
| Approach: | They propose an open-source instruction tuning tool for social science NLP tasks that captures implicit pragmatic cues from text. |
| Outcome: | The proposed model matches or improves on a state-of-the-art, multi-task finetuned model on 80% of social tasks. |
Residualized Similarity for Faithfully Explainable Authorship Verification (2025.findings-emnlp)
Copied to clipboard
Peter Zeng, Pegah Alipoormolabashi, Jihu Mun, Gourab Dey, Nikita Soni, Niranjan Balasubramanian, Owen Rambow, H. Schwartz
| Challenge: | Neural methods achieve high accuracy, but their representations lack direct interpretability. |
| Approach: | They propose a method that supplements systems using interpretable features with a neural network to improve their performance while maintaining interpretability. |
| Outcome: | The proposed method improves the performance of state-of-the-art models while maintaining interpretability. |
ALBA: Adaptive Language-Based Assessments for Mental Health (2024.naacl-long)
Copied to clipboard
| Challenge: | Adaptive language-based assessments require a substantial sample of words per person for accuracy. |
| Approach: | They propose an adaptive language-based assessment task that involves ordering questions and scoring latent psychological trait using limited language responses to previous questions. |
| Outcome: | The proposed methods improve over non-adaptive baselines, but are more accurate and scalable with fewer questions. |
Idiosyncratic Versus Normative Modeling of Atypical Speech Recognition: Dysarthric Case Studies (2025.emnlp-main)
Copied to clipboard
| Challenge: | Past studies have focused on fully personalized (or idiosyncratic) models for atypical speech . past studies focused on idiotic models, but current approaches focus on generalizing and handling idiomatic patterns . |
| Approach: | They compare four models that generalize and handle idiosyncrasy to find atypical speech . they find the dysarthric-idios-ync model performs better than the idioconic approach . |
| Outcome: | The proposed model generalizes and handles idiosyncrasy better than the idiocy model . the model requires less personalized data and reduces word error rate from 71% to 32% . |
Large Human Language Models: A Need and the Challenges (2024.naacl-long)
Copied to clipboard
| Challenge: | a growing recognition of the importance of modeling human and social factors into human-centered NLP models . authors advocate for three positions toward creating large human language models based on psychological and behavioral sciences . |
| Approach: | et al. advocate for three positions toward creating large human language models . they argue that LM training should include the human context and recognize that people are more than their group . |
| Outcome: | a new study shows that learning language from linguistic signals alone is not adequate, according to a recent paper . authors advocate for three positions toward creating large human language models . a human-centered model should include the human context, and account for the dynamic nature of the human environment, they say . |
Systematic Evaluation of Auto-Encoding and Large Language Model Representations for Capturing Author States and Traits (2025.findings-acl)
Copied to clipboard
Khushboo Singh, Vasudha Varadarajan, Adithya V Ganesan, August Håkan Nilsson, Nikita Soni, Syeda Mahwish, Pranav Chitale, Ryan L. Boyd, Lyle Ungar, Richard N Rosenthal, H. Schwartz
| Challenge: | Large Language Models (LLMs) are increasingly used in human-centered applications, yet their ability to model diverse psychological constructs is not well understood. |
| Approach: | They evaluated a range of Transformer-LMs to predict psychological variables across five major dimensions: affect, substance use, mental health, sociodemographics, and personality. |
| Outcome: | The models predict affect, substance use, mental health, sociodemographics, and personality across five major dimensions. |
Capturing Author Self Beliefs in Social Media Language (2025.acl-long)
Copied to clipboard
Siddharth Mangalik, Adithya V Ganesan, Abigail B. Wheeler, Nicholas Kerry, Jeremy D. W. Clifton, H. Schwartz, Ryan L. Boyd
| Challenge: | Existing methods for identifying self beliefs are limited. |
| Approach: | They propose a task that classifies language that contains explicit or implicit mentions of the author's self beliefs using an annotated set of 2,000 human-annotated self beliefs, 100,000 LLM-labeled examples, and 10,000 surveyed self belief paragraphs. |
| Outcome: | The proposed model outperforms OpenAI’s state-of-the-art GPT-4o model in the AUC of 0.944 and annotates 2,000 human-annotated self beliefs, 100,000 LLM-labeled examples, and 10,000 surveyed self belief paragraphs. |
Capturing Human Cognitive Styles with Language: Towards an Experimental Evaluation Paradigm (2025.naacl-short)
Copied to clipboard
Vasudha Varadarajan, Syeda Mahwish, Xiaoran Liu, Julia Buffolino, Christian Luhmann, Ryan L. Boyd, H. Schwartz
| Challenge: | While NLP models often capture cognitive states via language, validity of predicted states is determined by comparing annotations created without access to the cognitive states of the authors. |
| Approach: | They propose a framework for evaluating language-based cognitive style models against human behavior by using an experiment-based framework. |
| Outcome: | The proposed framework shows that language features can predict participants’ decision style with moderate-to-high accuracy (AUC 0.8), demonstrating that cognitive style can be partly captured and revealed by discourse patterns. |