Papers by H. Schwartz

10 papers
Evaluation of LLMs-based Hidden States as Author Representations for Psychological Human-Centered NLP Tasks (2025.findings-naacl)

Copied to clipboard

Challenge: Many human-centered NLP tasks focus on assessing human-attributes of a user based on their language.
Approach: They evaluate different ways of representing documents and users using different LM and HuLM architectures to predict task outcomes as dynamically changing states and averaged trait-like user-level attributes.
Outcome: The proposed representations predict valence, arousal, empathy, and distress as well as trait-like user-level attributes.
WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher Learning (2025.acl-long)

Copied to clipboard

Challenge: Current speech encoding pipelines rely on an additional text-based LM to get robust representations of human communication, even though speech-to-text models often have a LM within.
Approach: They propose to align Whisper's latent space with semantic representations from a text autoencoder and lexically derived embeddings of basic psychological dimensions: emotion and personality.
Outcome: The proposed approach surpasses current speech encoders over self-supervised affective tasks and downstream psychological tasks, achieving an error reduction of 73.4% and 83.8%, respectively.
SOCIALITE-LLAMA: An Instruction-Tuned Model for Social Scientific Tasks (2024.eacl-short)

Copied to clipboard

Challenge: Social science NLP tasks require large data to capture semantics and implicit pragmatics.
Approach: They propose an open-source instruction tuning tool for social science NLP tasks that captures implicit pragmatic cues from text.
Outcome: The proposed model matches or improves on a state-of-the-art, multi-task finetuned model on 80% of social tasks.
Residualized Similarity for Faithfully Explainable Authorship Verification (2025.findings-emnlp)

Copied to clipboard

Challenge: Neural methods achieve high accuracy, but their representations lack direct interpretability.
Approach: They propose a method that supplements systems using interpretable features with a neural network to improve their performance while maintaining interpretability.
Outcome: The proposed method improves the performance of state-of-the-art models while maintaining interpretability.
ALBA: Adaptive Language-Based Assessments for Mental Health (2024.naacl-long)

Copied to clipboard

Challenge: Adaptive language-based assessments require a substantial sample of words per person for accuracy.
Approach: They propose an adaptive language-based assessment task that involves ordering questions and scoring latent psychological trait using limited language responses to previous questions.
Outcome: The proposed methods improve over non-adaptive baselines, but are more accurate and scalable with fewer questions.
Idiosyncratic Versus Normative Modeling of Atypical Speech Recognition: Dysarthric Case Studies (2025.emnlp-main)

Copied to clipboard

Challenge: Past studies have focused on fully personalized (or idiosyncratic) models for atypical speech . past studies focused on idiotic models, but current approaches focus on generalizing and handling idiomatic patterns .
Approach: They compare four models that generalize and handle idiosyncrasy to find atypical speech . they find the dysarthric-idios-ync model performs better than the idioconic approach .
Outcome: The proposed model generalizes and handles idiosyncrasy better than the idiocy model . the model requires less personalized data and reduces word error rate from 71% to 32% .
Large Human Language Models: A Need and the Challenges (2024.naacl-long)

Copied to clipboard

Challenge: a growing recognition of the importance of modeling human and social factors into human-centered NLP models . authors advocate for three positions toward creating large human language models based on psychological and behavioral sciences .
Approach: et al. advocate for three positions toward creating large human language models . they argue that LM training should include the human context and recognize that people are more than their group .
Outcome: a new study shows that learning language from linguistic signals alone is not adequate, according to a recent paper . authors advocate for three positions toward creating large human language models . a human-centered model should include the human context, and account for the dynamic nature of the human environment, they say .
Systematic Evaluation of Auto-Encoding and Large Language Model Representations for Capturing Author States and Traits (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used in human-centered applications, yet their ability to model diverse psychological constructs is not well understood.
Approach: They evaluated a range of Transformer-LMs to predict psychological variables across five major dimensions: affect, substance use, mental health, sociodemographics, and personality.
Outcome: The models predict affect, substance use, mental health, sociodemographics, and personality across five major dimensions.
Capturing Author Self Beliefs in Social Media Language (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for identifying self beliefs are limited.
Approach: They propose a task that classifies language that contains explicit or implicit mentions of the author's self beliefs using an annotated set of 2,000 human-annotated self beliefs, 100,000 LLM-labeled examples, and 10,000 surveyed self belief paragraphs.
Outcome: The proposed model outperforms OpenAI’s state-of-the-art GPT-4o model in the AUC of 0.944 and annotates 2,000 human-annotated self beliefs, 100,000 LLM-labeled examples, and 10,000 surveyed self belief paragraphs.
Capturing Human Cognitive Styles with Language: Towards an Experimental Evaluation Paradigm (2025.naacl-short)

Copied to clipboard

Challenge: While NLP models often capture cognitive states via language, validity of predicted states is determined by comparing annotations created without access to the cognitive states of the authors.
Approach: They propose a framework for evaluating language-based cognitive style models against human behavior by using an experiment-based framework.
Outcome: The proposed framework shows that language features can predict participants’ decision style with moderate-to-high accuracy (AUC 0.8), demonstrating that cognitive style can be partly captured and revealed by discourse patterns.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations