Papers by Serina Chang
Valid Survey Simulations with Limited Human Data: The Roles of Prompting, Fine-Tuning, and Rectification (2026.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) are a cost-effective and time-consuming way to capture public opinion and behavior, but their outputs are often biased and yield invalid estimates. |
| Approach: | They propose to use large language models to generate survey responses and rectification methods that debias population estimates to find out how human responses are best allocated between them. |
| Outcome: | The proposed methods reduce bias below 5% and increase sample size by up to 14% under a fixed budget. |
Detecting Gang-Involved Escalation on Social Media Using Context (D18-1)
Copied to clipboard
Serina Chang, Ruiqi Zhong, Ethan Adams, Fei-Tzin Lee, Siddharth Varia, Desmond Patton, William Frey, Chris Kedzie, Kathy McKeown
| Challenge: | In cities such as Chicago, gang-involved youth have increasingly turned to social media to post about their experiences and intents online. |
| Approach: | They propose a system that uses domain-specific resources and contextual representations of the emotional and semantic content of the user’s recent tweets and their interactions with other users to detect Aggression and Loss in social media posts. |
| Outcome: | The proposed system improves on a large unlabeled dataset and incorporates contextual representations of the emotional and semantic content of the user’s recent tweets as well as their interactions with other users. |
ChatBench: From Static Benchmarks to Human-AI Evaluation (2025.acl-long)
Copied to clipboard
| Challenge: | In 2024, 40% of US adults reported using generative AI in their everyday lives, an unprecedented rate of adoption for a new technology. |
| Approach: | They propose to convert MMLU questions into user-AI conversations by seeding the user with the question and having them carry out a conversation with the LLM to answer their question. |
| Outcome: | The proposed model can estimate user-AI accuracy by fine-tuning a user simulator on a subset of ChatBench. |
Language Model Fine-Tuning on Scaled Survey Data for Predicting Distributions of Public Opinions (2025.acl-long)
Copied to clipboard
| Challenge: | Prior studies have failed to accurately predict distribution of survey responses from human subjects. |
| Approach: | They propose to fine-tune large language models to predict human response distributions by leveraging unique structural characteristics of survey data. |
| Outcome: | The proposed model can capture group-specific variability in public opinions, generalizing to unseen subpopulations, survey waves and question topics, and different survey families. |
Graph-Based Alternatives to LLMs for Human Simulation (2026.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) are a popular approach for simulating human behaviors, yet it remains unclear if they are necessary for all simulation tasks. |
| Approach: | They propose a graph neural network that can match or surpass strong LLMs for close-ended simulations. |
| Outcome: | The proposed model outperforms strongest LLM-based methods across three datasets and three evaluation settings. |
Automatically Inferring Gender Associations from Language (D19-1)
Copied to clipboard
| Challenge: | In this paper, we demonstrate that there are large-scale differences in the ways that people talk about women and men and that these differences vary across domains. |
| Approach: | They propose to integrate two datasets and a novel approach to automatically infer gender associations from language and find coherent word clusters and label clusters for the semantic concepts they represent. |
| Outcome: | The proposed methods outperform strong baselines in large-scale studies of how people talk about women and men in two different settings. |