Challenge: Prior studies have failed to accurately predict distribution of survey responses from human subjects.
Approach: They propose to fine-tune large language models to predict human response distributions by leveraging unique structural characteristics of survey data.
Outcome: The proposed model can capture group-specific variability in public opinions, generalizing to unseen subpopulations, survey waves and question topics, and different survey families.

Similar Papers

Specializing Large Language Models to Simulate Survey Response Distributions for Global Populations (2025.naacl-long)

Copied to clipboard

Challenge: Prior work has focused on using large language models to simulate human behaviors . but, LLMs are known to generate erroneous, stereotypical, or overconfident answers .
Approach: They propose to specialize large language models for simulating survey response distributions by first-token probabilities.
Outcome: The proposed model outperforms other methods and zero-shot classifiers on unseen questions, countries, and a completely unseened survey.
Valid Survey Simulations with Limited Human Data: The Roles of Prompting, Fine-Tuning, and Rectification (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are a cost-effective and time-consuming way to capture public opinion and behavior, but their outputs are often biased and yield invalid estimates.
Approach: They propose to use large language models to generate survey responses and rectification methods that debias population estimates to find out how human responses are best allocated between them.
Outcome: The proposed methods reduce bias below 5% and increase sample size by up to 14% under a fixed budget.
Fine-tuning Large Language Models with Limited Data: A Survey and Practical Guide (2026.tacl-1)

Copied to clipboard

Challenge: Pre-trained language models provide strong foundations, but effective adaptation under data scarcity requires efficient and efficient fine-tuning techniques.
Approach: They propose to review parameter-efficient fine-tuning techniques that lower training and deployment costs and domain and cross-lingual adaptation methods for both encoder and decoder models.
Outcome: The proposed techniques lower training and deployment costs, domain and cross-lingual adaptation methods, and model specialization strategies.
Survey Response Generation: Generating Closed-Ended Survey Responses In-Silico with Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing studies focus on generating closed-ended survey responses with large language models, whereas LLMs are typically trained to generate open-ended text.
Approach: They evaluate the impact of various Survey Response Generation Methods on simulated responses by generating closed-ended responses from large language models.
Outcome: The proposed methods perform best in individual-level and subpopulation-level alignment.
I Learn Better If You Speak My Language: Understanding the Superior Performance of Fine-Tuning Large Language Models with LLM-Generated Responses (2024.emnlp-main)

Copied to clipboard

Challenge: Recent research has demonstrated that a large language model (LLM) can generate training data for another LLM, or for creating supplementary training materials, such as rationales.
Approach: They conduct an in-depth investigation to understand why fine-tuning an LLM with responses generated by a LLM often yields better results than using responses generated from humans.
Outcome: The proposed approach can be used to transfer knowledge from a larger model to a smaller one, or for creating supplementary training materials, such as rationales.
Take the essence and discard the dross: A Rethinking on Data Selection for Fine-Tuning Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Existing studies focus on data selection but lack a clear, unified framework . variability in experimental settings complicates systematic comparisons .
Approach: They propose a three-stage scheme to standardize data selection for fine-tuning large language models . they propose unified comparison approach that incorporates ratio-based efficiency and ranking-based feasibility metrics to address inconsistencies across experiments.
Outcome: The proposed scheme outperforms existing methods in a dozen key studies and identifies key challenges.
Scaling Data Diversity for Fine-Tuning Language Models in Human Alignment (2024.lrec-main)

Copied to clipboard

Challenge: Large language models (LLMs) can reveal toxic or offensive content inadvertently or intentionally.
Approach: They propose to control the diversity of both sides according to the number of samples for fine-tuning, which can directly reflect their impact.
Outcome: The proposed approach improves the performance of large language models after fine-tuning.
Finetuning LLMs for Human Behavior Prediction in Social Science Experiments (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models can be used to simulate social science experiments . finetuning LLMs directly on individual-level responses from past experiments improves accuracy .
Approach: They propose to fine tune large language models directly on individual responses from past experiments to achieve multiple levels of generalization.
Outcome: The proposed model outperforms GPT-4o in completely unseen studies by 36% . the proposed model reduces demographic parity difference by 10.6% compared to GPT-4)
Distribution Shift Alignment Helps LLMs Simulate Survey Response Distributions (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods to simulate survey responses are based on zero-shot methods, but they are sensitive to prompt changes and deviate from the real-world distributions.
Approach: They propose a distribution shift alignment method that aligns both the output distributions and the distribution shifts across different backgrounds to provide results closer to the true distribution than the training data.
Outcome: The proposed method outperforms zero-shot methods on five public survey datasets and reduces the required real data by 53.48-69.12%.
Algorithmic Fidelity of Large Language Models in Generating Synthetic German Public Opinions: A Case Study (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models have generated significant interest in their potential for synthetic data generation across various domains.
Approach: They use open-ended survey data from the German Longitudinal Election Studies to prompt different LLMs to generate synthetic public opinions reflective of German subpopulations by incorporating demographic features into the persona prompts.
Outcome: The LLM performs better for supporters of left-leaning parties like The Greens and The Left compared to other parties, and matches the least with the right-party AfD.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations