Challenge: Using demographic factors, pre-trained language models can adapt to demographic changes.
Approach: They propose a framework to measure demographic alignment of language models with a target demographic for the first time.
Outcome: The proposed framework outperforms human-machine language models in age-related tasks and outperformed a typical 21-year-old at memorization.

Similar Papers

Can Demographic Factors Improve Text Classification? Revisiting Demographic Adaptation in the Age of Transformers (2023.findings-eacl)

Copied to clipboard

Challenge: Existing studies show that incorporating demographic factors in language representations improves performance on downstream NLP tasks.
Approach: They use continuous language modeling and dynamic multi-task learning to adapt pre-trained Transformers to incorporate demographic information into their representations.
Outcome: The proposed model shows that the results are consistent with previous studies.
Your Mileage May Vary: How Empathy and Demographics Shape Human Preferences in LLM Responses (2025.findings-emnlp)

Copied to clipboard

Challenge: large language models (LLMs) increasingly assist subjective decision-making . prior work uses aggregate human judgments, but demographic variation and its linguistic drivers remain underexplored.
Approach: They analyze how demographic background and empathy level correlate with LLM-generated dilemma responses . they also identify markers that predict group-level differences .
Outcome: The authors show that demographic background and empathy level correlate with LLM preferences . their findings highlight the need for demographically informed LLM evaluations.
Exploring the Reliability of Large Language Models as Customized Evaluators for Diverse NLP Tasks (2025.coling-main)

Copied to clipboard

Challenge: Existing work uses large language models (LLMs) to evaluate natural language process tasks, but there are shortcomings in current LLMs.
Approach: They examine the alignment between LLM evaluators and human annotators by comparing conventional and alignment tasks with different evaluation criteria.
Outcome: The proposed models excel in general criteria, such as fluency, but face challenges with complex criteria, including numerical reasoning.
Missing the Margins: A Systematic Literature Review on the Demographic Representativeness of LLMs (2025.findings-acl)

Copied to clipboard

Challenge: 211 studies on the demographic representativeness of large language models have conflicting results . 29% of the studies report positive conclusions on the representativeness, 30% do not evaluate LLMs across multiple demographic categories or within demographic subcategories.
Approach: 211 papers review the representativeness of large language models . authors recommend more precise evaluation methods and comprehensive documentation of demographic attributes .
Outcome: 211 studies on the representativeness of large language models are reviewed . 29% of the studies report positive conclusions, but 30% fail to specify subcategories . authors recommend more precise evaluation methods and documentation of demographic attributes .
Large Language Models for Psycholinguistic Plausibility Pretesting (2024.findings-eacl)

Copied to clipboard

Challenge: Psycholinguists typically use language models to create controlled materials . plausibility judgments are often based on coarse-grained judgements, but fine-grounded ones do not .
Approach: They investigate whether Language Models can be used to generate plausibility judgments . they find that plausible judgements from LMs are highly related to human judgements - whereas other LM models are not .
Outcome: The proposed language models can generate plausibility judgments from human evaluators . the proposed models do not provide satisfactory discriminative power .
The Generation Gap: Exploring Age Bias in the Value Systems of Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Using the World Value Survey, we find a general inclination of LLM values towards younger demographics, especially when compared to the US population.
Approach: They use data from the World Value Survey to examine the alignment of LLM values with specific age groups.
Outcome: The proposed model can be used to predict the value of a large language model and to assess its performance on 13 categories.
HumanRankEval: Automatic Evaluation of LMs as Conversational Assistants (2024.naacl-long)

Copied to clipboard

Challenge: Language models (LMs) are popular conversational assistants, but evaluation of such models is not scalable.
Approach: They propose a task that performs automatic evaluation using human judgement and a large-scale set of questions with multiple answers authored and scored by humans.
Outcome: The proposed task performs well with human judgements and is particularly responsive to model changes following instruction-tuning.
Predict the Next Word: <Humans exhibit uncertainty in this task and language models _____> (2024.eacl-short)

Copied to clipboard

Challenge: Language models (LMs) are statistical models trained to assign probability to human-generated text.
Approach: They evaluate language models' ability to reproduce variability that humans exhibit in the ‘next word prediction’ task.
Outcome: The language models are trained to assign probability to human-generated text . they exhibit low calibration to human uncertainty, and advise against it .
Aligning Language Models to User Opinions (2023.findings-emnlp)

Copied to clipboard

Challenge: Personality is a defining feature of human beings, shaped by a complex interplay of demographic characteristics, moral principles, and social experiences.
Approach: They use public opinion surveys to model past user opinions in addition to user demographics and ideology to achieve up to 7 points accuracy gains in predicting public opinions from survey questions.
Outcome: The proposed model achieves 7 points accuracy gains in predicting public opinions from public opinion surveys across a broad set of topics.
Unveiling Performance Challenges of Large Language Models in Low-Resource Healthcare: A Demographic Fairness Perspective (2025.coling-main)

Copied to clipboard

Challenge: Existing large language models (LLMs) are not effective in solving real-world healthcare tasks, but they are able to provide demographic information and provide biased health predictions.
Approach: They evaluate state-of-the-art LLMs with three prevalent learning frameworks across six diverse healthcare tasks and find significant challenges in applying LLM to real-world healthcare tasks.
Outcome: The proposed models perform poorly in real-world healthcare tasks and are inconsistent with existing learning frameworks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations