HumBEL: A Human-in-the-Loop Approach for Evaluating Demographic Factors of Language Models in Human-Machine Conversations (2024.eacl-long)
Copied to clipboard
| Challenge: | Using demographic factors, pre-trained language models can adapt to demographic changes. |
| Approach: | They propose a framework to measure demographic alignment of language models with a target demographic for the first time. |
| Outcome: | The proposed framework outperforms human-machine language models in age-related tasks and outperformed a typical 21-year-old at memorization. |
Similar Papers
Can Demographic Factors Improve Text Classification? Revisiting Demographic Adaptation in the Age of Transformers (2023.findings-eacl)
Copied to clipboard
| Challenge: | Existing studies show that incorporating demographic factors in language representations improves performance on downstream NLP tasks. |
| Approach: | They use continuous language modeling and dynamic multi-task learning to adapt pre-trained Transformers to incorporate demographic information into their representations. |
| Outcome: | The proposed model shows that the results are consistent with previous studies. |
Your Mileage May Vary: How Empathy and Demographics Shape Human Preferences in LLM Responses (2025.findings-emnlp)
Copied to clipboard
| Challenge: | large language models (LLMs) increasingly assist subjective decision-making . prior work uses aggregate human judgments, but demographic variation and its linguistic drivers remain underexplored. |
| Approach: | They analyze how demographic background and empathy level correlate with LLM-generated dilemma responses . they also identify markers that predict group-level differences . |
| Outcome: | The authors show that demographic background and empathy level correlate with LLM preferences . their findings highlight the need for demographically informed LLM evaluations. |
Exploring the Reliability of Large Language Models as Customized Evaluators for Diverse NLP Tasks (2025.coling-main)
Copied to clipboard
| Challenge: | Existing work uses large language models (LLMs) to evaluate natural language process tasks, but there are shortcomings in current LLMs. |
| Approach: | They examine the alignment between LLM evaluators and human annotators by comparing conventional and alignment tasks with different evaluation criteria. |
| Outcome: | The proposed models excel in general criteria, such as fluency, but face challenges with complex criteria, including numerical reasoning. |
Missing the Margins: A Systematic Literature Review on the Demographic Representativeness of LLMs (2025.findings-acl)
Copied to clipboard
| Challenge: | 211 studies on the demographic representativeness of large language models have conflicting results . 29% of the studies report positive conclusions on the representativeness, 30% do not evaluate LLMs across multiple demographic categories or within demographic subcategories. |
| Approach: | 211 papers review the representativeness of large language models . authors recommend more precise evaluation methods and comprehensive documentation of demographic attributes . |
| Outcome: | 211 studies on the representativeness of large language models are reviewed . 29% of the studies report positive conclusions, but 30% fail to specify subcategories . authors recommend more precise evaluation methods and documentation of demographic attributes . |
Large Language Models for Psycholinguistic Plausibility Pretesting (2024.findings-eacl)
Copied to clipboard
| Challenge: | Psycholinguists typically use language models to create controlled materials . plausibility judgments are often based on coarse-grained judgements, but fine-grounded ones do not . |
| Approach: | They investigate whether Language Models can be used to generate plausibility judgments . they find that plausible judgements from LMs are highly related to human judgements - whereas other LM models are not . |
| Outcome: | The proposed language models can generate plausibility judgments from human evaluators . the proposed models do not provide satisfactory discriminative power . |
The Generation Gap: Exploring Age Bias in the Value Systems of Large Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Using the World Value Survey, we find a general inclination of LLM values towards younger demographics, especially when compared to the US population. |
| Approach: | They use data from the World Value Survey to examine the alignment of LLM values with specific age groups. |
| Outcome: | The proposed model can be used to predict the value of a large language model and to assess its performance on 13 categories. |
HumanRankEval: Automatic Evaluation of LMs as Conversational Assistants (2024.naacl-long)
Copied to clipboard
| Challenge: | Language models (LMs) are popular conversational assistants, but evaluation of such models is not scalable. |
| Approach: | They propose a task that performs automatic evaluation using human judgement and a large-scale set of questions with multiple answers authored and scored by humans. |
| Outcome: | The proposed task performs well with human judgements and is particularly responsive to model changes following instruction-tuning. |
Predict the Next Word: <Humans exhibit uncertainty in this task and language models _____> (2024.eacl-short)
Copied to clipboard
| Challenge: | Language models (LMs) are statistical models trained to assign probability to human-generated text. |
| Approach: | They evaluate language models' ability to reproduce variability that humans exhibit in the ‘next word prediction’ task. |
| Outcome: | The language models are trained to assign probability to human-generated text . they exhibit low calibration to human uncertainty, and advise against it . |
Aligning Language Models to User Opinions (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Personality is a defining feature of human beings, shaped by a complex interplay of demographic characteristics, moral principles, and social experiences. |
| Approach: | They use public opinion surveys to model past user opinions in addition to user demographics and ideology to achieve up to 7 points accuracy gains in predicting public opinions from survey questions. |
| Outcome: | The proposed model achieves 7 points accuracy gains in predicting public opinions from public opinion surveys across a broad set of topics. |
Unveiling Performance Challenges of Large Language Models in Low-Resource Healthcare: A Demographic Fairness Perspective (2025.coling-main)
Copied to clipboard
| Challenge: | Existing large language models (LLMs) are not effective in solving real-world healthcare tasks, but they are able to provide demographic information and provide biased health predictions. |
| Approach: | They evaluate state-of-the-art LLMs with three prevalent learning frameworks across six diverse healthcare tasks and find significant challenges in applying LLM to real-world healthcare tasks. |
| Outcome: | The proposed models perform poorly in real-world healthcare tasks and are inconsistent with existing learning frameworks. |