Challenge: Existing methods for measuring identity fusion are limited and require controlled surveys or direct field contact.
Approach: They propose a new metric that integrates cognitive linguistics with large language models to measure identity fusion.
Outcome: The proposed metric outperforms existing methods and human annotations in violence risk assessment.

Similar Papers

Subtle Biases Need Subtler Measures: Dual Metrics for Evaluating Representative and Affinity Bias in Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Representative bias is a tendency of Large Language Models to generate outputs that mirror the experiences of certain identity groups, and affinity bias is an evaluative preference for specific narratives.
Approach: They propose two new metrics to measure representative bias and affinity bias within large language models and present a new set of tasks designed with customized rubrics to detect these biases.
Outcome: The proposed model identifies representative biases in prominent LLMs, with a preference for identities associated with being white, straight, and men.
Fusion-Eval: Integrating Assistant Evaluators with LLMs (2024.emnlp-industry)

Copied to clipboard

Challenge: Recent studies have employed large language models (LLMs) as reference-free metrics for NLG evaluation, enhancing adaptability to new tasks tasks.
Approach: They propose a method that leverages large language models to integrate insights from various assistant evaluators.
Outcome: The proposed approach achieves a 0.962 system-level Kendall-Tau correlation with humans on SummEval and a 0.7444 turn-level Spearman correlation on TopicalChat, which is significantly higher than baseline methods.
Value Portrait: Assessing Language Models’ Values through Psychometrically and Ecologically Valid Items (2025.acl-long)

Copied to clipboard

Challenge: Existing benchmarks rely on human annotations that are vulnerable to value-related biases.
Approach: They propose a value portrait benchmark that uses items that capture real-life user-LLM interactions and a rated item based on its similarity to their own thoughts to determine reliability.
Outcome: The proposed framework improves the relevance of assessment results to real-world LLM usage by allowing human subjects to rate items with similarity to their own thoughts and derived correlations between these ratings and the subjects’ actual value scores.
LLMs for Generating and Evaluating Counterfactuals: A Comprehensive Study (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown remarkable performance in NLP tasks, but their efficacy in generating high-quality CFs remains uncertain.
Approach: They compare LLMs' ability to generate CFs that flip the original label and human CF's.
Outcome: The proposed models generate fluent CFs, but struggle to keep the induced changes minimal.
Generative Psycho-Lexical Approach for Constructing Value Systems in Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have raised concerns regarding their intrinsic values.
Approach: They propose a psychologically grounded five-factor value system for Large Language Models that integrates psychological principles with cutting-edge AI priorities.
Outcome: The proposed value system meets standard psychological criteria, improves LLM safety prediction, and enhances Llm alignment, when compared to the canonical Schwartz’s values.
Value FULCRA: Mapping Large Language Models to the Multidimensional Spectrum of Basic Human Value (2024.naacl-long)

Copied to clipboard

Challenge: Existing work specifies values as risk criteria formulated in the AI community, e.g., fairness and privacy protection, suffering from poor clarity, adaptability and transparency.
Approach: They propose a value alignment paradigm based on Schwartz's Theory of Basic Values as an instantiation and propose 'BaseAlign' to support this paradigm.
Outcome: The proposed model covers existing risks and anticipates unidentified ones with a low-data set.
Systematic Evaluation of Auto-Encoding and Large Language Model Representations for Capturing Author States and Traits (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used in human-centered applications, yet their ability to model diverse psychological constructs is not well understood.
Approach: They evaluated a range of Transformer-LMs to predict psychological variables across five major dimensions: affect, substance use, mental health, sociodemographics, and personality.
Outcome: The models predict affect, substance use, mental health, sociodemographics, and personality across five major dimensions.
Measuring scalar constructs in social science with LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Valid scalar measurement of skalar constructs is a fundamental task in text analysis.
Approach: They evaluate four approaches to measuring scalar constructs using large language models . pairwise comparisons produced better measurements than prompting LLMs, they say . validation of skalar measurement enables wide range of substantive applications in social science research .
Outcome: The proposed methods improve on pairwise comparisons and finetuning . the proposed methods can be used in social science research .
Measuring Psychological Depth in Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Current evaluations of creative stories focus on objective properties of the text, such as its style, coherence, diversity, and creativity.
Approach: They propose a framework that measures an LLM's ability to produce authentic and narratively complex stories that provoke emotion, empathy, and engagement.
Outcome: The proposed framework shows that humans can consistently evaluate stories based on the PDS (0.72 Krippendorff’s alpha).
CogniVal in Action: An Interface for Customizable Cognitive Word Embedding Evaluation (2020.coling-demos)

Copied to clipboard

Challenge: Existing tools for evaluation of word embeddings are extrinsic and intrinsic methods, but they do not accurately reflect the meaning of words.
Approach: They present a command-line interface for CogniVal with multiple improvements over the original framework and the possibility to evaluate custom embeddings against custom cognitive data sources.
Outcome: The proposed system improves and extends the CogniVal framework and provides scalable and customized experiments.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations