The Personalization Trap: How User Memory Alters Emotional Reasoning in LLMs (2026.acl-short)
Copied to clipboard
| Challenge: | Using long-term memory, large language models can embed social hierarchies into their emotional reasoning. |
| Approach: | They evaluate 15 large language models on validated emotional intelligence tests to examine how user memory affects emotional intelligence. |
| Outcome: | The results show that the models with advantaged profiles receive more accurate emotional interpretations. |
Similar Papers
When Personalization Misleads: Understanding and Mitigating Hallucinations in Personalized LLMs (2026.findings-acl)
Copied to clipboard
| Challenge: | Personalization can inadvertently distort factual reasoning when faced with factual queries. |
| Approach: | They propose a lightweight inference-time approach that mitigates personalization-induced factual distortions while preserving personalized behavior. |
| Outcome: | Experiments across multiple LLM backbones and personalization methods show that FPPS significantly improves factual accuracy while maintaining personalized performance. |
PRIME: Large Language Model Personalization with Cognitive Dual-Memory and Personalized Thought Process (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language model (LLM) personalization aims to align outputs with individuals’ unique preferences and opinions. |
| Approach: | They integrate a cognitive dual-memory model into LLM personalization by mirroring episodic memory to historical user engagements and semantic memory to long-term, evolving user beliefs. |
| Outcome: | The proposed framework integrates the well-established cognitive dual-memory model into LLM personalization, using episodic and semanticmemories. |
Can LLM be a Personalized Judge? (2024.findings-emnlp)
Copied to clipboard
| Challenge: | a new study examines the reliability of large language models (LLMs) for personalization and role-playing evaluation without examining its validity. |
| Approach: | They investigate the reliability of LLM-as-a-Personalized-Judge for personalization . they find that personas provided to LLMs have limited predictive power . |
| Outcome: | The proposed model is less reliable than previously thought, the authors show . human annotation reveals that third-person crowd worker evaluations of personalized preferences are even worse than LLM predictions. |
Do Emotions Influence Moral Judgment in Large Language Models? (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent systems enforce explicit ethical constraints, but moral judgment rarely involves such clear-cut prohibitions. |
| Approach: | They develop an emotion-induction pipeline that infuses emotion into moral situations and evaluate shifts in moral acceptability across datasets and LLMs. |
| Outcome: | The proposed pipeline can infuses emotion into moral situations and evaluate moral acceptability shifts across datasets and LLMs. |
Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models can lead to undesired consequences when misaligned with human values . previous studies have shown misalignment of LLMs with human value using expert-designed or agent-based emulated bias scenarios . |
| Approach: | They investigate whether large language models (LLMs) are misaligned with human values . they find no significant differences in understanding of HVSB between LLMs . |
| Outcome: | The results show that large language models do not have lower misalignment rates and attack success rates . the study also shows that smaller language models have the ability to explain HVSB . |
Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis (2025.acl-long)
Copied to clipboard
| Challenge: | Personalized AI assistants are a challenging application that intertwines multiple problems in LLM research. |
| Approach: | They propose a Llama-3.2-based automated evaluation model that matches human preferences to a conversational dataset. |
| Outcome: | HiCUPID provides a conversational dataset tailored for personalization . the evaluation model closely mirrors human preferences, the researchers show . |
ReasoningRec: Bridging Personalized Recommendations and Human-Interpretable Explanations through LLM Reasoning (2025.findings-naacl)
Copied to clipboard
| Challenge: | Empirical evaluations demonstrate that ReasoningRec surpasses state-of-the-art methods by up to 12.5% in recommendation prediction while simultaneously providing human-intelligible explanations. |
| Approach: | They propose a reasoning-based recommendation framework that leverages Large Language Models to model users and items, focusing on preferences, aversions, and explanatory reasoning. |
| Outcome: | The proposed framework surpasses state-of-the-art methods by up to 12.5% in recommendation prediction while providing human-intelligible explanations. |
Memorization or Reasoning? Exploring the Idiom Understanding of LLMs (2025.emnlp-main)
Copied to clipboard
| Challenge: | idioms have long posed a challenge due to their unique linguistic properties, which set them apart from other common expressions. |
| Approach: | They propose to use a large-scale dataset of idioms in six languages to evaluate LLMs' idiomatic processing ability. |
| Outcome: | The proposed model integrates contextual cues and reasoning to improve idiom understanding in LLMs, suggesting that their performance is influenced by memorization and reasoning. |
Systematic Biases in LLM Simulations of Debates (2024.emnlp-main)
Copied to clipboard
| Challenge: | Current research suggests that LLM-based agents become increasingly human-like in their performance, sparking interest in using these AI agents as substitutes for human participants in behavioral studies. |
| Approach: | They propose to use LLMs to simulate political debates on topics that are important aspects of people’s day-to-day lives and decision-making processes. |
| Outcome: | The proposed model can simulate political debates on topics that are important aspects of people’s day-to-day lives and decision-making processes. |
“Going to a trap house” conveys more fear than “Going to a mall”: Benchmarking Emotion Context Sensitivity for LLMs (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a new benchmark evaluates whether large language models can understand emotion context sensitivity of humans. |
| Approach: | a new benchmark evaluates whether large language models can understand emotion context sensitivity of humans. |
| Outcome: | a new benchmark evaluates whether large language models can understand emotion context sensitivity of humans. |