PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing personalization benchmarks focus on chit-chat, non-conversational tasks, or narrow domains, failing to capture complexities of personalized task-oriented assistance. |
| Approach: | They propose a benchmark to evaluate personalization in task-oriented AI assistants . the benchmark features user profiles equipped with rich preferences and interaction histories . |
| Outcome: | The proposed benchmark features user profiles equipped with rich preferences and interaction histories . it also features a judge agent and user agent that employs the LLM-as-a-Judge paradigm . |
Similar Papers
Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis (2025.acl-long)
Copied to clipboard
| Challenge: | Personalized AI assistants are a challenging application that intertwines multiple problems in LLM research. |
| Approach: | They propose a Llama-3.2-based automated evaluation model that matches human preferences to a conversational dataset. |
| Outcome: | HiCUPID provides a conversational dataset tailored for personalization . the evaluation model closely mirrors human preferences, the researchers show . |
Personalized Benchmarking: Evaluating LLMs by Individual Preferences (2026.findings-acl)
Copied to clipboard
| Challenge: | Current benchmarks average preferences across all users to compute aggregate ratings . this overlooks individual user preferences when establishing model rankings . |
| Approach: | They compute personalized model rankings using ELO ratings and Bradley-Terry coefficients . they find users exhibit substantial heterogeneity in topical interests and communication styles . |
| Outcome: | The results show that individual rankings of LLM models diverge dramatically from aggregate rankings . a compact combination of topic and style features provides a useful feature space . |
PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits (2024.findings-naacl)
Copied to clipboard
| Challenge: | Recent studies have shown that LLMs can generate content that aligns with their assigned personality traits, but there is limited research on whether they consistently reflect specific personality traits. |
| Approach: | They propose to study the behavior of LLM-based agents which they refer to as LLM personas and simulate them to measure their personality traits. |
| Outcome: | The proposed model is based on the Big Five personality model and has been validated by human evaluations and automatic evaluations. |
Can LLM be a Personalized Judge? (2024.findings-emnlp)
Copied to clipboard
| Challenge: | a new study examines the reliability of large language models (LLMs) for personalization and role-playing evaluation without examining its validity. |
| Approach: | They investigate the reliability of LLM-as-a-Personalized-Judge for personalization . they find that personas provided to LLMs have limited predictive power . |
| Outcome: | The proposed model is less reliable than previously thought, the authors show . human annotation reveals that third-person crowd worker evaluations of personalized preferences are even worse than LLM predictions. |
Recent Trends in Personalized Dialogue Generation: A Review of Datasets, Methodologies, and Evaluations (2024.lrec-main)
Copied to clipboard
| Challenge: | Personalization is a multifaceted process that requires multiple definitions and varies between individuals. |
| Approach: | They propose to systemically survey the recent landscape of personalized dialogue generation including the datasets employed, methodologies developed, and evaluation metrics applied. |
| Outcome: | The proposed model can generate fluent and coherent responses to human queries in a language-based conversational agent. |
Characteristic AI Agents via Large Language Models (2024.lrec-main)
Copied to clipboard
| Challenge: | Commercial products have been devoted to creating character-driven chatbots using large language models, but academic research in this area remains relatively scarce. |
| Approach: | They investigate the performance of LLMs in constructing characteristic AI agents by simulating real-life individuals across different settings. |
| Outcome: | The proposed benchmark compared LLMs with real-life individuals in different settings and includes evaluation metrics. |
HumanRankEval: Automatic Evaluation of LMs as Conversational Assistants (2024.naacl-long)
Copied to clipboard
| Challenge: | Language models (LMs) are popular conversational assistants, but evaluation of such models is not scalable. |
| Approach: | They propose a task that performs automatic evaluation using human judgement and a large-scale set of questions with multiple answers authored and scored by humans. |
| Outcome: | The proposed task performs well with human judgements and is particularly responsive to model changes following instruction-tuning. |
LaMP: When Large Language Models Meet Personalization (2024.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks for personalization in large language models are understudied . |
| Approach: | They propose a benchmark for training and evaluating language models for producing personalized outputs using a set of seven personalized tasks . they propose two retrieval augmentation approaches that retrieve personal items from each user profile for personalizing language model outputs. |
| Outcome: | The proposed approach is effective for a set of zero-shot and fine-tuned language models and highlights the impact of personalization in various natural language tasks. |
AgentSense: Benchmarking Social Intelligence of Language Agents through Interactive Scenarios (2025.naacl-long)
Copied to clipboard
Xinyi Mou, Jingcong Liang, Jiayu Lin, Xinnong Zhang, Xiawei Liu, Shiyue Yang, Rong Ye, Lei Chen, Haoyu Kuang, Xuanjing Huang, Zhongyu Wei
| Challenge: | Large language models are increasingly employed to empower autonomous agents to simulate human behavior. |
| Approach: | They propose to evaluate LLM-driven agents through multi-turn interactions using a bottom-up approach to create diverse social scenarios constructed from extensive scripts. |
| Outcome: | The proposed model evaluates LLM-driven agents through multi-turn interactions emphasizing goal completion and implicit reasoning. |
Benchmarking and Improving LLM Robustness for Personalized Generation (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing evaluations focus on whether a model’s responses align with a user’s preferences, but factuality is an important yet overlooked dimension. |
| Approach: | They propose a scalable framework for evaluating robustness of large language models in personalization and a new dataset, PERGData. |
| Outcome: | The proposed framework improves robustness by 25% across models. |