Language Models Don’t Know What You Want: Evaluating Personalization in Deep Research Needs Real Users (2026.acl-long)
Copied to clipboard
Nishant Balepur, Malachi Hamada, Varsha Kishore, Sergey Feldman, Amanpreet Singh, Pao Siangliulue, Joseph Chee Chang, Eunsol Choi, Jordan Lee Boyd-Graber, Aakanksha Naik
| Challenge: | Earlier research used real users to push personalization, but easy-to-use judges have been criticized for not adopting online studies. |
| Approach: | They propose a personalized action-following tool that infers a user's research interests and proposes personalized actions for a query. |
| Outcome: | The proposed tool beats baselines in citation metrics and personalized action-following with an online version of MySQA. |
Similar Papers
Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis (2025.acl-long)
Copied to clipboard
| Challenge: | Personalized AI assistants are a challenging application that intertwines multiple problems in LLM research. |
| Approach: | They propose a Llama-3.2-based automated evaluation model that matches human preferences to a conversational dataset. |
| Outcome: | HiCUPID provides a conversational dataset tailored for personalization . the evaluation model closely mirrors human preferences, the researchers show . |
On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey (2024.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) provide a data-centric solution to alleviate limitations of real-world data with synthetic data generation. |
| Approach: | They propose a generic workflow for LLM-driven synthetic data generation. |
| Outcome: | The proposed workflows highlight gaps in existing research and outline avenues for future studies. |
Can LLM be a Personalized Judge? (2024.findings-emnlp)
Copied to clipboard
| Challenge: | a new study examines the reliability of large language models (LLMs) for personalization and role-playing evaluation without examining its validity. |
| Approach: | They investigate the reliability of LLM-as-a-Personalized-Judge for personalization . they find that personas provided to LLMs have limited predictive power . |
| Outcome: | The proposed model is less reliable than previously thought, the authors show . human annotation reveals that third-person crowd worker evaluations of personalized preferences are even worse than LLM predictions. |
Measuring What Makes You Unique: Difference-Aware User Modeling for Enhancing LLM Personalization (2025.findings-acl)
Copied to clipboard
| Challenge: | Extensive experiments on real-world datasets demonstrate that DPL significantly enhances LLM personalization. |
| Approach: | They propose a novel approach that emphasizes extracting inter-user differences to enhance LLM personalization. |
| Outcome: | The proposed approach extracts inter-user differences to enhance LLM personalization. |
Benchmarking and Improving LLM Robustness for Personalized Generation (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing evaluations focus on whether a model’s responses align with a user’s preferences, but factuality is an important yet overlooked dimension. |
| Approach: | They propose a scalable framework for evaluating robustness of large language models in personalization and a new dataset, PERGData. |
| Outcome: | The proposed framework improves robustness by 25% across models. |
Personalized Benchmarking: Evaluating LLMs by Individual Preferences (2026.findings-acl)
Copied to clipboard
| Challenge: | Current benchmarks average preferences across all users to compute aggregate ratings . this overlooks individual user preferences when establishing model rankings . |
| Approach: | They compute personalized model rankings using ELO ratings and Bradley-Terry coefficients . they find users exhibit substantial heterogeneity in topical interests and communication styles . |
| Outcome: | The results show that individual rankings of LLM models diverge dramatically from aggregate rankings . a compact combination of topic and style features provides a useful feature space . |
Evaluating Personalized Tool-Augmented LLMs from the Perspectives of Personalization and Proactivity (2025.acl-long)
Copied to clipboard
| Challenge: | Personalized tool utilization is essential for aligning large language models (LLMs) with user preference in interaction scenarios with various tools. |
| Approach: | They propose a key-point-based LLM evaluation method that mitigates biases by manually annotating key points for each test case and providing them to LLM as the reference. |
| Outcome: | The proposed method mitigates biases in the LLM-as-a-judge system by manually annotating key points for each test case and providing them to LLM as the reference. |
PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing personalization benchmarks focus on chit-chat, non-conversational tasks, or narrow domains, failing to capture complexities of personalized task-oriented assistance. |
| Approach: | They propose a benchmark to evaluate personalization in task-oriented AI assistants . the benchmark features user profiles equipped with rich preferences and interaction histories . |
| Outcome: | The proposed benchmark features user profiles equipped with rich preferences and interaction histories . it also features a judge agent and user agent that employs the LLM-as-a-Judge paradigm . |
Returning the N to NLP: Towards Contextually Personalized Classification Models (2020.acl-main)
Copied to clipboard
| Challenge: | a recent study shows that NLP models treat language as universal, but that it is based on sociolinguistic research. |
| Approach: | They propose to incorporate user-dependent, contextual personal and social aspects into neural NLP models by means of socially contextual personalization. |
| Outcome: | The proposed approach could be adapted to better personalize the language of users . it outlines a possible direction to incorporate these aspects into neural NLP models . |
When Personalization Meets Reality: A Multi-Faceted Analysis of Personalized Preference Learning (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Reinforcement Learning from Human Feedback assumes homogeneous preferences across users . personalization can introduce up to 20% safety misalignment . |
| Approach: | They propose a framework to assess personalized preference learning by tailoring preferences for users . they compare eight personalization methods across three preference datasets . |
| Outcome: | The proposed framework measures performance, fairness, unintended effects, adaptability across preferences . performance differences between personalization methods could reach 36% when users strongly disagree . |