Challenge: Earlier research used real users to push personalization, but easy-to-use judges have been criticized for not adopting online studies.
Approach: They propose a personalized action-following tool that infers a user's research interests and proposes personalized actions for a query.
Outcome: The proposed tool beats baselines in citation metrics and personalized action-following with an online version of MySQA.

Similar Papers

Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis (2025.acl-long)

Copied to clipboard

Challenge: Personalized AI assistants are a challenging application that intertwines multiple problems in LLM research.
Approach: They propose a Llama-3.2-based automated evaluation model that matches human preferences to a conversational dataset.
Outcome: HiCUPID provides a conversational dataset tailored for personalization . the evaluation model closely mirrors human preferences, the researchers show .
On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) provide a data-centric solution to alleviate limitations of real-world data with synthetic data generation.
Approach: They propose a generic workflow for LLM-driven synthetic data generation.
Outcome: The proposed workflows highlight gaps in existing research and outline avenues for future studies.
Can LLM be a Personalized Judge? (2024.findings-emnlp)

Copied to clipboard

Challenge: a new study examines the reliability of large language models (LLMs) for personalization and role-playing evaluation without examining its validity.
Approach: They investigate the reliability of LLM-as-a-Personalized-Judge for personalization . they find that personas provided to LLMs have limited predictive power .
Outcome: The proposed model is less reliable than previously thought, the authors show . human annotation reveals that third-person crowd worker evaluations of personalized preferences are even worse than LLM predictions.
Measuring What Makes You Unique: Difference-Aware User Modeling for Enhancing LLM Personalization (2025.findings-acl)

Copied to clipboard

Challenge: Extensive experiments on real-world datasets demonstrate that DPL significantly enhances LLM personalization.
Approach: They propose a novel approach that emphasizes extracting inter-user differences to enhance LLM personalization.
Outcome: The proposed approach extracts inter-user differences to enhance LLM personalization.
Benchmarking and Improving LLM Robustness for Personalized Generation (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluations focus on whether a model’s responses align with a user’s preferences, but factuality is an important yet overlooked dimension.
Approach: They propose a scalable framework for evaluating robustness of large language models in personalization and a new dataset, PERGData.
Outcome: The proposed framework improves robustness by 25% across models.
Personalized Benchmarking: Evaluating LLMs by Individual Preferences (2026.findings-acl)

Copied to clipboard

Challenge: Current benchmarks average preferences across all users to compute aggregate ratings . this overlooks individual user preferences when establishing model rankings .
Approach: They compute personalized model rankings using ELO ratings and Bradley-Terry coefficients . they find users exhibit substantial heterogeneity in topical interests and communication styles .
Outcome: The results show that individual rankings of LLM models diverge dramatically from aggregate rankings . a compact combination of topic and style features provides a useful feature space .
Evaluating Personalized Tool-Augmented LLMs from the Perspectives of Personalization and Proactivity (2025.acl-long)

Copied to clipboard

Challenge: Personalized tool utilization is essential for aligning large language models (LLMs) with user preference in interaction scenarios with various tools.
Approach: They propose a key-point-based LLM evaluation method that mitigates biases by manually annotating key points for each test case and providing them to LLM as the reference.
Outcome: The proposed method mitigates biases in the LLM-as-a-judge system by manually annotating key points for each test case and providing them to LLM as the reference.
PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants (2025.findings-acl)

Copied to clipboard

Challenge: Existing personalization benchmarks focus on chit-chat, non-conversational tasks, or narrow domains, failing to capture complexities of personalized task-oriented assistance.
Approach: They propose a benchmark to evaluate personalization in task-oriented AI assistants . the benchmark features user profiles equipped with rich preferences and interaction histories .
Outcome: The proposed benchmark features user profiles equipped with rich preferences and interaction histories . it also features a judge agent and user agent that employs the LLM-as-a-Judge paradigm .
Returning the N to NLP: Towards Contextually Personalized Classification Models (2020.acl-main)

Copied to clipboard

Challenge: a recent study shows that NLP models treat language as universal, but that it is based on sociolinguistic research.
Approach: They propose to incorporate user-dependent, contextual personal and social aspects into neural NLP models by means of socially contextual personalization.
Outcome: The proposed approach could be adapted to better personalize the language of users . it outlines a possible direction to incorporate these aspects into neural NLP models .
When Personalization Meets Reality: A Multi-Faceted Analysis of Personalized Preference Learning (2025.findings-emnlp)

Copied to clipboard

Challenge: Reinforcement Learning from Human Feedback assumes homogeneous preferences across users . personalization can introduce up to 20% safety misalignment .
Approach: They propose a framework to assess personalized preference learning by tailoring preferences for users . they compare eight personalization methods across three preference datasets .
Outcome: The proposed framework measures performance, fairness, unintended effects, adaptability across preferences . performance differences between personalization methods could reach 36% when users strongly disagree .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations