Challenge: Existing long-term open-domain dialogue datasets lack complex, real-world personalization and fail to capture implicit reasoning.
Approach: They propose a large-scale long-term dataset with 2,500 examples containing approximately 100 conversation sessions to study implicit reasoning in personalized dialogues.
Outcome: The proposed model improves the ability of LLMs to reason over long-term conversations with implicit contextual dependencies.

Similar Papers

Beyond Goldfish Memory: Long-Term Open-Domain Conversation (2022.acl-long)

Copied to clipboard

Challenge: Despite recent improvements in open-domain dialogue models, state of the art models are trained and evaluated on short conversations with little context.
Approach: They propose to use retrieval-augmented methods to summarize and recall past conversations to improve their models.
Outcome: The proposed models outperform the current state-of-the-art models on human-human chat sessions in both automatic and human evaluations.
Evaluating Very Long-Term Conversational Memory of LLM Agents (2024.acl-long)

Copied to clipboard

Challenge: Existing studies on long-term open-domain dialogues focus on evaluating responses within contexts spanning no more than five chat sessions.
Approach: They propose a machine-human pipeline to generate very long-term dialogues by leveraging LLMs and retrieval augmented generation techniques.
Outcome: The proposed pipeline generates very long-term dialogues using LLMs and RAGs . the generated conversations are verified and edited by human annotators for long-range consistency and grounding to the event graphs.
HyperMem: Hypergraph Memory for Long-Term Conversations (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to long-term memory management rely on pairwise relations, causing fragmented retrieval.
Approach: They propose a hypergraph-based hierarchical memory architecture that explicitly models high-order associations using hyperedges.
Outcome: Experiments show that HyperMem achieves state-of-the-art performance with 92.73% accuracy for long-term conversations.
Generalizing Conversational Dense Retrieval via LLM-Cognition Data Augmentation (2024.acl-long)

Copied to clipboard

Challenge: Existing conversational dense retrieval models view a conversation as a fixed sequence of questions and responses, and these alternate conversations are unrecorded.
Approach: They propose a framework for generalizing Conversational dense retrieval via LLM-cognition data Augmentation (ConvAug) they first generate multi-level augmented conversations to capture the diverse nature of conversational contexts.
Outcome: The proposed framework generalizes Conversational dense retrieval via LLM-cognition data Augmentation on four public datasets.
A Persona-Aware LLM-Enhanced Framework for Multi-Session Personalized Dialogue Generation (2025.findings-acl)

Copied to clipboard

Challenge: Existing personalized dialogue models focus on dialogue history and personality information, reducing the responses’ consistency.
Approach: They propose a Persona-Aware LLM-enAnCEd(PALACE) framework that generates responses consistent with dialogue history and personality information across multiple sessions to engage users’ interest in the dialogue.
Outcome: The proposed framework outperforms the state-of-the-art methods in automatic and human evaluation metrics on the MSC and DuLeMon datasets.
Hello Again! LLM-powered Personalized Agent for Long-term Dialogue (2025.naacl-long)

Copied to clipboard

Challenge: Existing dialogue systems focus on brief single-session interactions, neglecting real-world needs for long-term companionship and personalized interactions.
Approach: They propose a model-agnostic framework for long-term dialogue agents . they use event summary and persona management to enable reasoning .
Outcome: The proposed framework incorporates three independently tunable modules dedicated to event perception, persona extraction, and response generation.
Eliciting Knowledge from Large Pre-Trained Models for Unsupervised Knowledge-Grounded Conversation (2022.emnlp-main)

Copied to clipboard

Challenge: Recent advances in large-scale pre-training provide large models with the potential to learn knowledge from the raw text.
Approach: They propose a posterior-based reweighing and noisy training strategy to exploit generated knowledge in dialogue generation.
Outcome: Empirical results show that the proposed methods outperform the state-of-the-art methods in unsupervised knowledge-grounded conversation.
HiGMem: A Hierarchical and LLM-Guided Memory System for Long-Term Conversational Agents (2026.findings-acl)

Copied to clipboard

Challenge: Existing memory systems rely on vector similarity for retrieval, resulting in bloated evidence sets . existing systems produce little additional recall, but this approach lowers retrieval precision .
Approach: They propose a two-level event-turn memory system that uses event summaries as semantic anchors to predict which related turns are worth reading.
Outcome: The proposed system achieves the best F1 on four of five question categories and improves adversarial F1 from 0.54 to 0.78 over A-Mem while retrieving an order of magnitude fewer turns.
Stark: Social Long-Term Multi-Modal Conversation with Persona Commonsense Knowledge (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies focus on image-sharing behavior in singular sessions, leading to limited long-term social interaction.
Approach: They propose a large-scale long-term multi-modal dialogue dataset that generates long-time multi-modity dialogue distilled from ChatGPT and proposed image aligner.
Outcome: The proposed framework generates long-term multi-modal dialogue from ChatGPT and image aligner.
Inside Out: Evolving User-Centric Core Memory Trees for Long-Term Personalized Dialogue Systems (2026.acl-long)

Copied to clipboard

Challenge: Existing personalized dialogue systems struggle to reconcile unbounded interactions with finite context constraints.
Approach: They propose a framework that utilizes a globally maintained PersonaTree as the carrier of long-term user profiling.
Outcome: The proposed framework outperforms existing systems in suppressing contextual noise and persona inconsistency.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations