Challenge: Existing models fail to capture and model customer intention effectively because of insufficient information exploitation and only apparent information like descriptions and titles are used.
Approach: They propose to exploit existing session data to capture and model intention in E-commerce product purchase sessions using a multimodal benchmark.
Outcome: The proposed framework can bridge the gap between intention understanding in simplified research cases like co-buy intention and more complex yet practical scenarios like session history.

Similar Papers

EcomScriptBench: A Multi-task Benchmark for E-commerce Script Planning via Step-wise Intention-Driven Product Association (2025.acl-long)

Copied to clipboard

Challenge: Goal-oriented script planning is used by humans to plan for typical activities . however, this capability remains underexplored due to several challenges .
Approach: They propose a framework that enables product-enriched scripts by associating products with each step based on the semantic similarity between the actions and their purchase intentions.
Outcome: The proposed framework can generate product-enriched scripts from 2.4 million scripts . human annotations are conducted to provide gold labels for a sampled subset .
Intention Knowledge Graph Construction for User Intention Relation Modeling (2026.eacl-long)

Copied to clipboard

Challenge: Existing knowledge graphs focus on connecting intentions but lacks the ability to model the relationships between different intentions.
Approach: They propose a framework to automatically generate an intention knowledge graph, capturing connections between user intentions.
Outcome: The proposed model outperforms state-of-the-art methods and shows its utility.
IntentionQA: A Benchmark for Evaluating Purchase Intention Comprehension Abilities of Language Models in E-commerce (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches that distill intentions from LMs fail to generate meaningful and human-centric intentions applicable in real-world E-commerce contexts.
Approach: They propose a double-task multiple-choice question answering benchmark to evaluate LMs' comprehension of purchase intentions in E-commerce.
Outcome: The proposed benchmark consists of 4,360 carefully curated problems across three difficulty levels, constructed using an automated pipeline to ensure scalability on large E-commerce platforms.
Multi-Graph Co-Training for Capturing User Intent in Session-based Recommendation (2025.coling-main)

Copied to clipboard

Challenge: Existing methods rely on user actions within the current session, overlooking the wealth of auxiliary information available.
Approach: They propose a session-based recommendation model that leverages the current session graph and similar session graphs to capture the intrinsic relationships between items.
Outcome: The proposed model improves on the Diginetica dataset by 2.00% and 10.70% respectively.
ShopperBench: A Benchmark for Personalized Shopping with Persona-Guided Simulation (2026.eacl-industry)

Copied to clipboard

Challenge: Existing evaluation frameworks lack mechanisms to assess Personalized shopping agents' ability to adapt their strategies to heterogeneous user preferences and decisionmaking patterns.
Approach: They propose a persona-guided benchmark that augments shopping trajectories with personas . they propose persona Fidelity, Persona-Query Alignment, and Path Consistency .
Outcome: The proposed benchmark captures how shopper types navigate product search and selection . it measures persona Fidelity, Persona-Query Alignment, and Path Consistency .
Beyond IVR: Benchmarking Customer Support LLM Agents for Business-Adherence (2026.eacl-industry)

Copied to clipboard

Challenge: Existing benchmarks focus on tool usage or task completion, overlooking an agent’s capacity to adhere to multi-step policies, navigate task dependencies, and remain robust to unpredictable user or environment behavior.
Approach: They propose a benchmark to assess policy-aware agents in customer support using a dynamic-prompt agent and a static-promped agent that explicitly models policy control.
Outcome: The proposed benchmark assesses agent's ability to adhere to multi-step policies, navigate task dependencies, and remain robust to unpredictable user or environment behavior.
Dynabench: Rethinking Benchmarking in NLP (2021.naacl-main)

Copied to clipboard

Challenge: Dynabench is an open-source platform for dynamic dataset creation and model benchmarking.
Approach: They propose an open-source platform for dynamic dataset creation and model benchmarking.
Outcome: The proposed platform can be used to create models that fail on simple challenges and falter in real-world scenarios.
DeepResearch Retail: Benchmarking Tool-Augmented Deep Research in the E-Commerce Domain (2026.acl-industry)

Copied to clipboard

Challenge: Existing DR systems are largely web-centric and do not incorporate structured, domain-specific, and personalized information accessible through internal API tools.
Approach: They propose a framework grounded in real-world e-commerce data for assessing Deep Research with tools in realistic commercial settings.
Outcome: The proposed framework evaluates factual faithfulness and multidimensional response quality when reasoning over heterogeneous web and internal data sources.
Is Your Language Model Ready for Monetization Decisions? (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks focus on shopping-centric scenarios and user-facing data, overlooking intermediate decision stages and robustness considerations.
Approach: They propose a multi-task benchmark to evaluate large language models in real-world monetization contexts.
Outcome: The proposed benchmark covers intent understanding, commercial matching, and user behavior modeling.
Leveraging Product Catalog Patterns for Multilingual E-commerce Product Attribute Prediction (2025.emnlp-industry)

Copied to clipboard

Challenge: E-commerce stores increasingly use Large Language Models to improve catalog data quality . a critical challenge is accurately predicting missing structured attribute values .
Approach: They propose a retrieval-augmented system that leverages existing product catalog entries to guide LLM predictions for missing attributes.
Outcome: The proposed system improves catalog data quality by 34% and accuracy by 0.8% . the proposed model can predict missing attributes in multilingual product catalogs .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations