Challenge: Existing approaches to e-commerce relevance matching ignore bipartite graphs in logs . experimental results show that proposed method improves human relevance judgment .
Approach: They propose an efficient knowledge distillation framework for e-commerce relevance matching to exploit the advantages of Transformer-style and classical relevance matching models.
Outcome: The proposed method significantly improves human relevance judgment on large-scale real-world data.

Similar Papers

KD-Boost: Boosting Real-Time Semantic Matching in E-commerce with Knowledge Distillation (2023.emnlp-industry)

Copied to clipboard

Challenge: Existing SOTA techniques for semantic matching are mostly based on Siamese networks.
Approach: They propose a novel knowledge distillation algorithm designed for real-time semantic matching . they train low latency accurate student models by leveraging soft labels from a teacher model .
Outcome: The proposed algorithm outperforms teacher and SOTA knowledge distillation benchmarks on e-commerce datasets.
RTSM: Knowledge Distillation with Diverse Signals for Efficient Real-Time Semantic Matching in E-Commerce (2025.naacl-industry)

Copied to clipboard

Challenge: e-commerce product search is a key component of product discovery and sales in ecommerce . high computational demands of large transformer models pose challenges for their deployment in real-time scenarios.
Approach: They propose a framework for real-time semantic matching that leverages both soft labels from a teacher model and ground truth generated from pairwise query-product and query-query signals.
Outcome: The proposed framework outperforms teacher models and state-of-the-art models on e-commerce datasets.
Rationale-Guided Distillation for E-Commerce Relevance Classification: Bridging Large Language Models and Lightweight Cross-Encoders (2025.coling-industry)

Copied to clipboard

Challenge: Large-scale e-commerce search systems typically follow a multi-step process to retrieve relevant products for a given query.
Approach: They propose a distillation approach that uses "rationales" generated by Large Language Models to guide smaller cross-encoder models.
Outcome: The proposed model achieves ROC-AUC improvements of 1.4% on 9 multilingual e-commerce datasets, 2.4% on 3 ESCI datasets and 6% on GLUE datasets while being 50 times faster per sample.
CUPID: Curriculum Learning Based Real-Time Prediction using Distillation (2023.acl-industry)

Copied to clipboard

Challenge: Relevance in E-commerce Product Search is crucial for providing customers with accurate results that match their query intent.
Approach: They propose a curriculum learning based real-time relevance prediction using distillation . they propose e-commerce search systems that use transformers to predict relevance .
Outcome: The proposed model improves on english and Arabic in a bi-lingual relevance prediction task while maintaining low evaluation latency on CPUs.
Context-Aware Query Rewriting for Improving Users’ Search Experience on E-commerce Websites (2023.acl-industry)

Copied to clipboard

Challenge: Existing query rewriting models ignore user history behaviors and consider only the instant search query, which is often a short string offering limited information about the true shopping intent.
Approach: They propose an end-to-end context-aware query rewriting model that takes search context into account and builds a session graph using the history search queries and their contained words.
Outcome: The proposed model outperforms state-of-the-art models under various metrics.
Multi-stage Distillation Framework for Cross-Lingual Semantic Similarity Matching (2022.findings-naacl)

Copied to clipboard

Challenge: Existing studies have shown that cross-lingual knowledge distillation can improve the performance of pre-trained models for cross-linguistic similarity matching tasks.
Approach: They propose a multi-stage distillation framework for constructing a small-size but high-performance cross-lingual model using contrastive learning, bottleneck, and parameter recurrent strategies.
Outcome: The proposed model can compress the size of XLM-R and MiniLM by more than 50% while the performance is only reduced by about 1%.
Centrality-aware Product Retrieval and Ranking (2024.emnlp-industry)

Copied to clipboard

Challenge: Ambiguity and complexity of user queries often lead to mismatch between user’s intent and retrieved product titles or documents.
Approach: They propose a user-intent centrality optimization approach which optimizes for the user intent in semantic product search.
Outcome: The proposed approach improves product ranking efficiency for ambiguous queries and lexical terms with alphanumeric characters.
Neural Network based Extreme Classification and Similarity Models for Product Matching (N18-3)

Copied to clipboard

Challenge: Matching a seller listed item to an appropriate product has become a fundamental step for e-commerce platforms.
Approach: They propose to use a shallow neural network to match a seller's item to an appropriate product . they also propose a similarity approach based on deep siamese network to train and infer product information.
Outcome: The proposed models outperform the baseline models by more than 5% in terms of accuracy and are capable of efficient training and inference.
PairDistill: Pairwise Relevance Distillation for Dense Retrieval (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in dense retrieval have demonstrated remarkable efficacy compared to traditional sparse retrieval methods.
Approach: They propose to use pairwise relevance distillation to leverage pairwise reranking to enrich the training of dense retrieval models.
Outcome: The proposed method outperforms existing methods and achieves state-of-the-art results on multiple benchmarks.
Deep Metric Learning to Hierarchically Rank - An Application in Product Retrieval (2023.emnlp-industry)

Copied to clipboard

Challenge: e-commerce search engines use customer behavior signals to augment lexical matching and improve search relevance.
Approach: They propose a method to identify duplicate and near-duplicate products across stores . they use Hierarchical Ranked Multi Similarity Loss to learn hierarchical metric space .
Outcome: The proposed model outperforms baselines in terms of catalog coverage and precision of the mappings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations