Challenge: e-commerce platforms are encountering increasingly complex product categorization scenarios . multiple business domains correspond to different category taxonomies, with different depths and distinct literal expressions of category names.
Approach: They propose a taxonomy-agnostic framework that calculates semantic relatedness between product titles and category names in the vector space.
Outcome: The proposed framework outperforms strong baselineson three dynamic multi-domain product categorization tasks.

Similar Papers

Consistent Text Categorization using Data Augmentation in e-Commerce (2023.acl-industry)

Copied to clipboard

Challenge: Upon closer inspection, we found inconsistencies in the labeling of similar items.
Approach: They propose to improve an existing product categorization model that takes a product title as input and outputs the most suitable category out of thousands of available candidates.
Outcome: The proposed model is based on a product title and outputs the most suitable category out of thousands of available candidates.
TXtract: Taxonomy-Aware Knowledge Extraction for Thousands of Product Categories (2020.acl-main)

Copied to clipboard

Challenge: State-of-the-art methods for knowledge extraction are designed for a single category of product, but do not apply to real-life e-Commerce scenarios.
Approach: They propose a taxonomy-aware knowledge extraction model that applies to thousands of categories organized in a hierarchical taxonomies.
Outcome: The proposed model outperforms state-of-the-art methods on 4,000 categories in F1 and 15% across all categories.
Label-semantics Aware Generative Approach for Domain-Agnostic Multilabel Classification (2025.findings-acl)

Copied to clipboard

Challenge: Existing approaches to multi-label text classification are limited by textual data.
Approach: They propose a domain-agnostic generative model framework for multi-label text classification that generates predefined label descriptions and matches them to predefined labels.
Outcome: The proposed model achieves 13.94% and 24.85% performance over all datasets.
TEMP: Taxonomy Expansion with Dynamic Margin Loss through Taxonomy-Paths (2021.emnlp-main)

Copied to clipboard

Challenge: Existing taxonomies are unable to maintain coverage due to the rising of new concepts . TEMP uses pre-trained contextual encoders to predict the position of new ideas .
Approach: They propose a self-supervised taxonomy expansion method that ranks taxonomies by ranking them . they use pre-trained contextual encoders to train the model with dynamic margin loss .
Outcome: The proposed method outperforms state-of-the-art taxonomy expansion methods by 14.3% and 15.8% on public benchmarks.
FABRIC: Fully-Automated Broad Intent Categorization in E-commerce (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing query classification models have excellent predictive performance on single-intent queries, but there is little research on predicting multiple-intentions for broad queries.
Approach: They propose to combine user click data, query-item relevance and LLM judgments to create an automatic method for multi-label e-commerce query classification.
Outcome: The proposed method reduces the ambiguity of the annotations by blending the label assessment from three different sources: user click data, query-item relevance and LLM judgments.
Leveraging Taxonomy and LLMs for Improved Multimodal Hierarchical Classification (2025.coling-main)

Copied to clipboard

Challenge: Multi-level Hierarchical Classification (MLHC) is a critical tool in modern data analysis.
Approach: They propose a taxonomy-embedded transitional LLM-agnostic framework for multimodality classification that leverages large language models to enforce consistency across hierarchical levels.
Outcome: The proposed framework improves on the MEP-3M dataset with various hierarchical levels compared to conventional models.
Multi-Domain Named Entity Recognition with Genre-Aware and Agnostic Inference (2020.acl-main)

Copied to clipboard

Challenge: Named entity recognition (NER) is a key component of many text processing pipelines.
Approach: They propose a new architecture tailored to the task of identifying named entities with data from multiple genres.
Outcome: The proposed architecture outperforms baseline and competitive methods on all three setups with differences ranging between +1.95 to +3.11 average F1 across multiple genres when compared to standard approaches.
Can Large Language Models Serve as Effective Classifiers for Hierarchical Multi-Label Classification of Scientific Documents at Industrial Scale? (2025.coling-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated great potential in complex tasks such as multi-label classification, but the vast number of labels can exceed LLMs’ input limits.
Approach: They propose a method that integrates large language models with dense retrieval techniques to overcome these challenges.
Outcome: The proposed methods avoid frequent retraining by leveraging zero-shot and few-shot learning for real-time label assignment.
Centrality-aware Product Retrieval and Ranking (2024.emnlp-industry)

Copied to clipboard

Challenge: Ambiguity and complexity of user queries often lead to mismatch between user’s intent and retrieved product titles or documents.
Approach: They propose a user-intent centrality optimization approach which optimizes for the user intent in semantic product search.
Outcome: The proposed approach improves product ranking efficiency for ambiguous queries and lexical terms with alphanumeric characters.
Semantic-conditioned Dual Adaptation for Cross-domain Query-based Visual Segmentation (2023.findings-acl)

Copied to clipboard

Challenge: Existing approaches to visual segmentation from language queries require expensive labeling and degradation when deployed to an unseen domain.
Approach: They propose a task to adapt a visual segmentation model from a labeled domain to an unseen domain.
Outcome: The proposed framework achieves precise feature- and relation-invariant across domains via universal semantic structure.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations