Papers by Anoop Saladi

11 papers
SMART: Scalable Multilingual Approach for a Robust TOD System (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing TOD frameworks face significant challenges in handling unstructured information, providing multilingual support, and engaging proactively.
Approach: They propose a novel TOD framework that combines traditional pipeline elements with modern agent-based approaches and features a simplified dialogue state, intelligent clarification mechanisms, and a unified natural language generation component that eliminates response redundancy.
Outcome: The proposed framework outperforms baseline systems across key metrics and integrates in an e-commerce store.
AutoEval-ToD: Automated Evaluation of Task-oriented Dialog Systems (2025.naacl-long)

Copied to clipboard

Challenge: Current evaluation methodologies heavily depend on human annotators, which can be inefficient, subjective, and expensive to scale.
Approach: They propose an automated end-to-end evaluation framework that interacts with the ToD system and then assesses its performance across key dimensions.
Outcome: The proposed framework first interacts with the ToD system and assesses its performance across key dimensions by analyzing both its responses and internal states.
ASK: Aspects and Retrieval based Hybrid Clarification in Task Oriented Dialogue Systems (2025.acl-industry)

Copied to clipboard

Challenge: Ambiguous user queries pose a challenge in task-oriented dialogue systems . Large Language Models (LLMs) rely on the top-k retrieved documents for clarification . traditional approaches lack principled mechanisms to determine when to use broad domain knowledge vs specific retrieved document context for clarification.
Approach: They propose a hybrid approach that dynamically chooses between document-based or aspect-based clarification based on query ambiguity.
Outcome: The proposed approach shows significant improvements over baselines on product troubleshooting and product search datasets.
MIRAGE: Metadata-guided Image Retrieval and Answer Generation for E-commerce Troubleshooting (2026.eacl-industry)

Copied to clipboard

Challenge: Existing multimodal systems often associate text and images based on embedding similarity or simple co-location, but fail to ensure that the linked image accurately depicts the specific product or component mentioned in a troubleshooting instruction.
Approach: They propose a metadata-first paradigm that treats structured metadata as a modality for multimodal grounding.
Outcome: The proposed model uses a semantic schema to capture product attributes and visual aspects.
VADE: Visual Attention Guided Hallucination Detection and Elimination (2025.findings-acl)

Copied to clipboard

Challenge: Vision Language Models (VLMs) are prone to hallucinations, generating outputs that lack grounding in the actual visual data.
Approach: They propose a sequence modelling approach to learn complex sequential patterns from transformer attention maps.
Outcome: The proposed approach achieves an average PR-AUC of 80% in hallucination detection on M-HalDetect and an 5% improvement in hallucinosis mitigation on MSCOCO.
AutoChunker: Structured Text Chunking and its Evaluation (2025.acl-industry)

Copied to clipboard

Challenge: Existing methods for text chunking struggle with document structure and noise . Existing approaches struggle with maintaining semantic coherence while handling complex documents.
Approach: They propose a bottom-up approach to chunking that combines document structure awareness with noise elimination.
Outcome: The proposed method outperforms existing methods in noise reduction, completeness, context coherence, task relevance, and retrieval performance.
An Address Intelligence Framework for E-commerce Deliveries (2025.emnlp-industry)

Copied to clipboard

Challenge: a physical address is an important touchpoint between an e-commerce domain and its customers . incomplete or incorrect addresses can prevent delivery problems and improve the overall customer delivery experience.
Approach: They propose a language model to assist customers withaddress standardization and a Pareto-ensemble multi-task prediction algorithm that derives critical insights from customer addresses to minimize operational losses.
Outcome: The proposed system can minimize operational losses in an e-commerce domain.
AutoKB: Automated Creation of Structured Knowledge Bases for Domain-Specific Support (2025.naacl-industry)

Copied to clipboard

Challenge: Effective customer support requires domain-specific solutions tailored to users’ issues.
Approach: They propose an automated pipeline for building a domain-specific KB with a hierarchical tree structure that maps user issues to precise and domain-compliant solutions.
Outcome: Experiments in troubleshooting and medical domains show that the proposed pipeline outperforms LLMs and unstructured knowledge bases and is 75 times more cost-effective than manual methods.
MTIVE: Multi-Task Image Verification Engine Using Vision-Language Models for E-commerce (2026.acl-industry)

Copied to clipboard

Challenge: Vision-language models struggle with noisy real-world images and multi-task requirements.
Approach: They propose a curriculum learning framework that adapts vision-language models through three stages . MTIVE uses frozen base weights with stacked LoRA adapters for shared domain knowledge .
Outcome: MTIVE outperforms open-source and proprietary baselines in standard and continual learning settings.
GeoGround: Uncertainty-Weighted Multi-Task Learning for Geo-Alignment and Address Defect Detection (2026.acl-industry)

Copied to clipboard

Challenge: Address intelligence in e-commerce requires precise geocoding and proactive defect detection under strict sub-50 ms latency constraints.
Approach: They propose a multi-task learning framework that jointly models coordinate grounding and address defect detection.
Outcome: The proposed model achieves 5.86 gains in address defect detection precision and 4.86 improvements in location prediction accuracy over strong encoder baselines while remaining 75 more efficient than decoder LLMs such as Qwen2-1.5B.
VIT-Pro: Visual Instruction Tuning for Product Images (2025.naacl-industry)

Copied to clipboard

Challenge: general-purpose vision-language models struggle to understand and converse about real-world e-commerce product images.
Approach: a new approach is proposed to use large-scale image-text pairs to train a generative VLM for e-commerce product images.
Outcome: The proposed model outperforms general-purpose VLMs on multiple vision tasks in the e-commerce domain.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations