Papers by Manoj Kumar

5 papers
Unsupervised training data re-weighting for natural language understanding with local distribution approximation (2022.emnlp-industry)

Copied to clipboard

Challenge: a distribution mismatch between offline training and live data can cause biases . cyclic seasonality shifts, and changing pool of users can contribute to this problem .
Approach: They propose an unsupervised approach to mitigate offline training data sampling bias . they propose a local distribution approximation in the pre-trained embedding space .
Outcome: The proposed approach mitigates the offline training data sampling bias in multiple NLU tasks without additional annotation.
On Localizing and Deleting Toxic Memories in Large Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Existing methods to reduce toxic generation in large language models are not fully understood.
Approach: They propose to understand the mechanisms that drive toxic generation in large language models by using memory localization to reduce toxic generation.
Outcome: The proposed method reduces toxic generation from 62.86% to 28.61%, but it also improves generation quality.
Controlled Data Generation via Insertion Operations for NLU (2022.naacl-industry)

Copied to clipboard

Challenge: a new approach to annotate live traffic is emerging to be cost-effective and efficient . manual data annotation is expensive and not preferred for meeting customer privacy expectations .
Approach: They propose a targeted synthetic data generation technique by inserting tokens into a given semantic signature.
Outcome: The proposed approach achieves the same accuracy as training with all available data on a voice assistant dataset.
Improving Large-Scale Conversational Assistants using Model Interpretation based Training Sample Selection (2022.emnlp-industry)

Copied to clipboard

Challenge: Large-scale, voice-based conversational assistants process each utterance through a multi-stage pipeline that includes wakeword detection, automatic speech recognition (ASR), natural language understanding (NLU), entity resolution, and textto-speech.
Approach: They propose a method to identify customer implicitly satisfied with Alexa's responses by leveraging interpretations of model behavior.
Outcome: The proposed approach produces statistically significant improvements in both offline and online tests.
Chasing the Tail with Domain Generalization: A Case Study on Frequency-Enriched Datasets (2022.aacl-main)

Copied to clipboard

Challenge: In academic research, natural language understanding tasks are typically defined by creating annotated datasets in which each utterance is encountered once.
Approach: They propose a method that explicitly uses utterance frequency in training data to learn models that are more robust to unknown distributions.
Outcome: The proposed approach shows up to 7.02% relative improvement over baselines on the tail data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations