Papers by Sonal Gupta

14 papers
Conversational Semantic Parsing (2020.emnlp-main)

Copied to clipboard

Challenge: Structured representations for task-oriented assistant systems are limited due to the limitations of the representation.
Approach: They propose a semantic representation for task-oriented conversational systems that can represent co-reference and context carryover.
Outcome: The proposed model improves the best results on ATIS, SNIPS, TOP and DSTC2 by up to 5 points for slot-carryover.
CCQA: A New Web-Scale Question Answering Dataset for Model Pre-Training (2022.findings-naacl)

Copied to clipboard

Challenge: Existing approaches to answer open domain questions rely on unlabeled text or synthetically generated question-answer pairs.
Approach: They propose a large-scale open-domain question-answering dataset based on the Common Crawl project that can be used to in-domain pre-train popular language models.
Outcome: The proposed dataset achieves promising results in zero-shot, low resource and fine-tuned settings across multiple tasks, models and benchmarks.
Low-Resource Domain Adaptation for Compositional Task-Oriented Semantic Parsing (2020.emnlp-main)

Copied to clipboard

Challenge: Recent advances in deep learning have enabled several approaches to successfully parse more complex queries, but these models require a large amount of annotated training data to parser on new domains (e.g. reminder, music).
Approach: They propose a method that adapts task-oriented semantic parsers to low-resource domains and outperforms a supervised neural model at a 10-fold data reduction.
Outcome: The proposed method outperforms baseline methods on a newly collected multi-domain task-oriented semantic parsing dataset (TOPv2) .
Muppet: Massive Multi-task Representations with Pre-Finetuning (2021.emnlp-main)

Copied to clipboard

Challenge: Recent work shows gains from pre-training and fine-tuning that are multi-task . but it can be difficult to know which intermediate tasks will best transfer .
Approach: They propose a large-scale learning stage for pre-finetuning between pre-training and fine-tun.
Outcome: The proposed model improves performance on pretrained discriminators and generation models on a wide range of tasks while improving sample efficiency during fine-tuning.
Span-based Hierarchical Semantic Parsing for Task-Oriented Dialog (D19-1)

Copied to clipboard

Challenge: Existing semantic parsers score intents and slots as labels of nesting nodes, but decode a valid tree globally.
Approach: They propose a span-based semantic parser for parsing compositional utterances into Task Oriented Parse (TOP) the parsers score labels of the tree nodes covering each token span independently, but decode a valid tree globally.
Outcome: The proposed parser outperforms previous methods on the TOP dataset in accuracy and training speed.
Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning (2021.acl-long)

Copied to clipboard

Challenge: Pre-trained language models can be fine-tuned to produce state-of-the-art results for a wide range of language understanding tasks.
Approach: They propose to analyze fine-tuning through the lens of intrinsic dimension . they show that pre-trained models have a low intrinsic dimension reparameterization .
Outcome: The proposed model can achieve 90% of the full parameter performance levels on MRPC with low data regime.
MTOP: A Comprehensive Multilingual Task-Oriented Semantic Parsing Benchmark (2021.eacl-main)

Copied to clipboard

Challenge: Existing datasets for task-oriented dialog systems are limited and expensive . current models are based on the simple intent and slot detection paradigm for non-compositional queries.
Approach: They propose to use a multilingual dataset to scale semantic parsing models to new languages . they demonstrate an average improvement of +6.3 points on Slot F1 for existing datasets .
Outcome: The proposed model achieves an average improvement of +6.3 points on Slot F1 over existing models.
Salient Phrase Aware Dense Retrieval: Can a Dense Retriever Imitate a Sparse One? (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing sparse retrievers lack the ability to match salient phrases and rare entities in the query.
Approach: They introduce a dense Lexical Model that can be trained to imitate a sparse one.
Outcome: The proposed model outperforms sparse retrievers on a range of tasks including five question answering datasets and the MS MARCO passage retrieval.
El Volumen Louder Por Favor: Code-switching in Task-oriented Semantic Parsing (2021.eacl-main)

Copied to clipboard

Challenge: Code-switching (CS) is the alternation of languages within an utterance or conversation.
Approach: They propose to use translation-and-align and augment with a generation model followed by match-and filter to improve CS generalizability of cross-lingual models when data for only one language is available.
Outcome: The proposed models improve when only English data is available alongside zero or a few CS training instances.
Semantic Parsing for Task Oriented Dialog using Hierarchical Representations (D18-1)

Copied to clipboard

Challenge: Existing work on task oriented dialog systems has limited expressive power to one intent per query and one slot label per token.
Approach: They propose a hierarchical annotation scheme for semantic parsing that allows representation of compositional queries.
Outcome: The proposed representation outperforms sequence-to-sequence approaches on a 44k annotated query dataset.
UniK-QA: Unified Representations of Structured and Unstructured Knowledge for Open-Domain Question Answering (2022.findings-naacl)

Copied to clipboard

Challenge: a recent study aims to answer factual questions using a structured knowledge base (KBQA).
Approach: They propose a unifying approach that homogenizes all knowledge sources by reducing them to text . they demonstrate that UniK-QA is a simple and yet effective way to combine heterogeneous sources of knowledge.
Outcome: The proposed approach improves state-of-the-art results on knowledge-base QA tasks by 11 points compared to graph-based methods.
Cross-lingual Transfer Learning for Multilingual Task Oriented Dialog (N19-1)

Copied to clipboard

Challenge: a lack of multilingual training data has hindered development of conversational AI models for task-oriented tasks . a new data set of 57k annotated utterances in english, spanish, and Thai is used to evaluate cross-lingual methods .
Approach: They present a data set of 57k annotated utterances in English, Spanish and Thai . they evaluate three different cross-lingual transfer methods to identify user intents and slots .
Outcome: The proposed model outperforms existing methods in English, Spanish and Thai . the proposed model is based on training data from three languages .
Sound Natural: Content Rephrasing in Dialog Systems (2020.emnlp-main)

Copied to clipboard

Challenge: Currently, virtual assistants work in the paradigm of intent-slot tagging and the slot values are directly passed as-is to the execution engine.
Approach: They propose to use BART to rephrase a query to make it more natural . they propose to add a copy-pointer and copy loss to it to improve performance .
Outcome: The proposed model improves on existing models by adding a copy-pointer and copy loss.
Domain-matched Pre-training Tasks for Dense Retrieval (2022.findings-naacl)

Copied to clipboard

Challenge: Existing approaches to improve performance of pre-training tasks are needed.
Approach: They propose to pre-train large bi-encoder models on a recently released set of 65 millionsynthetically generated questions and 200 million post-comment pairs from a preexisting reddit conversation dataset.
Outcome: The proposed model can be pre-trained on a set of 65 millionsynthetically generated questions and 200 million post-comment pairs from a preexisting dataset of Reddit conversations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations