Papers by Anuj Goyal

7 papers
Controlled Text Generation for Data Augmentation in Intelligent Artificial Agents (D19-56)

Copied to clipboard

Challenge: Data availability is a bottleneck during early stages of development of new capabilities for intelligent artificial agents.
Approach: They propose to use conditional variational auto-encoders to augment training data of a popular commercial artificial agent with a small set of phrase templates to generate new semantically similar phrases.
Outcome: The proposed approach outperforms the previous controlled text generation techniques with limited data and significantly outperformed the previous methods.
Semantic Span Annotation: An Exploratory Study of LLM Annotation (2026.acl-srw)

Copied to clipboard

Challenge: Structured span extraction research is siloed by context length, annotation task, and domain . Identifying a span within a natural language text and affixing it with a semantic label has been considered a core task in NLP .
Approach: They propose a framework for structured span annotation that integrates five datasets under a common JSONL format with character-level offsets.
Outcome: The proposed framework can generalize across four domains under three prompting configurations.
Toward More Accurate and Generalizable Evaluation Metrics for Task-Oriented Dialogs (2023.acl-industry)

Copied to clipboard

Challenge: Existing methods to dialog quality estimation focus on evaluating individual turns or collecting dialog-level quality measurements from end users immediately following an interaction.
Approach: They propose a dialog-level annotation workflow called Dialog Quality Annotation . they propose to annotate dialogs for attributes such as goal completion and user sentiment .
Outcome: The proposed model outperforms existing methods for dialog quality estimation . it shows that high-quality human-annotated data is important for dialog-quality evaluation .
Simple Question Answering with Subgraph Ranking and Joint-Scoring (N19-1)

Copied to clipboard

Challenge: Knowledge graph based simple question answering is a major area of research in question answering.
Approach: They propose a framework to describe and analyze existing knowledge graph based simple question answering approaches.
Outcome: The proposed model achieves a state-of-the-art (85.44% accuracy) on the SimpleQuestions dataset.
Alexa Conversations: An Extensible Data-driven Approach for Building Task-oriented Dialogue Systems (2021.naacl-demos)

Copied to clipboard

Challenge: Traditional goal-oriented dialogue systems require annotations which are hard to obtain for every new domain, limiting scalability.
Approach: They propose a data-driven approach to building goal-oriented dialogue systems . they use a seed dialogue simulator to generate annotated conversations instead of collecting annotations .
Outcome: The proposed system improves turn-level action signature prediction accuracy by 50% . the system is scalable, extensible and data efficient .
MultiWOZ 2.1: A Consolidated Multi-Domain Dialogue Dataset with State Corrections and State Tracking Baselines (2020.lrec-1)

Copied to clipboard

Challenge: MultiWOZ 2.0 has substantial noise in dialogue state annotations and dialogue utterances . follow-up work has augmented the original dataset with user dialogue acts .
Approach: They propose to reannotate dialogue state and utterances based on original dataset . they then compare their results to other datasets to improve their models .
Outcome: The proposed dataset improves on the noise in the dialogue state annotations and dialogue utterances.
OodGAN: Generative Adversarial Network for Out-of-Domain Data Generation (2021.naacl-industry)

Copied to clipboard

Challenge: Existing models for OOD detection work with text, but they do not work directly with the text.
Approach: They propose to use a sequential generative adversarial network (SeqGAN) based model to generate OOD data for a given domain automatically.
Outcome: The proposed model outperforms state-of-the-art in OOD detection metrics for ROSTD and OSQ datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations