Papers by Srinivasan Iyer

12 papers
Do Explanations Help Users Detect Errors in Open-Domain QA? An Evaluation of Spoken vs. Visual Explanations (2021.findings-acl)

Copied to clipboard

Challenge: despite interest in explainable AI, there is increasing skepticism as to whether explanations are useful to end-users in downstream applications.
Approach: They conduct user studies to measure whether explanations help users decide when to accept or reject an ODQA system's answer.
Outcome: The proposed study shows that explanations outperform baselines across modalities but the best strategy varies with the modality.
Efficient One-Pass End-to-End Entity Linking for Questions (2020.emnlp-main)

Copied to clipboard

Challenge: Existing models for entity linking are limited to entity disambiguation and require mention boundaries to be given in the input.
Approach: They propose a fast end-to-end entity linking model that uses a biencoder to jointly detect mentions and link in one pass.
Outcome: The proposed model outperforms the current state of the art on WebQSP and GraphQuestions with extended annotations that cover multiple entities per question.
Complementary Explanations for Effective In-Context Learning (2023.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have remarkable capabilities in learning from expla- nations in prompts, but there has been limited understanding of exactly how these explana- tions function or why they are effective.
Approach: They propose a maximal marginal relevance-based exemplar selection approach to construct exemplar sets that are both relevant and comple- mentary.
Outcome: The proposed model improves in- context learning performance across three tasks on multiple LLMs.
Efficient Large Scale Language Modeling with Mixtures of Experts (2022.emnlp-main)

Copied to clipboard

Challenge: Mixture of Experts layers (MoEs) enable efficient scaling of language models . large autoregressive language models such as GPT-3 can be adapted to a wide range of tasks .
Approach: They propose to use Mixture of Experts layers to enable efficient scaling of language models . they find that MoEs are substantially more compute efficient than dense models compared to MoE models - but only when they are more modestly trained .
Outcome: The proposed model outperforms dense models in a wide range of tasks and domains.
RECONSIDER: Improved Re-Ranking using Span-Focused Cross-Attention for Open Domain Question Answering (2021.naacl-main)

Copied to clipboard

Challenge: State-of-the-art Machine Reading Comprehension (MRC) models for Open-domain Question Answering (QA) achieve high recall amongst top few predictions, but low overall accuracy, motivating the need for answer re-ranking.
Approach: They propose a method to make answer re-ranking successful for span-extraction tasks even beyond large pre-training.
Outcome: The proposed approach achieves 45.5% Exact Match accuracy on Natural Questions and 61.7% on TriviaQA.
JuICe: A Large Scale Distantly Supervised Dataset for Open Domain Context-based Code Generation (D19-1)

Copied to clipboard

Challenge: Interactive programming with interleaved code snippet cells and natural language markdown is gaining popularity in the form of Jupyter notebooks.
Approach: They propose to train code generation models based on a corpus of 1.5 million examples with a curated test set of 3.7K instances based off online programming assignments.
Outcome: The proposed model generates code cells based on the NL-Code history and human-curated data.
Methods for Measuring, Updating, and Visualizing Factual Beliefs in Language Models (2023.eacl-main)

Copied to clipboard

Challenge: Pretrained language models store a large amount of factual information that can be elicited by prompting or finetuning.
Approach: They propose methods to measure model factual beliefs and update incorrect beliefs in models . they propose a new visualization tool that shows relationships between stored model beliefs .
Outcome: The proposed methods improve models' consistency and accuracy, the authors show . their methods outperform existing methods in more difficult settings, the paper shows .
Learning to Map Context-Dependent Sentences to Executable Formal Queries (N18-1)

Copied to clipboard

Challenge: Existing models that map utterances to executable queries are context-dependent and can incorporate interaction history.
Approach: They propose a context-dependent model that maps utterances to executable queries . their approach combines implicit and explicit modeling of references between utterations .
Outcome: The proposed model can map utterances to executable queries based on interaction history . key to mapping utterrances to queries is resolving references .
Learning Programmatic Idioms for Scalable Semantic Parsing (D19-1)

Copied to clipboard

Challenge: In state-of-the-art semantic parsers map natural language instructions to source code . idioms improve the accuracy of semantic parses, allowing for faster decoding .
Approach: They propose an iterative method to extract code idioms from large source code corpora . they use most-frequent subtrees of their syntax trees to train semantic parsers to apply them .
Outcome: The proposed method improves the state-of-the-art semantic parsers' accuracy and training time by more than 50%.
Neural Semantic Parsing (P18-5)

Copied to clipboard

Challenge: Semantic parsing is the study of translating natural language utterances into machine-executable programs.
Approach: They will describe the various approaches researchers have taken to translate natural language into a formal language . they will also discuss why much recent work has chosen to use standard programming languages instead of more linguistically-motivated representations.
Outcome: This paper will describe the various approaches researchers have taken to translate natural language into a formal language.
ToKen: Task Decomposition and Knowledge Infusion for Few-Shot Hate Speech Detection (2022.emnlp-main)

Copied to clipboard

Challenge: Hate speech detection is complex and requires commonsense reasoning and social nuance . prior work has shown that even humans cannot achieve a high agreement on whether a post constitutes HS .
Approach: They frame a few-shot learning task to decompose a hate speech detection task into its "constituent" parts. they show that infusing commonsense knowledge from reasoning datasets improves the performance even further.
Outcome: The proposed method outperforms baseline methods in the 16-shot case.
Mapping Language to Code in Programmatic Context (D18-1)

Copied to clipboard

Challenge: Existing approaches for automatically mapping natural language to executable code have considered limited language or code environments.
Approach: They propose a task of generating class member functions given English documentation and the programmatic context provided by the rest of the class.
Outcome: The proposed model can generate member functions from documentation and the class environment.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations