Papers by Kaheer Suleman

13 papers
The KnowRef Coreference Corpus: Removing Gender and Number Cues for Difficult Pronominal Anaphora Resolution (P19-1)

Copied to clipboard

Challenge: Existing methods for coreference resolution exploit the number and gender of antecedents or have been handcrafted and do not reflect the diversity of naturally occurring text.
Approach: They propose a trick to improve resolution by antecedent switching to target common-sense understanding and world knowledge.
Outcome: The proposed method achieves state-of-the-art results on the GAP coreference task.
Deconstructing NLG Evaluation: Evaluation Practices, Assumptions, and Their Implications (2022.naacl-main)

Copied to clipboard

Challenge: Evaluating natural language generation systems is difficult, as there are many ways to express similar things in text.
Approach: They combine interviews with NLG practitioners to examine ethical considerations and their implications for NLG evaluation.
Outcome: The findings of the study surface goals, community practices, assumptions, and constraints that shape NLG evaluations, and examine their implications and how they embody ethical considerations.
A Knowledge Hunting Framework for Common Sense Reasoning (D18-1)

Copied to clipboard

Challenge: a new system that uses common sense to solve a common sense problem is developed . a winograd schema challenge and a choice of plausible alternatives are popular tests .
Approach: They propose an automatic system that achieves state-of-the-art results on the Winograd Schema Challenge . they use a knowledge hunting module to gather web text for problem resolutions .
Outcome: The proposed system achieves state-of-the-art on the Winograd Schema Challenge . it improves F1 performance on the full WSC by 0.21 over the previous best .
The KITMUS Test: Evaluating Knowledge Integration from Multiple Sources (2023.acl-long)

Copied to clipboard

Challenge: Existing models that make inferences using information from multiple sources are largely understudied .
Approach: They propose a test suite of coreference resolution subtasks that require reasoning over multiple facts and introduce subtask where knowledge is present only at inference time using fictional knowledge.
Outcome: The proposed subtasks differ in terms of which knowledge sources contain the relevant facts and where knowledge is present only at inference time using fictional knowledge.
A Generalized Knowledge Hunting Framework for the Winograd Schema Challenge (N18-4)

Copied to clipboard

Challenge: a new system that performs well on common-sense reasoning tasks is developed . the Winograd Schema Challenge (WSC) is a popular alternative to the Turing test .
Approach: They propose an automatic system that performs well on two common-sense reasoning tasks.
Outcome: The proposed system improves performance on the Winograd Schema Challenge and COPA by 0.16 over the previous best.
How Reasonable are Common-Sense Reasoning Tasks: A Case-Study on the Winograd Schema Challenge and SWAG (D19-1)

Copied to clipboard

Challenge: a recent study has improved the state-of-the-art on common-sense reasoning benchmarks . a san francisco-based approach to common-ense reasoning is challenging .
Approach: They propose to use common-sense reasoning benchmarks to test machine learning's common-sentence inference task SWAG to test common-mind systems.
Outcome: a new study shows that improved performance on common-sense reasoning benchmarks is genuine . the proposed task is more difficult than the current one, but it is more efficient than the previous one.
On the Systematicity of Probing Contextualized Word Representations: The Case of Hypernymy in BERT (2020.starsem-1)

Copied to clipboard

Challenge: Existing studies have found that BERT can correctly retrieve noun hypernyms in cloze tasks, but this does not correspond to systematic knowledge in BERT.
Approach: They propose to use BERT to probe for hypernymy knowledge encoded in representations for cloze tasks to find out whether it is systematic or not .
Outcome: The proposed model can retrieve hypernyms in cloze tasks, but not systematic knowledge in BERT.
Modeling Event Plausibility with Consistent Conceptual Abstraction (2021.naacl-main)

Copied to clipboard

Challenge: Understanding natural language requires common sense, one aspect of which is the ability to discern the plausibility of events.
Approach: They propose a method of forcing model consistency that improves correlation with human plausibility judgements.
Outcome: The proposed method improves correlation with human plausibility judgements.
TopiOCQA: Open-domain Conversational Question Answering with Topic Switching (2022.tacl-1)

Copied to clipboard

Challenge: Current datasets for conversational question answering do not contain topic switches . people often engage in information-seeking conversations to discover new knowledge .
Approach: They propose an open-domain conversational dataset with topic switches based on Wikipedia.
Outcome: The proposed dataset achieves an F1 of 55.8, falling short of human performance by 14.2 points, indicating the difficulty of the dataset.
An Analysis of Dataset Overlap on Winograd-Style Tasks (2020.coling-main)

Copied to clipboard

Challenge: a large number of test instances overlap considerably with pretraining corpora, a study finds . for a number of years, models struggled to exceed chance-level performance .
Approach: They analyze the effects of varying degrees of overlaps that occur between pretraining corpora and test instances in WSC-style tasks.
Outcome: The WSC-Web dataset is the largest to date and has lower overlaps with current pretraining corpora.
Challenges to Evaluating the Generalization of Coreference Resolution Models: A Measurement Modeling Perspective (2024.findings-acl)

Copied to clipboard

Challenge: a recent study shows that evaluations of CR models on multiple datasets conflate different factors concerning what is being measured.
Approach: They propose to view evaluations through the lens of measurement modeling . they show that evaluations risk conflating different factors concerning what is being measured .
Outcome: The evaluations on seven datasets show that models that reflect coreference generalization are often correlated with differences in how coreference is defined and operationalized.
ADEPT: An Adjective-Dependent Plausibility Task (2021.acl-long)

Copied to clipboard

Challenge: ADEPT is a large-scale semantic plausibility task that requires a significant degree of world knowledge and common-sense reasoning.
Approach: They propose a large-scale semantic plausibility task that pairs 16 thousand sentences with slightly modified versions obtained by adding an adjective to a noun.
Outcome: The proposed task is easier for humans (85% accuracy), but more difficult for transformer-based models (71% accuracy).
Can a Gorilla Ride a Camel? Learning Semantic Plausibility from Text (D19-60)

Copied to clipboard

Challenge: Existing work on modeling semantic plausibility has focused on physical plausability but distributional methods fail when tested in supervised settings.
Approach: They propose to use large pretrained language models to model plausibility in supervised settings by extracting attested events from a large corpus and injecting explicit commonsense knowledge into a distributional model.
Outcome: The proposed model is effective in modeling plausibility in a supervised setting.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations