Papers by Sameer Pradhan

8 papers
SPLICE: A Singleton-Enhanced PipeLIne for Coreference REsolution (2024.lrec-main)

Copied to clipboard

Challenge: Existing attempts to integrate singleton mention detection into end-to-end coreference resolution for English have been hampered by the lack of singletont mention spans in the OntoNotes benchmark.
Approach: They propose a two-step neural mention and coreference resolution system that integrates singleton mentions with OntoNotes syntax trees to achieve a near approximation of the Ontonotes dataset with all singletont mentions.
Outcome: The proposed system achieves 94% recall on a sample of gold singletons.
The Universal Anaphora Scorer (2022.lrec-1)

Copied to clipboard

Challenge: a new version of the Reference Coreference Scorer is proposed to evaluate anaphoric interpretations . the proposed approach to evaluation of split antecedent anaphorisms is entirely novel .
Approach: They propose an extended version of the Reference Coreference Scorer to evaluate anaphoric interpretations . the UA scorer supports the evaluation of split antecedent anaphorisms and discourse deixis .
Outcome: The proposed method can be used to evaluate anaphoric interpretations in an extended range of anas . it supports evaluations of split antecedent anaphorisms and discourse deixis, for which no tools exist .
My Science Tutor (MyST)–a Large Corpus of Children’s Conversational Speech (2024.lrec-main)

Copied to clipboard

Challenge: a 13-year project was conducted between 2007 and 2019 to improve students' learning proficiency in elementary school science using conversational multimedia virtual tutor, Marni.
Approach: They propose to use the corpus-name corpus to improve automatic speech recognition models and algorithms by training and developing a model on the training and development portion of the corpuse.
Outcome: The corpus comprises 400 hours of speech, spanning some 230K utterances spread across about 10,500 virtual tutor sessions.
PropBank Comes of Age—Larger, Smarter, and more Diverse (2022.starsem-1)

Copied to clipboard

Challenge: The PropBank has been used for semantic role labeling for over 20 years . it includes non-verbal predicates, adjectives, prepositions and multi-word expressions .
Approach: They describe the evolution of the PropBank approach to semantic role labeling over the last 20 years . they describe the substantial effort that has gone into ensuring consistency and reliability of the various annotated datasets and resources .
Outcome: The PropBank has been used for more than 20 years to test semantic role labeling systems.
Universal Anaphora: The First Three Years (2024.lrec-main)

Copied to clipboard

Challenge: Universal Anaphora initiative aims to push forward the state of the art in anaphora and anaphorism resolution by expanding the aspects of anaphonic interpretation which are or can be reliably annotated in an anagraphic corpora.
Approach: They propose to develop a standard for anaphoric annotations and a method for evaluating models that can carry out this type of interpretation.
Outcome: The Universal Anaphora initiative aims to push forward the state of the art in anaphora and anaphorism resolution by producing unified standards to annotate and encode annotations, delivering datasets encoded according to these standards, and developing methods for evaluating models that carry out this type of interpretation.
OntoGUM: Evaluating Contextualized SOTA Coreference Resolution on 12 More Genres (2021.acl-short)

Copied to clipboard

Challenge: Existing methods for coreference resolution are unable to evaluate generalizability to open domain data.
Approach: They propose to make an OntoNotes-like coreference dataset publicly available and convert it into an English corpus.
Outcome: The proposed dataset is the largest human-annotated coreference corpus following the OntoNotes guidelines and the first to be evaluated for consistency with the OnToNote's scheme.
Annotating Chinese Word Senses with English WordNet: A Practice on OntoNotes Chinese Sense Inventories (2024.lrec-main)

Copied to clipboard

Challenge: a recent study has shown that large language models can be useful for cross-lingual applications.
Approach: They propose to annotate Chinese word senses using English WordNet synsets . they examine the relationship between two annotators and find patterns among synset .
Outcome: The proposed method shows that the annotators agree on 38% of the synsets compared with the original synset . the results highlight similarities between the synnotated synset and the WordNet structure .
The New Propbank: Aligning Propbank with AMR through POS Unification (L18-1)

Copied to clipboard

Challenge: Existing Propbank corpus converts sense labels to a format which is more compatible with AMR and more robust to sparsity.
Approach: They propose a corpus which converts existing Propbank sense labels to a new unified format which is more compatible with AMR and more robust to sparsity.
Outcome: The proposed format is more compatible with AMR and robust to sparsity.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations