Papers by Silviu Paun

10 papers
The Universal Anaphora Scorer (2022.lrec-1)

Copied to clipboard

Challenge: a new version of the Reference Coreference Scorer is proposed to evaluate anaphoric interpretations . the proposed approach to evaluation of split antecedent anaphorisms is entirely novel .
Approach: They propose an extended version of the Reference Coreference Scorer to evaluate anaphoric interpretations . the UA scorer supports the evaluation of split antecedent anaphorisms and discourse deixis .
Outcome: The proposed method can be used to evaluate anaphoric interpretations in an extended range of anas . it supports evaluations of split antecedent anaphorisms and discourse deixis, for which no tools exist .
Crowdsourcing and Aggregating Nested Markable Annotations (P19-1)

Copied to clipboard

Challenge: Existing methods for identifying markables for coreference annotation are task and language-independent and can be used for a variety of other annotation tasks.
Approach: They propose a method for identifying markables for coreference annotation that combines automatic markable detectors with checking with a Game-With-A-Purpose (GWAP) and aggregation using a Bayesian annotation model.
Outcome: The proposed method improves mention boundaries on news and other genres by over seven percentage points compared with state-of-the-art, domain-independent automatic mention detectors and almost three points over an in-domain mention detector.
Beyond Black & White: Leveraging Annotator Disagreement via Soft-Label Multi-Task Learning (2021.naacl-main)

Copied to clipboard

Challenge: Prior work shows that disagreement between annotators can be useful in training models.
Approach: They propose to use disagreements as an auxiliary task in a multi-task neural network to incorporate disagreements into models.
Outcome: The proposed method significantly improves performance on NLP tasks beyond the standard approach and prior work.
Stay Together: A System for Single and Split-antecedent Anaphora Resolution (2021.naacl-main)

Copied to clipboard

Challenge: Recent work on single-antecedent anaphora has greatly improved . attention has now turned to more complex cases of anaphorisms such as split-antevore anaprs .
Approach: They propose a system that resolves both single and split-antecedent anaphors and evaluates it in a more realistic setting that uses predicted mentions.
Outcome: The proposed system resolves both single and split-antecedent anaphora and evaluates it in a more realistic setting that uses predicted mentions.
A Probabilistic Annotation Model for Crowdsourcing Coreference (D18-1)

Copied to clipboard

Challenge: Existing methods to generate annotated corpora for coreference are expensive and limited.
Approach: They propose a model of annotation for aggregating crowdsourced anaphoric annotations.
Outcome: The proposed model can extract from crowdsourced annotations coreference chains comparable to those obtained with expert annotation.
Aggregating Crowdsourced and Automatic Judgments to Scale Up a Corpus of Anaphoric Reference for Fiction and Wikipedia Texts (2023.eacl-main)

Copied to clipboard

Challenge: Existing approaches to scale up anaphoric annotation have not overcome these limitations.
Approach: They propose to use a game-with-a-purpose to ‘complete’ markable annotations by using an anaphoric resolver and an aggregation method for anaphorism.
Outcome: The proposed method could be adopted to greatly speed up annotation time in other projects involving games-with-a-purpose.
Universal Anaphora: The First Three Years (2024.lrec-main)

Copied to clipboard

Challenge: Universal Anaphora initiative aims to push forward the state of the art in anaphora and anaphorism resolution by expanding the aspects of anaphonic interpretation which are or can be reliably annotated in an anagraphic corpora.
Approach: They propose to develop a standard for anaphoric annotations and a method for evaluating models that can carry out this type of interpretation.
Outcome: The Universal Anaphora initiative aims to push forward the state of the art in anaphora and anaphorism resolution by producing unified standards to annotate and encode annotations, delivering datasets encoded according to these standards, and developing methods for evaluating models that carry out this type of interpretation.
Free the Plural: Unrestricted Split-Antecedent Anaphora Resolution (2020.coling-main)

Copied to clipboard

Challenge: a limitation of coreference resolution models is the focus on single-antecedent anaphors.
Approach: They propose a model for unrestricted resolution of split-antecedent anaphors with multiple antecedents . they use auxiliary corpora where split-antcedent ananaphor was annotated by crowd .
Outcome: The proposed model significantly improves on a baseline enhanced by BERT embeddings on anaphoric reference corpus.
Aggregating and Learning from Multiple Annotators (2021.eacl-tutorials)

Copied to clipboard

Challenge: NLP is based on high-quality annotated datasets, but many other tasks have unique characteristics not considered by standard models of annotation.
Approach: tutorial aims to connect NLP researchers with state-of-the-art aggregation models for canonical language annotation tasks.
Outcome: This tutorial aims to connect NLP researchers with state-of-the-art aggregation models for a diverse set of canonical language annotation tasks.
A Crowdsourced Corpus of Multiple Judgments and Disagreement on Anaphoric Interpretation (N19-1)

Copied to clipboard

Challenge: a corpus of anaphoric information (coreference) is crowdsourced through a game-with-a-purpose . its main feature is the large number of judgments per markable: 20 on average, and over 2.2M in total.
Approach: They propose to crowdsource anaphoric information corpus by a game-with-a-purpose and to use it to train a coreference resolver.
Outcome: The proposed corpus contains annotations for 108,000 markables and 20 judgments per markable, and 2.2M in total.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations