Papers by Silviu Paun
The Universal Anaphora Scorer (2022.lrec-1)
Copied to clipboard
| Challenge: | a new version of the Reference Coreference Scorer is proposed to evaluate anaphoric interpretations . the proposed approach to evaluation of split antecedent anaphorisms is entirely novel . |
| Approach: | They propose an extended version of the Reference Coreference Scorer to evaluate anaphoric interpretations . the UA scorer supports the evaluation of split antecedent anaphorisms and discourse deixis . |
| Outcome: | The proposed method can be used to evaluate anaphoric interpretations in an extended range of anas . it supports evaluations of split antecedent anaphorisms and discourse deixis, for which no tools exist . |
Crowdsourcing and Aggregating Nested Markable Annotations (P19-1)
Copied to clipboard
| Challenge: | Existing methods for identifying markables for coreference annotation are task and language-independent and can be used for a variety of other annotation tasks. |
| Approach: | They propose a method for identifying markables for coreference annotation that combines automatic markable detectors with checking with a Game-With-A-Purpose (GWAP) and aggregation using a Bayesian annotation model. |
| Outcome: | The proposed method improves mention boundaries on news and other genres by over seven percentage points compared with state-of-the-art, domain-independent automatic mention detectors and almost three points over an in-domain mention detector. |
Beyond Black & White: Leveraging Annotator Disagreement via Soft-Label Multi-Task Learning (2021.naacl-main)
Copied to clipboard
| Challenge: | Prior work shows that disagreement between annotators can be useful in training models. |
| Approach: | They propose to use disagreements as an auxiliary task in a multi-task neural network to incorporate disagreements into models. |
| Outcome: | The proposed method significantly improves performance on NLP tasks beyond the standard approach and prior work. |
Stay Together: A System for Single and Split-antecedent Anaphora Resolution (2021.naacl-main)
Copied to clipboard
| Challenge: | Recent work on single-antecedent anaphora has greatly improved . attention has now turned to more complex cases of anaphorisms such as split-antevore anaprs . |
| Approach: | They propose a system that resolves both single and split-antecedent anaphors and evaluates it in a more realistic setting that uses predicted mentions. |
| Outcome: | The proposed system resolves both single and split-antecedent anaphora and evaluates it in a more realistic setting that uses predicted mentions. |
A Probabilistic Annotation Model for Crowdsourcing Coreference (D18-1)
Copied to clipboard
| Challenge: | Existing methods to generate annotated corpora for coreference are expensive and limited. |
| Approach: | They propose a model of annotation for aggregating crowdsourced anaphoric annotations. |
| Outcome: | The proposed model can extract from crowdsourced annotations coreference chains comparable to those obtained with expert annotation. |
Aggregating Crowdsourced and Automatic Judgments to Scale Up a Corpus of Anaphoric Reference for Fiction and Wikipedia Texts (2023.eacl-main)
Copied to clipboard
Juntao Yu, Silviu Paun, Maris Camilleri, Paloma Garcia, Jon Chamberlain, Udo Kruschwitz, Massimo Poesio
| Challenge: | Existing approaches to scale up anaphoric annotation have not overcome these limitations. |
| Approach: | They propose to use a game-with-a-purpose to ‘complete’ markable annotations by using an anaphoric resolver and an aggregation method for anaphorism. |
| Outcome: | The proposed method could be adopted to greatly speed up annotation time in other projects involving games-with-a-purpose. |
Universal Anaphora: The First Three Years (2024.lrec-main)
Copied to clipboard
Massimo Poesio, Maciej Ogrodniczuk, Vincent Ng, Sameer Pradhan, Juntao Yu, Nafise Sadat Moosavi, Silviu Paun, Amir Zeldes, Anna Nedoluzhko, Michal Novák, Martin Popel, Zdeněk Žabokrtský, Daniel Zeman
| Challenge: | Universal Anaphora initiative aims to push forward the state of the art in anaphora and anaphorism resolution by expanding the aspects of anaphonic interpretation which are or can be reliably annotated in an anagraphic corpora. |
| Approach: | They propose to develop a standard for anaphoric annotations and a method for evaluating models that can carry out this type of interpretation. |
| Outcome: | The Universal Anaphora initiative aims to push forward the state of the art in anaphora and anaphorism resolution by producing unified standards to annotate and encode annotations, delivering datasets encoded according to these standards, and developing methods for evaluating models that carry out this type of interpretation. |
Free the Plural: Unrestricted Split-Antecedent Anaphora Resolution (2020.coling-main)
Copied to clipboard
| Challenge: | a limitation of coreference resolution models is the focus on single-antecedent anaphors. |
| Approach: | They propose a model for unrestricted resolution of split-antecedent anaphors with multiple antecedents . they use auxiliary corpora where split-antcedent ananaphor was annotated by crowd . |
| Outcome: | The proposed model significantly improves on a baseline enhanced by BERT embeddings on anaphoric reference corpus. |
Aggregating and Learning from Multiple Annotators (2021.eacl-tutorials)
Copied to clipboard
| Challenge: | NLP is based on high-quality annotated datasets, but many other tasks have unique characteristics not considered by standard models of annotation. |
| Approach: | tutorial aims to connect NLP researchers with state-of-the-art aggregation models for canonical language annotation tasks. |
| Outcome: | This tutorial aims to connect NLP researchers with state-of-the-art aggregation models for a diverse set of canonical language annotation tasks. |
A Crowdsourced Corpus of Multiple Judgments and Disagreement on Anaphoric Interpretation (N19-1)
Copied to clipboard
| Challenge: | a corpus of anaphoric information (coreference) is crowdsourced through a game-with-a-purpose . its main feature is the large number of judgments per markable: 20 on average, and over 2.2M in total. |
| Approach: | They propose to crowdsource anaphoric information corpus by a game-with-a-purpose and to use it to train a coreference resolver. |
| Outcome: | The proposed corpus contains annotations for 108,000 markables and 20 judgments per markable, and 2.2M in total. |