Papers by Massimo Poesio
Conceptual Pacts for Reference Resolution Using Small, Dynamically Constructed Language Models: A Study in Puzzle Building Dialogues (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing large language models can be fine-tuned offline but are large and resource-intensive. |
| Approach: | They propose to use a simple reference resolver to simulate a conceptual pact process over time with different conversation pairs. |
| Outcome: | The proposed model performs better than a pre-trained model with exhaustive retraining after each prediction, while being more transparent, faster and less resource-intensive. |
The Universal Anaphora Scorer (2022.lrec-1)
Copied to clipboard
| Challenge: | a new version of the Reference Coreference Scorer is proposed to evaluate anaphoric interpretations . the proposed approach to evaluation of split antecedent anaphorisms is entirely novel . |
| Approach: | They propose an extended version of the Reference Coreference Scorer to evaluate anaphoric interpretations . the UA scorer supports the evaluation of split antecedent anaphorisms and discourse deixis . |
| Outcome: | The proposed method can be used to evaluate anaphoric interpretations in an extended range of anas . it supports evaluations of split antecedent anaphorisms and discourse deixis, for which no tools exist . |
Crowdsourcing and Aggregating Nested Markable Annotations (P19-1)
Copied to clipboard
| Challenge: | Existing methods for identifying markables for coreference annotation are task and language-independent and can be used for a variety of other annotation tasks. |
| Approach: | They propose a method for identifying markables for coreference annotation that combines automatic markable detectors with checking with a Game-With-A-Purpose (GWAP) and aggregation using a Bayesian annotation model. |
| Outcome: | The proposed method improves mention boundaries on news and other genres by over seven percentage points compared with state-of-the-art, domain-independent automatic mention detectors and almost three points over an in-domain mention detector. |
Beyond Black & White: Leveraging Annotator Disagreement via Soft-Label Multi-Task Learning (2021.naacl-main)
Copied to clipboard
| Challenge: | Prior work shows that disagreement between annotators can be useful in training models. |
| Approach: | They propose to use disagreements as an auxiliary task in a multi-task neural network to incorporate disagreements into models. |
| Outcome: | The proposed method significantly improves performance on NLP tasks beyond the standard approach and prior work. |
Low-Hallucination and Efficient Coreference Resolution with LLMs (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large Language Models have shown promising results in coreference resolution, but they face a critical issue: hallucinations. |
| Approach: | They propose a low-hallucination and efficient solution to the problem of hallucinations . they propose efficient constrained decoding for coreference resolution . |
| Outcome: | The proposed approach achieves better performance on the English OntoNotes development set. |
ArMIS - The Arabic Misogyny and Sexism Corpus with Annotator Subjective Disagreements (2022.lrec-1)
Copied to clipboard
| Challenge: | sexist and misogynistic language is increasingly used on social media in recent years . there are few benchmarks for misogorical annotations in Arabic . a dataset characterized by religious beliefs does not reconcile disagreements . |
| Approach: | They propose to use an Arabic misogyny and sexism dataset to analyze disagreements between religious annotators. |
| Outcome: | The proposed dataset shows that disagreements between annotators with different religious beliefs can be reconciled . |
Hard and Soft Evaluation of NLP models with BOOtSTrap SAmpling - BooStSa (2022.acl-demo)
Copied to clipboard
| Challenge: | Developing better methods for a task is a common feature of the computational linguistics literature. |
| Approach: | They propose to use bootstrap to compute significance levels with the BOOtSTrap SAmpling procedure to evaluate models that predict hard labels and soft labels as well. |
| Outcome: | The proposed method can be used to evaluate models that predict hard labels and soft labels on benchmark data sets. |
Stay Together: A System for Single and Split-antecedent Anaphora Resolution (2021.naacl-main)
Copied to clipboard
| Challenge: | Recent work on single-antecedent anaphora has greatly improved . attention has now turned to more complex cases of anaphorisms such as split-antevore anaprs . |
| Approach: | They propose a system that resolves both single and split-antecedent anaphors and evaluates it in a more realistic setting that uses predicted mentions. |
| Outcome: | The proposed system resolves both single and split-antecedent anaphora and evaluates it in a more realistic setting that uses predicted mentions. |
BERTective: Language Models and Contextual Information for Deception Detection (2021.eacl-main)
Copied to clipboard
| Challenge: | Existing methods to classify texts as truthful or deceptive are limited by the context of the text being analyzed. |
| Approach: | They propose to use a corpus of Italian dialogues to classify texts as truthful or deceptive. |
| Outcome: | The proposed models show that not all contexts are equally useful to the task. |
A Probabilistic Annotation Model for Crowdsourcing Coreference (D18-1)
Copied to clipboard
| Challenge: | Existing methods to generate annotated corpora for coreference are expensive and limited. |
| Approach: | They propose a model of annotation for aggregating crowdsourced anaphoric annotations. |
| Outcome: | The proposed model can extract from crowdsourced annotations coreference chains comparable to those obtained with expert annotation. |
Aggregating Crowdsourced and Automatic Judgments to Scale Up a Corpus of Anaphoric Reference for Fiction and Wikipedia Texts (2023.eacl-main)
Copied to clipboard
Juntao Yu, Silviu Paun, Maris Camilleri, Paloma Garcia, Jon Chamberlain, Udo Kruschwitz, Massimo Poesio
| Challenge: | Existing approaches to scale up anaphoric annotation have not overcome these limitations. |
| Approach: | They propose to use a game-with-a-purpose to ‘complete’ markable annotations by using an anaphoric resolver and an aggregation method for anaphorism. |
| Outcome: | The proposed method could be adopted to greatly speed up annotation time in other projects involving games-with-a-purpose. |
Universal Anaphora: The First Three Years (2024.lrec-main)
Copied to clipboard
Massimo Poesio, Maciej Ogrodniczuk, Vincent Ng, Sameer Pradhan, Juntao Yu, Nafise Sadat Moosavi, Silviu Paun, Amir Zeldes, Anna Nedoluzhko, Michal Novák, Martin Popel, Zdeněk Žabokrtský, Daniel Zeman
| Challenge: | Universal Anaphora initiative aims to push forward the state of the art in anaphora and anaphorism resolution by expanding the aspects of anaphonic interpretation which are or can be reliably annotated in an anagraphic corpora. |
| Approach: | They propose to develop a standard for anaphoric annotations and a method for evaluating models that can carry out this type of interpretation. |
| Outcome: | The Universal Anaphora initiative aims to push forward the state of the art in anaphora and anaphorism resolution by producing unified standards to annotate and encode annotations, delivering datasets encoded according to these standards, and developing methods for evaluating models that carry out this type of interpretation. |
Free the Plural: Unrestricted Split-Antecedent Anaphora Resolution (2020.coling-main)
Copied to clipboard
| Challenge: | a limitation of coreference resolution models is the focus on single-antecedent anaphors. |
| Approach: | They propose a model for unrestricted resolution of split-antecedent anaphors with multiple antecedents . they use auxiliary corpora where split-antcedent ananaphor was annotated by crowd . |
| Outcome: | The proposed model significantly improves on a baseline enhanced by BERT embeddings on anaphoric reference corpus. |
Assessing the Capabilities of Large Language Models in Coreference: An Evaluation (2024.lrec-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are a new approach to coreference resolution, but their performance is not yet fully understood. |
| Approach: | They propose that future efforts should improve scope, data, and evaluation methods of traditional coreference research to adapt to the development of LLMs. |
| Outcome: | The proposed methods improve scope, data, and evaluation methods of traditional coreference research to adapt to the development of LLMs. |
A Cluster Ranking Model for Full Anaphora Resolution (2020.lrec-1)
Copied to clipboard
| Challenge: | Anaphora resolution systems designed for CONLL 2012 dataset can handle key aspects of the full anaphora task such as the identification of singletons and of certain types of non-referring expressions. |
| Approach: | They propose an architecture to identify non-referring expressions and build coreference chains, including singletons, using system mentions. |
| Outcome: | The proposed model performs better on the CONLL 2012 dataset than the state-of-the-art system. |
Neural Mention Detection (2020.lrec-1)
Copied to clipboard
| Challenge: | Mention detection is an important preprocessing step for downstream applications such as NER and coreference resolution. |
| Approach: | They propose and compare three approaches to mention detection using ELMO embeddings and a biaffine classifier. |
| Outcome: | The proposed model outperforms state-of-the-art models on the GENIA corpora and improves on mention recall. |
Using Automatically Extracted Minimum Spans to Disentangle Coreference Evaluation from Boundary Detection (P19-1)
Copied to clipboard
| Challenge: | Existing methods to evaluate maximum spans tangle coreference evaluation with mention boundary detection . however, this method is costly and does not scale to large corpora. |
| Approach: | They propose an algorithm for automatically extracting minimum spans to benefit from minimum span evaluation in all corpora. |
| Outcome: | The proposed algorithm is consistent with those manually annotated by experts. |
Assessing Polyseme Sense Similarity through Co-predication Acceptability and Contextualised Embedding Distance (2020.starsem-1)
Copied to clipboard
| Challenge: | Co-predication is a commonly used linguistic test to tell apart shifts in polysemic sense from changes in homonymic meaning. |
| Approach: | They examine how co-predication acceptability relates to explicit ratings of polyseme word sense similarity and how well they can be predicted through the distance between target words’ contextualised word embeddings. |
| Outcome: | The proposed measures can be predicted through the distance between target words’ contextualised word embeddings. |
Patterns of Polysemy and Homonymy in Contextualised Language Models (2021.findings-emnlp)
Copied to clipboard
| Challenge: | a recent study has focused on homonymy, a variety of multiplicity of meanings exemplified by word forms with unrelated meanings. |
| Approach: | They investigate the extent to which contextualised embeddings reflect traditional distinctions of polysemy and homonymy. |
| Outcome: | The proposed model shows that it can distinguish between polysemy and homonymy . it shows that the model fails to replicate the results of the human-annotated dataset . |
Multitask Learning-Based Neural Bridging Reference Resolution (2020.coling-main)
Copied to clipboard
| Challenge: | Existing models for bridging references lack large corpora annotated with briding references . second challenge is different definitions of bridding used in different corpors . |
| Approach: | They propose a multi-task learning-based neural model for bridging reference resolution . they show substantial improvements of up to 8 p.p. on full briding resolution compared to previous models . |
| Outcome: | The proposed model outperforms the best reported results on full bridging resolution by up to 8 p.p. |
Polysemy through the lens of psycholinguistic variables: a dataset and an evaluation of static and contextualized language models (2024.starsem-1)
Copied to clipboard
| Challenge: | Polysemes are words that can have different senses depending on context . traditionally, NLP models assume that each sense should be given a separate representation in a lexicon, thus limiting the amount of evidence that can be gained from their use. |
| Approach: | They propose a framework to model polysemes as a continuous variation in psycholinguistic properties of a word in context without postulating jumps between senses. |
| Outcome: | The proposed framework accommodates different sense interpretations, without postulating clear-cut jumps between senses. |
Named Entity Recognition as Dependency Parsing (2020.acl-main)
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is a fundamental task in Natural Language Processing, concerned with identifying spans of text expressing references to entities. |
| Approach: | They propose a method to handle both types of NEs in one system by using a biaffine dependency parsing model which scores pairs of start and end tokens in a sentence. |
| Outcome: | The proposed model performs well on 8 corpora and achieves accuracy gains of up to 2.2 percentage points. |
Cross-lingual Zero Pronoun Resolution (2020.lrec-1)
Copied to clipboard
| Challenge: | In pronoun-dropping languages, predicate arguments are not realized instead of being realized as overt pronounos. |
| Approach: | They propose a BERT-based model for zero pronoun resolution in Arabic and Chinese . they also evaluate BERT feature extraction and fine-tune models on the task . |
| Outcome: | The proposed model outperforms the state-of-the-art model for Arabic and Chinese on OntoNotes 5.0. |
A Crowdsourced Corpus of Multiple Judgments and Disagreement on Anaphoric Interpretation (N19-1)
Copied to clipboard
| Challenge: | a corpus of anaphoric information (coreference) is crowdsourced through a game-with-a-purpose . its main feature is the large number of judgments per markable: 20 on average, and over 2.2M in total. |
| Approach: | They propose to crowdsource anaphoric information corpus by a game-with-a-purpose and to use it to train a coreference resolver. |
| Outcome: | The proposed corpus contains annotations for 108,000 markables and 20 judgments per markable, and 2.2M in total. |