Papers by Piek Vossen
Measuring the Diversity of Automatic Image Descriptions (C18-1)
Copied to clipboard
| Challenge: | a lack of diversity in automatic image description systems is a general problem in natural language generation . authors use established metrics to evaluate system performance on the head of the vocabulary . automatic image descriptions are difficult because of the unbounded range of variation in natural languages . |
| Approach: | They propose to frame automatic image description as a word recall task to quantify the production of generic sentences as 'undiversity' they propose to use established metrics to evaluate the diversity of the output . |
| Outcome: | The proposed metrics evaluate the diversity of sentences generated by state-of-the-art systems on a MS COCO dataset. |
A Shared Task of a New, Collaborative Type to Foster Reproducibility: A First Exercise in the Area of Language Science and Technology with REPROLANG2020 (2020.lrec-1)
Copied to clipboard
António Branco, Nicoletta Calzolari, Piek Vossen, Gertjan Van Noord, Dieter van Uytvanck, João Silva, Luís Gomes, André Moreira, Willem Elbers
| Challenge: | Scientific knowledge is grounded on falsifiable predictions and therefore its credibility and raison d'être rely on the possibility of repeating experiments and getting similar results as originally obtained and reported. |
| Approach: | They propose a collaborative task which is collaborative rather than competitive and supports reproduction of research results. |
| Outcome: | The proposed task is called REPROLANG-The Shared Task on the Reproduction of Research Results in Science and Technology of Natural Language Processing (LREC2020). |
Do Differences in Values Influence Disagreements in Online Discussions? (2023.emnlp-main)
Copied to clipboard
| Challenge: | Disagreement is an important aspect of online discussions since it can drive novel ideas, incentivize evaluation of the proposed ideas, and avoid echo chambers. |
| Approach: | They propose to use human-annotated agreement labels to estimate personal values and to include value information in agreement prediction to improve performance. |
| Outcome: | The proposed models show that dissimilarity of value profiles correlates with disagreement in specific cases and that including value information in agreement prediction improves performance. |
Language Models Lack Temporal Generalization and Bigger is Not Better (2025.findings-acl)
Copied to clipboard
| Challenge: | 450 encoder models are fine-tuned on 15 data splits on a task to detect events in Early Modern Dutch archival texts. |
| Approach: | They propose to fine tune six encoder models that have been pretrained with very different data on a task in Early Modern Dutch archival texts. |
| Outcome: | The proposed model is fine tuned with 5 seeds on 15 different data splits and reaches highest F1 performance. |
Annotating Perspectives on Vaccination (2020.lrec-1)
Copied to clipboard
| Challenge: | Vaccination corpus is a corpus of texts related to the online vaccination debate . it contains documents from the Internet which reflect different views on vaccinations . |
| Approach: | They present a corpus of texts related to the online vaccination debate annotated with perspectives about attribution, claims and opinions. |
| Outcome: | The Vaccination Corpus contains 294 documents from the Internet which reflect different views on vaccinations. |
Reasoning about Ambiguous Definite Descriptions (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing resources to evaluate reasoning are not well suited to investigate the capability of resolving ambiguities by explicit reasoning. |
| Approach: | They propose to use ambiguous definite descriptions to create a benchmark dataset which requires models to resolve ambiguity by explicit reasoning. |
| Outcome: | The proposed model includes all information required to resolve the ambiguity in the prompt, which means a model does not require anything but reasoning to do well. |
Scoring and Classifying Implicit Positive Interpretations: A Challenge of Class Imbalance (C18-1)
Copied to clipboard
| Challenge: | a reimplementation of a system on detecting implicit positive meaning from negated statements is reported . a baseline taking the mean score or most frequent class is hard to beat because of class imbalance in the dataset. |
| Approach: | They propose a system to detect implicit positive meaning from negated statements . they convert the scores into classes and report their results on regression and classification tasks . |
| Outcome: | The proposed system is hard to beat because of class imbalance in the dataset. |
Systematic Study of Long Tail Phenomena in Entity Linking (C18-1)
Copied to clipboard
| Challenge: | Existing systems for entity linking are based on frequent 'head' cases, while performance drops when moving towards rare 'long tail' entities. |
| Approach: | They propose to use a long tail to investigate the properties of entity linking datasets. |
| Outcome: | The proposed systems overfit to popular/frequent and non-ambiguous cases and find the most difficult cases among the infrequent candidates of ambiguous forms. |
An Empirical Analysis of Diversity in Argument Summarization (2024.eacl-long)
Copied to clipboard
| Challenge: | Current methods for summarizing arguments miss an important aspect of diversity . authors examine three aspects of diversity in argument summarization . |
| Approach: | They propose three aspects of diversity that are important for accommodating multiple perspectives. |
| Outcome: | The proposed models lack the diversity of opinions, sources, and annotators. |
A Deep Dive into Word Sense Disambiguation with LSTM (C18-1)
Copied to clipboard
| Challenge: | LSTM-based language models have been shown effective in Word Sense Disambiguation (WSD) but neither the training data nor the source code was released. |
| Approach: | They propose to use LSTM-based language models to perform Word Sense Disambiguation (WSD) using openly available datasets and software. |
| Outcome: | The proposed method returned state-of-the-art performance in several benchmarks, but neither the training data nor the source code were released. |
Modeling Dutch Medical Texts for Detecting Functional Categories and Levels of COVID-19 Patients (2022.lrec-1)
Copied to clipboard
Jenia Kim, Stella Verkijk, Edwin Geleijn, Marieke van der Leeden, Carel Meskers, Caroline Meskers, Sabina van der Veen, Piek Vossen, Guy Widdershoven
| Challenge: | Electronic Health Records contain a lot of information in natural language that is not expressed in structured clinical data. |
| Approach: | They propose a Dutch language model that can determine the functional level of patients according to a WHO coding framework. |
| Outcome: | The proposed model can determine the functional level of patients according to a WHO coding framework. |
Would you describe a leopard as yellow? Evaluating crowd-annotations with justified and informative disagreement (2020.coling-main)
Copied to clipboard
| Challenge: | Existing evaluation methods rely on agreement between annotators, which implies a single correct interpretation. |
| Approach: | They propose an agreement-independent quality metric based on answer-coherence to evaluate on expected disagreement. |
| Outcome: | The proposed model shows that agreement is the most important indicator of quality in semantic annotation tasks. |
Don’t Annotate, but Validate: a Data-to-Text Method for Capturing Event Data (L18-1)
Copied to clipboard
| Challenge: | Existing methods to create event data are limited by ambiguity and variation in the data. |
| Approach: | They propose a method to obtain large volumes of text corpora with event data . they use a tool to annotate texts and enrich the reference texts with event coreference annotations. |
| Outcome: | The proposed method obtains large volumes of high-quality text corpora with event data . the data obtained with this method have high precision and at a large scale . |
Unknown Script: Impact of Script on Cross-Lingual Transfer (2024.naacl-srw)
Copied to clipboard
| Challenge: | Existing models for high-resource languages are not available for all languages, and the vast majority of the world's languages are excluded from these models. |
| Approach: | They propose to use pre-trained models to analyze the effect of the target language and its script on cross-lingual transfer. |
| Outcome: | The proposed model is based on six models pre-trained on NER and POS tasks in the original script and romanized version. |
The Circumstantial Event Ontology (CEO) and ECB+/CEO: an Ontology and Corpus for Implicit Causal Relations between Events (L18-1)
Copied to clipboard
| Challenge: | a new ontology for calamity events models semantic circumstantial relations between event classes . a circumstancial relation makes clear "why" something happened, without necessarily predicting it. |
| Approach: | They propose a circumstantial event ontology that models semantic circumstancial relations between event classes . they propose ECB+ annotated corpus for circumstantal relations and a meta model . |
| Outcome: | The proposed model captures that the change yielded by one event explains to people the happening of the next event when observed. |
Introducing Frege to Fillmore: A FrameNet Dataset that Captures both Sense and Reference (2022.lrec-1)
Copied to clipboard
| Challenge: | a widely supported claim in the fields of semantics and philosophy is that meaning arises from the combination of sense and reference. |
| Approach: | They propose a tool that facilitates both referential- and frame annotations of language-independent corpora. |
| Outcome: | The Dutch FrameNet annotation tool facilitates both referential- and frame annotations of language-independent corpora. |
Resource Interoperability for Sustainable Benchmarking: The Case of Events (L18-1)
Copied to clipboard
| Challenge: | Despite efforts to improve interoperability, there are still problems with benchmark corpora that are hampered by too laborious conversion steps. |
| Approach: | They assess aspects of interoperability at the document-level across 20 annotated corpora and compare their compatibility and consistency across the corpors. |
| Outcome: | The proposed framework enables the analysis of document intersections between the corpora and shows their compatibility and consistency across the corpus. |
Large-scale Cross-lingual Language Resources for Referencing and Framing (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing corpora that capture language but do not represent actual situations hinder development of systems to resolve cross-document coreference. |
| Approach: | They introduce the concept of cross-lingual referential corpora and propose a framework to analyze framing . they expect to capture larger variation in framation compared to traditional approaches . |
| Outcome: | The proposed project will analyze the framing of incidents in different languages and texts . it expects to capture larger variation in framation compared to traditional approaches . |
Efficiently and Thoroughly Anonymizing a Transformer Language Model for Dutch Electronic Health Records: a Two-Step Method (2022.lrec-1)
Copied to clipboard
| Challenge: | Neural Networks (NNs) are used to model large amounts of data, such as text data, and have shown to be very useful for language modelling. |
| Approach: | They propose to use a Dutch language model for hospital notes to anonymize a model trained on large amounts of data and publish it online. |
| Outcome: | The proposed method predicts a name-like token 0.2% of the time, compared to the original training data. |