Papers by Nelleke Oostdijk

6 papers
Discourse Realization of Generics in Human and LLM-generated Texts (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models produce texts that appear coherent and credible, even when their factual reliability is uncertain.
Approach: They propose a text-level genericity score derived from clause-level annotations and apply it to argumentative essays produced by humans and LLMs.
Outcome: The proposed model is less generic than LLM-produced arguments, the study shows . higher genericity correlates with less structured, paratactic structures, the research shows a.
The CLARIN Knowledge Centre for Atypical Communication Expertise (2020.lrec-1)

Copied to clipboard

Challenge: ACE is a new knowledge center for Atypical communication experts . it is located at the Centre for Language and Speech Technology (CLST) at Radboud University .
Approach: They introduce a new CLARIN Knowledge Center called the K-Centre for Atypical Communication Expertise (ACE) ACE closely collaborates with The Language Archive at the Max Planck Institute for Psycholinguistics to safeguard GDPR-compliant data storage and access.
Outcome: The new CLARIN Knowledge Center is the K-Centre for Atypical Communication Expertise (ACE) ACE closely collaborates with The Language Archive (TLA) at the Max Planck Institute for Psycholinguistics in order to safeguard GDPR-compliant data storage and access.
The Causal News Corpus: Annotating Causal Relations in Event Sentences from News (2022.lrec-1)

Copied to clipboard

Challenge: Existing annotation guidelines for event causality focus on only explicit relations or clauses.
Approach: They propose an annotation schema for event causality that addresses these concerns . they annotated 3,559 event sentences from protest event news with labels on whether it contains causal relations or not.
Outcome: The proposed annotation schema for event causality addresses these concerns . it performs well with 81.20% F1 score on test set and 83.46% in 5-folds cross-validation .
Enhancing Discourse Parsing for Local Structures from Social Media with LLM-Generated Data (2025.coling-main)

Copied to clipboard

Challenge: Existing discourse parsers do not generalize well across genres and text types.
Approach: They propose to integrate large language models into RST discourse parsers to improve parser performance in a social media context.
Outcome: The proposed model improves parser performance in a social media context without pre-identified discourse units.
The Connection between the Text and Images of News Articles: New Insights for Multimedia Analysis (2020.lrec-1)

Copied to clipboard

Challenge: a case study of text and images reveals the inadequacy of simplistic assumptions about their connection and interplay.
Approach: They propose to use a case study to analyze 1000 flood-related news articles . they find that articles cluster into seven categories related to different topical aspects of flooding .
Outcome: The results show that flood-related news articles do not consistently report on a single, currently unfolding flooding event.
Metadata Collection Records for Language Resources (L18-1)

Copied to clipboard

Challenge: a pilot project aimed at bringing metadata records to the CLARIN context has been conducted . a virtual language observatory (VLO) was developed to provide an entry point to the language resources available in the infrastructure.
Approach: They propose to implement a CMDI profile for Dutch language resources . they propose an interface for creating, editing, listing, copying and exporting metadata records .
Outcome: The proposed interface is validated in a pilot with 45 Dutch language resources . the proposed interface provides a user interface for creating, editing, listing, copying and exporting descriptions of metadata collection records.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations