Papers by Merel Scholman

5 papers
Design Choices in Crowdsourcing Discourse Relation Annotations: The Effect of Worker Selection and Training (2022.lrec-1)

Copied to clipboard

Challenge: Recent methods have obtained promising results by extracting relation labels from participants . obtaining linguistic annotations from novice crowdworkers is difficult . crowdsourcing allows for fast and cost-effective collection of labelled data, but because tasks need to be intuitive, crowdworker cannot be asked to perform them.
Approach: They propose to use a selection-only approach to obtain linguistic annotations from novices . current study shows that the method is cost- and time-intensive .
Outcome: The current study shows that selection and training improves the agreement between workers and gold labels, but the method is cost- and time-intensive.
Modeling Orthographic Variation Improves NLP Performance for Nigerian Pidgin (2024.lrec-main)

Copied to clipboard

Challenge: Existing datasets for Nigerian Pidgin are characterised by noise in the form of orthographic variations.
Approach: They propose a phonetic-theoretic framework to generate orthographic variations to augment training data.
Outcome: The proposed framework improves machine translation and sentiment analysis by combining real and synthesized orthographic variations.
DiscoGeM 2.0: A Parallel Corpus of English, German, French and Czech Implicit Discourse Relations (2024.lrec-main)

Copied to clipboard

Challenge: DiscoGeM 2.0 is a crowdsourced, parallel corpus of 12,834 implicit discourse relations . implicit discourse relationships are highly ambiguous and can have various interpretations .
Approach: They propose a crowdsourced annotation method that can be extended to other languages . they propose to annotate 12,834 implicit discourse relations in German, German, French and Czech data .
Outcome: The proposed method can be extended to other languages and reveals that implicit relations inferred in one language may differ from those inferted in the translation.
Establishing Annotation Quality in Multi-label Annotations (2022.coling-1)

Copied to clipboard

Challenge: Multi-label annotations allow multiple interpretations of a single item, but they also affect the chance that two coders agree with each other.
Approach: They propose a bootstrapped method to obtain chance agreement for each measure and a method to get an adjusted agreement coefficient that is more interpretable.
Outcome: The proposed method allows for an adjusted agreement coefficient that is more interpretable on simulated datasets.
DiscoGeM: A Crowdsourced Corpus of Genre-Mixed Implicit Discourse Relations (2022.lrec-1)

Copied to clipboard

Challenge: DiscoGeM is a crowdsourced corpus of 6,505 implicit discourse relations . the results show that a significant proportion of discourse relations are ambiguous . text genre is crucially affected by the distribution of discourse relation labels .
Approach: They propose to use crowdsourced corpus of 6,505 implicit discourse relations to classify relations . they propose to include genre as a factor in automatic relation classification .
Outcome: The proposed dataset shows that a significant proportion of discourse relations are ambiguous and can express multiple relation senses.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations