Papers by Merel Scholman
Design Choices in Crowdsourcing Discourse Relation Annotations: The Effect of Worker Selection and Training (2022.lrec-1)
Copied to clipboard
| Challenge: | Recent methods have obtained promising results by extracting relation labels from participants . obtaining linguistic annotations from novice crowdworkers is difficult . crowdsourcing allows for fast and cost-effective collection of labelled data, but because tasks need to be intuitive, crowdworker cannot be asked to perform them. |
| Approach: | They propose to use a selection-only approach to obtain linguistic annotations from novices . current study shows that the method is cost- and time-intensive . |
| Outcome: | The current study shows that selection and training improves the agreement between workers and gold labels, but the method is cost- and time-intensive. |
Modeling Orthographic Variation Improves NLP Performance for Nigerian Pidgin (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing datasets for Nigerian Pidgin are characterised by noise in the form of orthographic variations. |
| Approach: | They propose a phonetic-theoretic framework to generate orthographic variations to augment training data. |
| Outcome: | The proposed framework improves machine translation and sentiment analysis by combining real and synthesized orthographic variations. |
DiscoGeM 2.0: A Parallel Corpus of English, German, French and Czech Implicit Discourse Relations (2024.lrec-main)
Copied to clipboard
| Challenge: | DiscoGeM 2.0 is a crowdsourced, parallel corpus of 12,834 implicit discourse relations . implicit discourse relationships are highly ambiguous and can have various interpretations . |
| Approach: | They propose a crowdsourced annotation method that can be extended to other languages . they propose to annotate 12,834 implicit discourse relations in German, German, French and Czech data . |
| Outcome: | The proposed method can be extended to other languages and reveals that implicit relations inferred in one language may differ from those inferted in the translation. |
Establishing Annotation Quality in Multi-label Annotations (2022.coling-1)
Copied to clipboard
| Challenge: | Multi-label annotations allow multiple interpretations of a single item, but they also affect the chance that two coders agree with each other. |
| Approach: | They propose a bootstrapped method to obtain chance agreement for each measure and a method to get an adjusted agreement coefficient that is more interpretable. |
| Outcome: | The proposed method allows for an adjusted agreement coefficient that is more interpretable on simulated datasets. |
DiscoGeM: A Crowdsourced Corpus of Genre-Mixed Implicit Discourse Relations (2022.lrec-1)
Copied to clipboard
| Challenge: | DiscoGeM is a crowdsourced corpus of 6,505 implicit discourse relations . the results show that a significant proportion of discourse relations are ambiguous . text genre is crucially affected by the distribution of discourse relation labels . |
| Approach: | They propose to use crowdsourced corpus of 6,505 implicit discourse relations to classify relations . they propose to include genre as a factor in automatic relation classification . |
| Outcome: | The proposed dataset shows that a significant proportion of discourse relations are ambiguous and can express multiple relation senses. |