Proceedings of the First Workshop on Aggregating and Analysing Crowdsourced Annotations for NLP
Dependency Tree Annotation with Mechanical Turk (D19-59)
Copied to clipboard
| Challenge: | a recent study shows that crowdsourcing is often used to obtain linguistic annotations but is rarely used for parsing. |
| Approach: | They propose to use Mechanical Turk to crowdsource parse trees using an interactive graphical dependency tree editor. |
| Outcome: | The proposed method is the first published use of Mechanical Turk to crowdsource parse trees . the authors find that the workers achieve high levels of accuracy on 72% of the sentences . |
Word Familiarity Rate Estimation Using a Bayesian Linear Mixed Model (D19-59)
Copied to clipboard
| Challenge: | 96,557 words were rated using the ‘Word List by Semantic Principles’ . 96 participants were surveyed using Yahoo! crowdsourcing . |
| Approach: | They used Bayesian linear mixed models to estimate word familiarity rates using the ‘Word List by Semantic Principles’ and the semantic labels used in the study. |
| Outcome: | The proposed method estimated word familiarity rates using Bayesian linear mixed models and semantic labels. |
Leveraging syntactic parsing to improve event annotation matching (D19-59)
Copied to clipboard
| Challenge: | Evaluating annotator consistency is crucial when building datasets for mention detection. |
| Approach: | They propose to use different fuzzy-matching functions to resolve this ambiguity by extracting syntactic heads present in annotations and using the Dice coefficient to measure similarity between sets. |
| Outcome: | The proposed functions are tested against the judgment of a human evaluator and show that the best-performing function agrees with the human . |
A Dataset of Crowdsourced Word Sequences: Collections and Answer Aggregation for Ground Truth Creation (D19-59)
Copied to clipboard
| Challenge: | Existing work on answer aggregation for labels is limited . existing work on label aggregations is limited to label . |
| Approach: | They propose three approaches to extractive word sequence aggregation from translated sentences generated by multiple workers. |
| Outcome: | The proposed dataset contains translated sentences generated from multiple workers. |
Crowd-sourcing annotation of complex NLU tasks: A case study of argumentative content annotation (D19-59)
Copied to clipboard
| Challenge: | Recent advances in machine reading and listening comprehension involve the annotation of long texts. |
| Approach: | They propose a way to perform a sentence-by-sentence annotation task with crowd annotators. |
| Outcome: | The proposed approach can be used to identify claims in a debate speech. |
Computer Assisted Annotation of Tension Development in TED Talks through Crowdsourcing (D19-59)
Copied to clipboard
| Challenge: | Using a neural network, we annotate whether tension is increasing, decreasing, or staying unchanged. |
| Approach: | They propose a machine-assisted method for the identification of tension development using a neural network based prediction model. |
| Outcome: | The proposed method is compared with other methods in in-house and crowdsourced environments. |
CoSSAT: Code-Switched Speech Annotation Tool (D19-59)
Copied to clipboard
| Challenge: | Code-switching is a phenomenon that occurs in multilingual societies where speakers who are fluent in two or more languages switch between these languages in the same conversation or utterance. |
| Approach: | They propose an interface which helps annotators transcribe code-switched speech faster, more easily and more accurately than a traditional interface. |
| Outcome: | The proposed interface can be used by 10 users to transcribe Hindi-English code-switched speech faster, easier and more accurately than a traditional interface. |