Proceedings of the First Workshop on Aggregating and Analysing Crowdsourced Annotations for NLP

7 papers
Dependency Tree Annotation with Mechanical Turk (D19-59)

Copied to clipboard

Challenge: a recent study shows that crowdsourcing is often used to obtain linguistic annotations but is rarely used for parsing.
Approach: They propose to use Mechanical Turk to crowdsource parse trees using an interactive graphical dependency tree editor.
Outcome: The proposed method is the first published use of Mechanical Turk to crowdsource parse trees . the authors find that the workers achieve high levels of accuracy on 72% of the sentences .
Word Familiarity Rate Estimation Using a Bayesian Linear Mixed Model (D19-59)

Copied to clipboard

Challenge: 96,557 words were rated using the ‘Word List by Semantic Principles’ . 96 participants were surveyed using Yahoo! crowdsourcing .
Approach: They used Bayesian linear mixed models to estimate word familiarity rates using the ‘Word List by Semantic Principles’ and the semantic labels used in the study.
Outcome: The proposed method estimated word familiarity rates using Bayesian linear mixed models and semantic labels.
Leveraging syntactic parsing to improve event annotation matching (D19-59)

Copied to clipboard

Challenge: Evaluating annotator consistency is crucial when building datasets for mention detection.
Approach: They propose to use different fuzzy-matching functions to resolve this ambiguity by extracting syntactic heads present in annotations and using the Dice coefficient to measure similarity between sets.
Outcome: The proposed functions are tested against the judgment of a human evaluator and show that the best-performing function agrees with the human .
A Dataset of Crowdsourced Word Sequences: Collections and Answer Aggregation for Ground Truth Creation (D19-59)

Copied to clipboard

Challenge: Existing work on answer aggregation for labels is limited . existing work on label aggregations is limited to label .
Approach: They propose three approaches to extractive word sequence aggregation from translated sentences generated by multiple workers.
Outcome: The proposed dataset contains translated sentences generated from multiple workers.
Crowd-sourcing annotation of complex NLU tasks: A case study of argumentative content annotation (D19-59)

Copied to clipboard

Challenge: Recent advances in machine reading and listening comprehension involve the annotation of long texts.
Approach: They propose a way to perform a sentence-by-sentence annotation task with crowd annotators.
Outcome: The proposed approach can be used to identify claims in a debate speech.
Computer Assisted Annotation of Tension Development in TED Talks through Crowdsourcing (D19-59)

Copied to clipboard

Challenge: Using a neural network, we annotate whether tension is increasing, decreasing, or staying unchanged.
Approach: They propose a machine-assisted method for the identification of tension development using a neural network based prediction model.
Outcome: The proposed method is compared with other methods in in-house and crowdsourced environments.
CoSSAT: Code-Switched Speech Annotation Tool (D19-59)

Copied to clipboard

Challenge: Code-switching is a phenomenon that occurs in multilingual societies where speakers who are fluent in two or more languages switch between these languages in the same conversation or utterance.
Approach: They propose an interface which helps annotators transcribe code-switched speech faster, more easily and more accurately than a traditional interface.
Outcome: The proposed interface can be used by 10 users to transcribe Hindi-English code-switched speech faster, easier and more accurately than a traditional interface.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations