Incorporating Worker Perspectives into MTurk Annotation Practices for NLP (2023.emnlp-main)
Copied to clipboard
| Challenge: | Current approaches to data collection for natural language processing on Amazon Mechanical Turk (MTurk) are susceptible to issues regarding workers’ rights and poor response quality without considering the perspectives of MTurq workers. |
| Approach: | They conducted a critical literature review and a survey of MTurk workers to address open questions regarding fair payment, worker privacy, data quality, and considering worker incentives. |
| Outcome: | The findings suggest that future studies may better account for MTurk workers’ experiences in order to respect workers' rights and improve response quality. |
Similar Papers
A Needle in a Haystack: An Analysis of High-Agreement Workers on MTurk for Summarization (2023.acl-long)
Copied to clipboard
Lining Zhang, Simon Mille, Yufang Hou, Daniel Deutsch, Elizabeth Clark, Yixin Liu, Saad Mahamood, Sebastian Gehrmann, Miruna Clinciu, Khyathi Raghavi Chandu, João Sedoc
| Challenge: | Using crowdsourcing, it is difficult to obtain high-quality annotations for difficult tasks. |
| Approach: | They propose a recruitment pipeline to recruit high-quality Amazon Mechanical Turk workers . they filter out subpar workers before they carry out the evaluations . |
| Outcome: | The proposed method can filter out subpar workers before they carry out evaluations and obtain high-agreement annotations with similar constraints on resources. |
The Perils of Using Mechanical Turk to Evaluate Open-Ended Text Generation (2021.emnlp-main)
Copied to clipboard
| Challenge: | Recent research has focused on open-ended text generation tasks because they are difficult to evaluate automatically. |
| Approach: | They conduct a survey of 45 open-ended text generation papers to determine whether models are reproducible . they then run story evaluation experiments with AMT workers and English teachers . |
| Outcome: | The results show that AMT workers and English teachers perform better when shown model-generated output alongside human-generated references. |
Beyond Fair Pay: Ethical Implications of NLP Crowdsourcing (2021.naacl-main)
Copied to clipboard
| Challenge: | Ethical considerations regarding the use of crowdworkers are limited to labor conditions . the Final Rule did not anticipate the use online crowdsourcing platforms for data collection . |
| Approach: | They propose to reopen discussion regarding ethical use of crowdworkers in NLP research . they propose to use online crowdsourcing platforms to evaluate risk of harm . |
| Outcome: | The proposed study identifies common scenarios where crowdworkers performing NLP tasks are at risk of harm. |
Mining Crowdsourcing Problems from Discussion Forums of Workers (2020.coling-main)
Copied to clipboard
| Challenge: | Among the most widely used platforms are Upwork, Appen, and above all Amazon Mechanical Turk (MTurk) which host annotation tasks and collect huge sets of annotated data from workers. |
| Approach: | They propose to use topic modeling to analyze workers' complaints from a new English corpus of workers’ forum discussions to identify problems in task design, task operation, and task evaluation that workers face with requesters in crowdsourcing processes. |
| Outcome: | The findings form the basis for future research on how to improve crowdsourcing processes. |
Speak: A Toolkit Using Amazon Mechanical Turk to Collect and Validate Speech Audio Recordings (2022.lrec-1)
Copied to clipboard
| Challenge: | Speak is a toolkit that allows researchers to crowdsource speech recordings using Amazon Mechanical Turk (MTurk). |
| Approach: | They propose to use Amazon Mechanical Turk to crowdsource speech recordings . they use various measures to ensure that the recordings are of adequate quality . |
| Outcome: | Speak is an open-source toolkit that allows researchers to crowdsource speech recordings using Amazon Mechanical Turk (MTurk). |
Quantifying and Avoiding Unfair Qualification Labour in Crowdsourcing (2021.acl-short)
Copied to clipboard
| Challenge: | Existing research suggests that crowd workers need to complete a substantial amount of poorly paid work to earn a fair wage. |
| Approach: | They propose to use a qualification that requires workers to have completed a certain number of tasks to earn a fair wage. |
| Outcome: | The proposed qualification reduces the burden on workers while still collecting high quality data. |
Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce (2025.acl-long)
Copied to clipboard
| Challenge: | Language is a powerful means of communication and should be regarded as more than just a collection of tokens. |
| Approach: | They collect feedback from individuals directly involved in and impacted by NLP artefacts for medium- and low-resource languages and highlight key issues related to data quality, cultural appropriateness and ethics of common annotation practices. |
| Outcome: | The findings highlight key issues related to data quality, cultural appropriateness, and ethics of common annotation practices. |
Proposal: From One-Fit-All to Perspective Aware Modeling (2025.acl-srw)
Copied to clipboard
| Challenge: | Variation in human annotation and human perspectives has drawn increasing attention in natural language processing research. |
| Approach: | They propose to use annotation formats that better capture granularity and uncertainty of individual judgments and annotation modeling that leverages socio-demographic features to better represent and predict underrepresented or minority perspectives. |
| Outcome: | The proposed tasks aim to advance natural language processing research towards more faithfully reflecting the diversity of human interpretation, enhancing both inclusiveness and fairness in language technologies. |
Whose Preferences? Differences in Fairness Preferences and Their Impact on the Fairness of AI Utilizing Human Feedback (2024.acl-long)
Copied to clipboard
| Challenge: | a growing body of work on learning from human feedback to align various aspects of machine learning systems with human values and preferences is focusing on the setting of fairness in content moderation. |
| Approach: | They propose to use human feedback to determine how two comments should be treated in content moderation to learn about human values and preferences. |
| Outcome: | The proposed approach is promising, as human preferences can often not be A: Some ladies like smaller men. B: Some men like smaller guys. Figure 1 shows that the proposed approach performs better for demographic intersections than a single classifier that gives equal weight to each annotation. |
EasyTurk: A User-Friendly Interface for High-Quality Linguistic Annotation with Amazon Mechanical Turk (2021.eacl-demos)
Copied to clipboard
| Challenge: | Amazon Mechanical Turk (AMT) is one of the most popular crowd-sourcing platforms, allowing researchers from all over the world to create linguistic datasets quickly and at a relatively low cost. |
| Approach: | They propose to improve the potential of Amazon Mechanical Turk by adding some new features to the tool. |
| Outcome: | The proposed tool improves the performance of Amazon Mechanical Turk by adding new features. |