Does Putting a Linguist in the Loop Improve NLU Data Collection? (2021.findings-emnlp)
Copied to clipboard
Alicia Parrish, William Huang, Omar Agha, Soo-Hwan Lee, Nikita Nangia, Alexia Warstadt, Karmanya Aggarwal, Emily Allaway, Tal Linzen, Samuel R. Bowman
| Challenge: | Many datasets for training and evaluating natural language understanding (NLU) models contain systematic artifacts that are identified only after data collection is complete. |
| Approach: | They propose to have linguists identify artifacts and gaps in the data and communicate with non-expert crowdworkers to adjust task instructions and incentives. |
| Outcome: | The proposed protocol does not increase accuracy on out-of-domain test sets, and adds a chatroom does not. |
Similar Papers
What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks? (2021.acl-long)
Copied to clipboard
| Challenge: | Despite the importance of datasets for natural language understanding, there has been little attention on crowdsourcing methods for collecting datasets. |
| Approach: | They compare the effectiveness of crowdsourcing methods for boosting NLU example difficulty with training crowdworkers instead of expert judgments. |
| Outcome: | The proposed method is ineffective for boosting NLU example difficulty, but it is not effective for training crowdworkers and qualifying workers based on expert judgments. |
Adversarial NLI: A New Benchmark for Natural Language Understanding (2020.acl-main)
Copied to clipboard
| Challenge: | a new large-scale NLI benchmark dataset is presented to test models on a variety of popular NLIs. |
| Approach: | They propose a large-scale NLI benchmark dataset that is iteratively compared with a human-and-model-in-the-loop procedure. |
| Outcome: | The proposed method can be applied in a never-ending learning scenario, becoming a moving target for NLU, rather than a static benchmark that will quickly saturate. |
Beyond Counting Datasets: A Survey of Multilingual Dataset Construction and Necessary Resources (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies have examined the quality of labeled data in non-English languages. |
| Approach: | They annotate how datasets are created, input text and label sources, tools used to build them and what they study. |
| Outcome: | The results show that language-proficient NLP researchers' estimated availability correlates with dataset availability. |
New Protocols and Negative Results for Textual Entailment Data Collection (2020.emnlp-main)
Copied to clipboard
| Challenge: | Natural language inference data has proven useful in benchmarking and as pretraining data for tasks requiring language understanding. |
| Approach: | They propose four alternative protocols to improve annotation quality and diversity . they use 8.5k-example training sets to compare different protocols . |
| Outcome: | The proposed protocols improve the ease of training and quality of the examples. |
Towards an Automatic Assessment of Crowdsourced Data for NLU (L18-1)
Copied to clipboard
| Challenge: | Recent development of spoken dialog systems aims at allowing a natural input style. |
| Approach: | They investigate how crowdsourced data can be assessed with respect to its naturalness and usefulness by using a word based language model to identify valid data. |
| Outcome: | The proposed methods show that valid data can be identified with the help of a word based language model. |
Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges (2025.acl-long)
Copied to clipboard
Bolei Ma, Yuting Li, Wei Zhou, Ziwei Gong, Yang Janet Liu, Katja Jasinskaja, Annemarie Friedrich, Julia Hirschberg, Frauke Kreuter, Barbara Plank
| Challenge: | linguistics studies how context influences meaning of language and how people use it to convey implied meanings, emotions, and intentions. |
| Approach: | They analyze task designs, data collection methods, evaluation approaches and their relevance to real-world applications. |
| Outcome: | The findings highlight emerging trends, challenges, and gaps in existing benchmarks . the findings will contribute to more nuanced and context-aware NLP models . |
Deep Learning for Natural Language Inference (N19-5)
Copied to clipboard
| Challenge: | This tutorial discusses cutting-edge research on NLI, including recent advance on dataset development, cutting- edge deep learning models, and highlights from recent research on using NLI to understand capabilities and limits of deep learning for language understanding and reasoning. |
| Approach: | This tutorial discusses cutting-edge research on NLI, including recent advance on dataset development and cutting- edge deep learning models. |
| Outcome: | This tutorial discusses cutting-edge research on NLI, including recent advance on dataset development, cutting- edge deep learning models, and highlights from recent research on using NLI to understand capabilities and limits of deep learning model for language understanding and reasoning. |
Crowdsourcing Beyond Annotation: Case Studies in Benchmark Data Collection (2021.emnlp-tutorials)
Copied to clipboard
| Challenge: | Developing a theory of crowdsourcing for practical language problems remains an open challenge . |
| Approach: | This tutorial exposes NLP researchers to data collection crowdsourcing methods and principles through case studies. |
| Outcome: | This tutorial exposes NLP researchers to various data collection crowdsourcing methods and practices through case studies. |
What Can We Learn from Collective Human Opinions on Natural Language Inference Data? (2020.emnlp-main)
Copied to clipboard
| Challenge: | Despite the subjective nature of many NLU evaluations, little attention has been paid to the distribution of human opinions. |
| Approach: | They use a dataset with 464,500 annotations to study Collective HumAn OpinionS . they argue that models lack the ability to recover the distribution over human labels . |
| Outcome: | The proposed dataset examines the distribution of human opinions in NLU evaluation datasets. |
TuringAdvice: A Generative and Dynamic Evaluation of Language Use (2021.naacl-main)
Copied to clipboard
| Challenge: | Empirical results show that today’s language models struggle at TuringAdvice . language models are getting ever-larger, and are being trained on ever-increasing quantities of text . |
| Approach: | They propose a task task that requires models to generate helpful advice in natural language. |
| Outcome: | The proposed model outperforms even multibillion parameter models on 600k in-domain training examples. |