Introducing NIEUW: Novel Incentives and Workflows for Eliciting Linguistic Data (L18-1)
Copied to clipboard
| Challenge: | a 2010 survey found that the language of the European Union, not even English, was not fully supplied . the absence of Language Resources stifles teaching and technology building, authors say . |
| Approach: | They propose to harness the power of alternative incentives to elicit linguistic data and annotation . they also describe changes to the workflows necessary to collect data from workforces attracted by incentives . |
| Outcome: | a new initiative to harness incentives to elicit linguistic data and annotation improves language resources . the NIEUW project is funded by the u.s. national science foundation . |
Similar Papers
Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce (2025.acl-long)
Copied to clipboard
| Challenge: | Language is a powerful means of communication and should be regarded as more than just a collection of tokens. |
| Approach: | They collect feedback from individuals directly involved in and impacted by NLP artefacts for medium- and low-resource languages and highlight key issues related to data quality, cultural appropriateness and ethics of common annotation practices. |
| Outcome: | The findings highlight key issues related to data quality, cultural appropriateness, and ethics of common annotation practices. |
Beyond Counting Datasets: A Survey of Multilingual Dataset Construction and Necessary Resources (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies have examined the quality of labeled data in non-English languages. |
| Approach: | They annotate how datasets are created, input text and label sources, tools used to build them and what they study. |
| Outcome: | The results show that language-proficient NLP researchers' estimated availability correlates with dataset availability. |
CLDFBench: Give Your Cross-Linguistic Data a Lift (2020.lrec-1)
Copied to clipboard
| Challenge: | despite the increasing amount of cross-linguistic data, most datasets are not FAIR (findable, accessible, interoperable, and reproducible) . with the Cross-Linguistic Data Formats initiative, first standards for cross-language data have been presented and successfully tested. |
| Approach: | They propose a framework for the retro-standardization of legacy data and the curation of new datasets that drastically simplifies the creation of CLDFs. |
| Outcome: | The proposed framework simplifies the creation of CLDFs by providing a consistent, reproducible workflow that supports version control and long term archiving of research data and code. |
Unmasking the Myth of Effortless Big Data - Making an Open Source Multi-lingual Infrastructure and Building Language Resources from Scratch (2022.lrec-1)
Copied to clipboard
Linda Wiechetek, Katri Hiovain-Asikainen, Inga Lill Sigga Mikkelsen, Sjur Moshagen, Flammie Pirinen, Trond Trosterud, Børre Gaup
| Challenge: | During the last two decades, machine learning approaches have dominated the field of natural language processing (NLP) weak literary traditions give rise to corpora too unreliable to function as a model for NLP tools. |
| Approach: | They propose an alternative to corpus-based language technology that can provide language technology solutions for minority languages. |
| Outcome: | The proposed approach can provide language technology solutions for minority languages outside the reach of corpus-based language technology. |
Connecting Language Technologies with Rich, Diverse Data Sources Covering Thousands of Languages (2024.lrec-main)
Copied to clipboard
Daan van Esch, Sandy Ritchie, Sebastian Ruder, Julia Kreutzer, Clara Rivera, Ishank Saxena, Isaac Caswell
| Challenge: | Existing data sources for many thousands of languages are rich and diverse . Efforts are ongoing to extend technology to many more of the world's languages . |
| Approach: | They provide an overview of some of the major online data sources available for thousands of languages. |
| Outcome: | The proposed language technologies are based on the data available for thousands of languages. |
Does Putting a Linguist in the Loop Improve NLU Data Collection? (2021.findings-emnlp)
Copied to clipboard
Alicia Parrish, William Huang, Omar Agha, Soo-Hwan Lee, Nikita Nangia, Alexia Warstadt, Karmanya Aggarwal, Emily Allaway, Tal Linzen, Samuel R. Bowman
| Challenge: | Many datasets for training and evaluating natural language understanding (NLU) models contain systematic artifacts that are identified only after data collection is complete. |
| Approach: | They propose to have linguists identify artifacts and gaps in the data and communicate with non-expert crowdworkers to adjust task instructions and incentives. |
| Outcome: | The proposed protocol does not increase accuracy on out-of-domain test sets, and adds a chatroom does not. |
Converting Legacy Data to CLDF: A FAIR Exit Strategy for Linguistic Web Apps (2024.lrec-main)
Copied to clipboard
| Challenge: | a number of web applications that enabled comparative linguistics research became obsolete . cross-linguistic data formats (CLDF) are available for use in linguistic research . |
| Approach: | a new standard allows researchers to convert legacy linguistic web apps into FAIR data . the standard uses W3C recommendations Model for Tabular Data and Metadata on the Web and MetaData Vocabulary for Tabulary . |
| Outcome: | a new standard can be used to convert legacy linguistic web apps into FAIR datasets . the standard is built on the W3C recommendations Model for Tabular Data and Metadata on the Web and MetaData Vocabulary for Tabulary on the web . |
Dim Wihl Gat Tun: The Case for Linguistic Expertise in NLP for Under-Documented Languages (2022.findings-acl)
Copied to clipboard
Clarissa Forbes, Farhan Samir, Bruce Oliver, Changbing Yang, Edith Coates, Garrett Nicolai, Miikka Silfverberg
| Challenge: | Recent progress in NLP is driven by pretrained models leveraging massive datasets. |
| Approach: | They argue that IGT data can be leveraged provided target language expertise is available and that it can be used to create effective models. |
| Outcome: | The proposed model can be leveraged provided that target language expertise is available. |
XTREME-UP: A User-Centric Scarce-Data Benchmark for Under-Represented Languages (2023.findings-emnlp)
Copied to clipboard
Sebastian Ruder, Jonathan Clark, Alexander Gutkin, Mihir Kale, Min Ma, Massimo Nicosia, Shruti Rijhwani, Parker Riley, Jean-Michel Sarr, Xinyi Wang, John Wieting, Nitish Gupta, Anna Katanova, Christo Kirov, Dana Dickinson, Brian Roark, Bidisha Samanta, Connie Tao, David Adelani, Vera Axelrod, Isaac Caswell, Colin Cherry, Dan Garrette, Reeve Ingle, Melvin Johnson, Dmitry Panteleev, Partha Talukdar
| Challenge: | Existing datasets are often informed by established research directions in the NLP community. |
| Approach: | They propose a benchmark to evaluate the capabilities of language models across 88 under-represented languages over 9 key user-centric technologies including ASR, OCR, MT, and information access tasks. |
| Outcome: | The proposed benchmark evaluates the capabilities of language models across 88 under-represented languages over 9 key user-centric technologies including ASR, OCR, MT, and information access tasks. |
Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges (2025.acl-long)
Copied to clipboard
Bolei Ma, Yuting Li, Wei Zhou, Ziwei Gong, Yang Janet Liu, Katja Jasinskaja, Annemarie Friedrich, Julia Hirschberg, Frauke Kreuter, Barbara Plank
| Challenge: | linguistics studies how context influences meaning of language and how people use it to convey implied meanings, emotions, and intentions. |
| Approach: | They analyze task designs, data collection methods, evaluation approaches and their relevance to real-world applications. |
| Outcome: | The findings highlight emerging trends, challenges, and gaps in existing benchmarks . the findings will contribute to more nuanced and context-aware NLP models . |