Challenge: a 2010 survey found that the language of the European Union, not even English, was not fully supplied . the absence of Language Resources stifles teaching and technology building, authors say .
Approach: They propose to harness the power of alternative incentives to elicit linguistic data and annotation . they also describe changes to the workflows necessary to collect data from workforces attracted by incentives .
Outcome: a new initiative to harness incentives to elicit linguistic data and annotation improves language resources . the NIEUW project is funded by the u.s. national science foundation .

Similar Papers

Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce (2025.acl-long)

Copied to clipboard

Challenge: Language is a powerful means of communication and should be regarded as more than just a collection of tokens.
Approach: They collect feedback from individuals directly involved in and impacted by NLP artefacts for medium- and low-resource languages and highlight key issues related to data quality, cultural appropriateness and ethics of common annotation practices.
Outcome: The findings highlight key issues related to data quality, cultural appropriateness, and ethics of common annotation practices.
Beyond Counting Datasets: A Survey of Multilingual Dataset Construction and Necessary Resources (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have examined the quality of labeled data in non-English languages.
Approach: They annotate how datasets are created, input text and label sources, tools used to build them and what they study.
Outcome: The results show that language-proficient NLP researchers' estimated availability correlates with dataset availability.
CLDFBench: Give Your Cross-Linguistic Data a Lift (2020.lrec-1)

Copied to clipboard

Challenge: despite the increasing amount of cross-linguistic data, most datasets are not FAIR (findable, accessible, interoperable, and reproducible) . with the Cross-Linguistic Data Formats initiative, first standards for cross-language data have been presented and successfully tested.
Approach: They propose a framework for the retro-standardization of legacy data and the curation of new datasets that drastically simplifies the creation of CLDFs.
Outcome: The proposed framework simplifies the creation of CLDFs by providing a consistent, reproducible workflow that supports version control and long term archiving of research data and code.
Unmasking the Myth of Effortless Big Data - Making an Open Source Multi-lingual Infrastructure and Building Language Resources from Scratch (2022.lrec-1)

Copied to clipboard

Challenge: During the last two decades, machine learning approaches have dominated the field of natural language processing (NLP) weak literary traditions give rise to corpora too unreliable to function as a model for NLP tools.
Approach: They propose an alternative to corpus-based language technology that can provide language technology solutions for minority languages.
Outcome: The proposed approach can provide language technology solutions for minority languages outside the reach of corpus-based language technology.
Connecting Language Technologies with Rich, Diverse Data Sources Covering Thousands of Languages (2024.lrec-main)

Copied to clipboard

Challenge: Existing data sources for many thousands of languages are rich and diverse . Efforts are ongoing to extend technology to many more of the world's languages .
Approach: They provide an overview of some of the major online data sources available for thousands of languages.
Outcome: The proposed language technologies are based on the data available for thousands of languages.
Does Putting a Linguist in the Loop Improve NLU Data Collection? (2021.findings-emnlp)

Copied to clipboard

Challenge: Many datasets for training and evaluating natural language understanding (NLU) models contain systematic artifacts that are identified only after data collection is complete.
Approach: They propose to have linguists identify artifacts and gaps in the data and communicate with non-expert crowdworkers to adjust task instructions and incentives.
Outcome: The proposed protocol does not increase accuracy on out-of-domain test sets, and adds a chatroom does not.
Converting Legacy Data to CLDF: A FAIR Exit Strategy for Linguistic Web Apps (2024.lrec-main)

Copied to clipboard

Challenge: a number of web applications that enabled comparative linguistics research became obsolete . cross-linguistic data formats (CLDF) are available for use in linguistic research .
Approach: a new standard allows researchers to convert legacy linguistic web apps into FAIR data . the standard uses W3C recommendations Model for Tabular Data and Metadata on the Web and MetaData Vocabulary for Tabulary .
Outcome: a new standard can be used to convert legacy linguistic web apps into FAIR datasets . the standard is built on the W3C recommendations Model for Tabular Data and Metadata on the Web and MetaData Vocabulary for Tabulary on the web .
Dim Wihl Gat Tun: The Case for Linguistic Expertise in NLP for Under-Documented Languages (2022.findings-acl)

Copied to clipboard

Challenge: Recent progress in NLP is driven by pretrained models leveraging massive datasets.
Approach: They argue that IGT data can be leveraged provided target language expertise is available and that it can be used to create effective models.
Outcome: The proposed model can be leveraged provided that target language expertise is available.
XTREME-UP: A User-Centric Scarce-Data Benchmark for Under-Represented Languages (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing datasets are often informed by established research directions in the NLP community.
Approach: They propose a benchmark to evaluate the capabilities of language models across 88 under-represented languages over 9 key user-centric technologies including ASR, OCR, MT, and information access tasks.
Outcome: The proposed benchmark evaluates the capabilities of language models across 88 under-represented languages over 9 key user-centric technologies including ASR, OCR, MT, and information access tasks.
Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges (2025.acl-long)

Copied to clipboard

Challenge: linguistics studies how context influences meaning of language and how people use it to convey implied meanings, emotions, and intentions.
Approach: They analyze task designs, data collection methods, evaluation approaches and their relevance to real-world applications.
Outcome: The findings highlight emerging trends, challenges, and gaps in existing benchmarks . the findings will contribute to more nuanced and context-aware NLP models .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations