TeDDi Sample: Text Data Diversity Sample for Language Comparison and Multilingual NLP (2022.lrec-1)
Copied to clipboard
| Challenge: | a deep understanding of language is being achieved through increased access to data from minority and low-resource languages. |
| Approach: | They present a diversity sample of text data for language comparison and multilingual natural language processing. |
| Outcome: | The TeDDi sample features 89 languages based on the typological diversity sample in the World Atlas of Language Structures . |
Similar Papers
A Measure for Transparent Comparison of Linguistic Diversity in Multilingual NLP Data Sets (2024.findings-naacl)
Copied to clipboard
| Challenge: | a new study aims to assess linguistic diversity of multilingual data sets against a reference language sample . linguistic diversity is typically measured as the number of languages included in the data set . but such measures do not consider structural properties of the included languages . |
| Approach: | They propose to measure linguistic diversity against a reference language sample to maximise linguistic diversity. |
| Outcome: | The proposed measure can be used to identify the types of languages that are not represented in a data set. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
What is ”Typological Diversity” in NLP? (2024.emnlp-main)
Copied to clipboard
| Challenge: | linguistic typology is commonly used to motivate language selections, but there are no set definitions or criteria for such claims. |
| Approach: | They propose to use linguistic typology to motivate language selections on the basis that a broad typological sample ought to imply generalization across a wide range of languages. |
| Outcome: | The proposed measures show that skewed language selection can lead to overestimated multilingual performance. |
Taxi1500: A Dataset for Multilingual Text Classification in 1500 Languages (2025.naacl-short)
Copied to clipboard
| Challenge: | a large-scale text classification dataset encompassing 1504 languages is needed to address this gap . low-resource languages are often overlooked due to the scarcity of evaluation datasets. |
| Approach: | They propose to use translations of the Bible to construct a large-scale text classification dataset that covers 1504 languages and annotate them using crowdsourcing. |
| Outcome: | The proposed dataset covers 1504 languages and is available to the public. |
DELTA: A Toolkit for Measuring Linguistic Diversity in Dependency-Parsed Corpora (2026.eacl-demo)
Copied to clipboard
| Challenge: | Existing tools for measuring diversity of specific linguistic phenomena are limited . we present an open-source framework for measuring linguistic diversity . |
| Approach: | They propose an open-source framework that integrates dependency tree querying with diversity computation. |
| Outcome: | The proposed framework can measure diversity across multiple linguistic levels and dimensions. |
Writing System and Speaker Metadata for 2,800+ Language Varieties (2022.lrec-1)
Copied to clipboard
| Challenge: | Currently, language technologies are easily available in only a small minority of the world's 7,000+ language varieties. |
| Approach: | They propose to use an open-source dataset to provide the writing system(s) for each of the 2,800+ languages used in the world today and an estimated speaker count for each. |
| Outcome: | The dataset provides the attested writing system(s) for each of these 2,800+ varieties, as well as an estimated speaker count for each variety. |
TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languages (2020.tacl-1)
Copied to clipboard
Jonathan H. Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, Jennimaria Palomaki
| Challenge: | Existing models for multilingual modeling are based on a set of typological features that are used to express meaning in languages such as English. |
| Approach: | They present a question-answer-typed question-referenced dataset that covers 11 typologically diverse languages with 204K question-and-answered pairs. |
| Outcome: | The proposed dataset covers 11 typologically diverse languages with 204K question-answer pairs. |
Beyond Counting Datasets: A Survey of Multilingual Dataset Construction and Necessary Resources (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies have examined the quality of labeled data in non-English languages. |
| Approach: | They annotate how datasets are created, input text and label sources, tools used to build them and what they study. |
| Outcome: | The results show that language-proficient NLP researchers' estimated availability correlates with dataset availability. |
A Multilingual Dataset for Evaluating Parallel Sentence Extraction from Comparable Corpora (L18-1)
Copied to clipboard
| Challenge: | BUCC Shared Task aims to extract parallel sentences from comparable corporad . resulting corpus contains about 3.5 million distinct sentences in english, french, german, Russian, and Chinese . |
| Approach: | They present challenges faced to build a parallel sentences dataset from comparable corporad . they emphasize issues faced to include Chinese as one of the languages . |
| Outcome: | The 2017 BUCC Shared Task was a first for this task . the dataset contains 3.5 million sentences in English, French, German, Russian, and Chinese . |
ParCourE: A Parallel Corpus Explorer for a Massively Multilingual Corpus (2021.acl-demo)
Copied to clipboard
| Challenge: | 7000 languages worldwide are spoken, but most research is focused on English . multilinguality is essential for multilingual research, and is a key component of the process. |
| Approach: | They propose a wordaligned parallel corpus that can be browsed using an online tool . they use the word alignment tools SimAlign and BabelNet to find the alignments . |
| Outcome: | The proposed tool can be set up for any parallel corpus and explores its quality and properties. |