Challenge: a deep understanding of language is being achieved through increased access to data from minority and low-resource languages.
Approach: They present a diversity sample of text data for language comparison and multilingual natural language processing.
Outcome: The TeDDi sample features 89 languages based on the typological diversity sample in the World Atlas of Language Structures .

Similar Papers

A Measure for Transparent Comparison of Linguistic Diversity in Multilingual NLP Data Sets (2024.findings-naacl)

Copied to clipboard

Challenge: a new study aims to assess linguistic diversity of multilingual data sets against a reference language sample . linguistic diversity is typically measured as the number of languages included in the data set . but such measures do not consider structural properties of the included languages .
Approach: They propose to measure linguistic diversity against a reference language sample to maximise linguistic diversity.
Outcome: The proposed measure can be used to identify the types of languages that are not represented in a data set.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
What is ”Typological Diversity” in NLP? (2024.emnlp-main)

Copied to clipboard

Challenge: linguistic typology is commonly used to motivate language selections, but there are no set definitions or criteria for such claims.
Approach: They propose to use linguistic typology to motivate language selections on the basis that a broad typological sample ought to imply generalization across a wide range of languages.
Outcome: The proposed measures show that skewed language selection can lead to overestimated multilingual performance.
Taxi1500: A Dataset for Multilingual Text Classification in 1500 Languages (2025.naacl-short)

Copied to clipboard

Challenge: a large-scale text classification dataset encompassing 1504 languages is needed to address this gap . low-resource languages are often overlooked due to the scarcity of evaluation datasets.
Approach: They propose to use translations of the Bible to construct a large-scale text classification dataset that covers 1504 languages and annotate them using crowdsourcing.
Outcome: The proposed dataset covers 1504 languages and is available to the public.
DELTA: A Toolkit for Measuring Linguistic Diversity in Dependency-Parsed Corpora (2026.eacl-demo)

Copied to clipboard

Challenge: Existing tools for measuring diversity of specific linguistic phenomena are limited . we present an open-source framework for measuring linguistic diversity .
Approach: They propose an open-source framework that integrates dependency tree querying with diversity computation.
Outcome: The proposed framework can measure diversity across multiple linguistic levels and dimensions.
Writing System and Speaker Metadata for 2,800+ Language Varieties (2022.lrec-1)

Copied to clipboard

Challenge: Currently, language technologies are easily available in only a small minority of the world's 7,000+ language varieties.
Approach: They propose to use an open-source dataset to provide the writing system(s) for each of the 2,800+ languages used in the world today and an estimated speaker count for each.
Outcome: The dataset provides the attested writing system(s) for each of these 2,800+ varieties, as well as an estimated speaker count for each variety.
TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languages (2020.tacl-1)

Copied to clipboard

Challenge: Existing models for multilingual modeling are based on a set of typological features that are used to express meaning in languages such as English.
Approach: They present a question-answer-typed question-referenced dataset that covers 11 typologically diverse languages with 204K question-and-answered pairs.
Outcome: The proposed dataset covers 11 typologically diverse languages with 204K question-answer pairs.
Beyond Counting Datasets: A Survey of Multilingual Dataset Construction and Necessary Resources (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have examined the quality of labeled data in non-English languages.
Approach: They annotate how datasets are created, input text and label sources, tools used to build them and what they study.
Outcome: The results show that language-proficient NLP researchers' estimated availability correlates with dataset availability.
A Multilingual Dataset for Evaluating Parallel Sentence Extraction from Comparable Corpora (L18-1)

Copied to clipboard

Challenge: BUCC Shared Task aims to extract parallel sentences from comparable corporad . resulting corpus contains about 3.5 million distinct sentences in english, french, german, Russian, and Chinese .
Approach: They present challenges faced to build a parallel sentences dataset from comparable corporad . they emphasize issues faced to include Chinese as one of the languages .
Outcome: The 2017 BUCC Shared Task was a first for this task . the dataset contains 3.5 million sentences in English, French, German, Russian, and Chinese .
ParCourE: A Parallel Corpus Explorer for a Massively Multilingual Corpus (2021.acl-demo)

Copied to clipboard

Challenge: 7000 languages worldwide are spoken, but most research is focused on English . multilinguality is essential for multilingual research, and is a key component of the process.
Approach: They propose a wordaligned parallel corpus that can be browsed using an online tool . they use the word alignment tools SimAlign and BabelNet to find the alignments .
Outcome: The proposed tool can be set up for any parallel corpus and explores its quality and properties.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations