Elizabeth Salesky, Eleanor Chodroff, Tiago Pimentel, Matthew Wiesner, Ryan Cotterell, Alan W Black, Jason Eisner
| Challenge: | Existing multilingual speech corpora have limited data in many languages . existing corpus is limited to a small number of languages with available data . |
| Approach: | They propose a large-scale phonetic typology corpus with phoneme-level labels and phoneme alignments in 690 readings spanning 635 languages. |
| Outcome: | The proposed corpus covers 635 languages and includes acoustic-phonetic measures of vowels and sibilants. |
Similar Papers
VoxCommunis: A Corpus for Cross-linguistic Phonetic Analysis (2022.lrec-1)
Copied to clipboard
| Challenge: | Until recently, the movement towards large-scale cross-linguistic phonetic research has been limited. |
| Approach: | They propose to use the VoxCommunis Corpus to facilitate cross-linguistic phonetic research . corpus contains acoustic models, pronunciation lexicons, word- and phone-level alignments . |
| Outcome: | The VoxCommunis Corpus contains acoustic models, pronunciation lexicons, word- and phone-level alignments . the corpus is free to download and use under a CC0 license . |
Phonetic Segmentation of the UCLA Phonetics Lab Archive (2024.lrec-main)
Copied to clipboard
| Challenge: | ''big data'' does not exist for the majority of the world's languages . a corpus of audited phonetic transcriptions and phone-level alignments is available for free . |
| Approach: | They present a corpus of audited phonetic transcriptions and phone-level alignments from the UCLA Phonetics Lab Archive . they discuss the utility of the corpus for general research and pedagogy in crosslinguistic phonetics . |
| Outcome: | The VoxAngeles corpus improves the original corpus for phonetic typology and word- and phone duration measurements. |
From Phonology to Syntax: Unsupervised Linguistic Typology at Different Levels with Language Embeddings (N18-1)
Copied to clipboard
| Challenge: | linguistic typology is the classification of languages according to their linguistic properties. |
| Approach: | They learn distributed language representations which can be used to predict typological properties on a massively multilingual scale. |
| Outcome: | The proposed model can predict typological properties on a massively multilingual scale. |
A Deep Generative Model of Vowel Formant Typology (N18-1)
Copied to clipboard
| Challenge: | a recent study has investigated the nature of vowel inventories, i.e., which vowels a language contains . a probabilistic approach does not rule out linguistic systems completely, but it can position phenomena on a scale from very common to very improbable. |
| Approach: | They propose a generative probability model of vowel inventory typology based on acoustic information rather than discrete symbols from the international phonetic alphabet. |
| Outcome: | The proposed model uses acoustic information rather than discrete symbols from the phonetic alphabet. |
Writing System and Speaker Metadata for 2,800+ Language Varieties (2022.lrec-1)
Copied to clipboard
| Challenge: | Currently, language technologies are easily available in only a small minority of the world's 7,000+ language varieties. |
| Approach: | They propose to use an open-source dataset to provide the writing system(s) for each of the 2,800+ languages used in the world today and an estimated speaker count for each. |
| Outcome: | The dataset provides the attested writing system(s) for each of these 2,800+ varieties, as well as an estimated speaker count for each variety. |
The Johns Hopkins University Bible Corpus: 1600+ Tongues for Typological Exploration (2020.lrec-1)
Copied to clipboard
Arya D. McCarthy, Rachel Wicks, Dylan Lewis, Aaron Mueller, Winston Wu, Oliver Adams, Garrett Nicolai, Matt Post, David Yarowsky
| Challenge: | Our corpus spans 1611 diverse written languages, with constituents of more than 90 language families. |
| Approach: | They propose to scrape and merge online resources and merge them with existing corpora to create a verse-parallel scheme for all translations. |
| Outcome: | The results show that the Bible provides high coverage of core vocabulary. |
The Past, Present, and Future of Typological Databases in NLP (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Typological information is inconsistent with each other and other sources of typological information, such as linguistic grammars. |
| Approach: | They propose to examine disagreements between typological databases and their uses in NLP by exploring disagreements across databases and resources. |
| Outcome: | The proposed view of typology has significant potential in the future, including in language modeling in low-resource scenarios. |
On the Relation between Linguistic Typology and (Limitations of) Multilingual Language Modeling (D18-1)
Copied to clipboard
| Challenge: | a key challenge in cross-lingual NLP is developing general language-independent architectures that are equally applicable to any language. |
| Approach: | They propose to use a full-vocabulary setup to test the performance of language modeling (LM) on 50 typologically diverse languages. |
| Outcome: | The proposed language modeling task is based on a full vocabulary setup focused on word-level prediction on 50 typologically diverse languages. |
LinguaMeta: Unified Metadata for Thousands of Languages (2024.lrec-main)
Copied to clipboard
| Challenge: | LinguaMeta is a unified repository of language metadata for thousands of languages. |
| Approach: | They introduce LinguaMeta, a unified resource for language metadata for thousands of languages. |
| Outcome: | The proposed resource is intended for use by researchers and organizations who aim to extend technology to thousands of languages. |
BDPROTO: A Database of Phonological Inventories from Ancient and Reconstructed Languages (L18-1)
Copied to clipboard
| Challenge: | BDPROTO is a database of phonological inventory data from 137 ancient and reconstructed languages. |
| Approach: | They present a database of phonological inventory data from 137 ancient and reconstructed languages. |
| Outcome: | The BDPROTO database is a publicly available, unicode-compliant resource . it contains phonological inventory data from 137 ancient and reconstructed languages . |