A Corpus for Large-Scale Phonetic Typology (2020.acl-main)

Copied to clipboard

Challenge: Existing multilingual speech corpora have limited data in many languages . existing corpus is limited to a small number of languages with available data .
Approach: They propose a large-scale phonetic typology corpus with phoneme-level labels and phoneme alignments in 690 readings spanning 635 languages.
Outcome: The proposed corpus covers 635 languages and includes acoustic-phonetic measures of vowels and sibilants.

Similar Papers

VoxCommunis: A Corpus for Cross-linguistic Phonetic Analysis (2022.lrec-1)

Copied to clipboard

Challenge: Until recently, the movement towards large-scale cross-linguistic phonetic research has been limited.
Approach: They propose to use the VoxCommunis Corpus to facilitate cross-linguistic phonetic research . corpus contains acoustic models, pronunciation lexicons, word- and phone-level alignments .
Outcome: The VoxCommunis Corpus contains acoustic models, pronunciation lexicons, word- and phone-level alignments . the corpus is free to download and use under a CC0 license .
Phonetic Segmentation of the UCLA Phonetics Lab Archive (2024.lrec-main)

Copied to clipboard

Challenge: ''big data'' does not exist for the majority of the world's languages . a corpus of audited phonetic transcriptions and phone-level alignments is available for free .
Approach: They present a corpus of audited phonetic transcriptions and phone-level alignments from the UCLA Phonetics Lab Archive . they discuss the utility of the corpus for general research and pedagogy in crosslinguistic phonetics .
Outcome: The VoxAngeles corpus improves the original corpus for phonetic typology and word- and phone duration measurements.
From Phonology to Syntax: Unsupervised Linguistic Typology at Different Levels with Language Embeddings (N18-1)

Copied to clipboard

Challenge: linguistic typology is the classification of languages according to their linguistic properties.
Approach: They learn distributed language representations which can be used to predict typological properties on a massively multilingual scale.
Outcome: The proposed model can predict typological properties on a massively multilingual scale.
A Deep Generative Model of Vowel Formant Typology (N18-1)

Copied to clipboard

Challenge: a recent study has investigated the nature of vowel inventories, i.e., which vowels a language contains . a probabilistic approach does not rule out linguistic systems completely, but it can position phenomena on a scale from very common to very improbable.
Approach: They propose a generative probability model of vowel inventory typology based on acoustic information rather than discrete symbols from the international phonetic alphabet.
Outcome: The proposed model uses acoustic information rather than discrete symbols from the phonetic alphabet.
Writing System and Speaker Metadata for 2,800+ Language Varieties (2022.lrec-1)

Copied to clipboard

Challenge: Currently, language technologies are easily available in only a small minority of the world's 7,000+ language varieties.
Approach: They propose to use an open-source dataset to provide the writing system(s) for each of the 2,800+ languages used in the world today and an estimated speaker count for each.
Outcome: The dataset provides the attested writing system(s) for each of these 2,800+ varieties, as well as an estimated speaker count for each variety.
The Johns Hopkins University Bible Corpus: 1600+ Tongues for Typological Exploration (2020.lrec-1)

Copied to clipboard

Challenge: Our corpus spans 1611 diverse written languages, with constituents of more than 90 language families.
Approach: They propose to scrape and merge online resources and merge them with existing corpora to create a verse-parallel scheme for all translations.
Outcome: The results show that the Bible provides high coverage of core vocabulary.
The Past, Present, and Future of Typological Databases in NLP (2023.findings-emnlp)

Copied to clipboard

Challenge: Typological information is inconsistent with each other and other sources of typological information, such as linguistic grammars.
Approach: They propose to examine disagreements between typological databases and their uses in NLP by exploring disagreements across databases and resources.
Outcome: The proposed view of typology has significant potential in the future, including in language modeling in low-resource scenarios.
On the Relation between Linguistic Typology and (Limitations of) Multilingual Language Modeling (D18-1)

Copied to clipboard

Challenge: a key challenge in cross-lingual NLP is developing general language-independent architectures that are equally applicable to any language.
Approach: They propose to use a full-vocabulary setup to test the performance of language modeling (LM) on 50 typologically diverse languages.
Outcome: The proposed language modeling task is based on a full vocabulary setup focused on word-level prediction on 50 typologically diverse languages.
LinguaMeta: Unified Metadata for Thousands of Languages (2024.lrec-main)

Copied to clipboard

Challenge: LinguaMeta is a unified repository of language metadata for thousands of languages.
Approach: They introduce LinguaMeta, a unified resource for language metadata for thousands of languages.
Outcome: The proposed resource is intended for use by researchers and organizations who aim to extend technology to thousands of languages.
BDPROTO: A Database of Phonological Inventories from Ancient and Reconstructed Languages (L18-1)

Copied to clipboard

Challenge: BDPROTO is a database of phonological inventory data from 137 ancient and reconstructed languages.
Approach: They present a database of phonological inventory data from 137 ancient and reconstructed languages.
Outcome: The BDPROTO database is a publicly available, unicode-compliant resource . it contains phonological inventory data from 137 ancient and reconstructed languages .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations