Challenge: Speech and language technology is currently only available for a tiny fraction of the world's languages.
Approach: They propose to map phonetic segments in International Phonetic Alphabet into multiple articulatory feature representations using a free-source library.
Outcome: The proposed library can be easily modified to support new phonological segment inventories.

Similar Papers

VoxCommunis: A Corpus for Cross-linguistic Phonetic Analysis (2022.lrec-1)

Copied to clipboard

Challenge: Until recently, the movement towards large-scale cross-linguistic phonetic research has been limited.
Approach: They propose to use the VoxCommunis Corpus to facilitate cross-linguistic phonetic research . corpus contains acoustic models, pronunciation lexicons, word- and phone-level alignments .
Outcome: The VoxCommunis Corpus contains acoustic models, pronunciation lexicons, word- and phone-level alignments . the corpus is free to download and use under a CC0 license .
Towards Universal Segmentations: UniSegments 1.0 (2022.lrec-1)

Copied to clipboard

Challenge: Existing data resources for morphological segmentation are limited to 32 languages . a large number of word forms exist, with some sub-parts being "recycled" many times .
Approach: They propose a multilingual data resource for morphological segmentation in 32 languages . they analyze diversity of how individual linguistic phenomena are captured across them .
Outcome: The proposed scheme is based on 17 existing data resources relevant for segmentation in 32 languages.
PWESuite: Phonetic Word Embeddings and Tasks They Facilitate (2024.lrec-main)

Copied to clipboard

Challenge: Existing word embedding methods overlook phonetic information that is crucial for many tasks.
Approach: They propose three methods that use articulatory features to build phonetically informed word embeddings.
Outcome: The proposed methods improve word retrieval and correlation with sound similarity and on rhyme and cognate detection tasks.
A Corpus for Large-Scale Phonetic Typology (2020.acl-main)

Copied to clipboard

Challenge: Existing multilingual speech corpora have limited data in many languages . existing corpus is limited to a small number of languages with available data .
Approach: They propose a large-scale phonetic typology corpus with phoneme-level labels and phoneme alignments in 690 readings spanning 635 languages.
Outcome: The proposed corpus covers 635 languages and includes acoustic-phonetic measures of vowels and sibilants.
SegBo: A Database of Borrowed Sounds in the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: phonological segment borrowing is a process through which languages acquire new contrastive speech sounds as the result of borrowing words from other languages.
Approach: They propose to use a database to aggregate borrowed phonological segments from languages to create a new contrastive sound.
Outcome: The proposed database is based on a cross-linguistic database of borrowed phonological segments.
Optical Character Recognition for the International Phonetic Alphabet (2026.eacl-short)

Copied to clipboard

Challenge: Grammar books are increasingly used as additional reference resources for low-resource languages . a significant portion of these documents come from scans and require an OCR tool .
Approach: They compare two neural OCR frameworks and a large vision-language model with a synthetic dataset based on Wiktionary to study the International Phonetic Alphabet (IPA).
Outcome: The proposed model improves on the International Phonetic Alphabet (IPA) character set.
Phonotomizer: A Compact, Unsupervised, Online Training Approach to Real-Time, Multilingual Phonetic Segmentation (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to phonetic segmentation are hierarchical and end-to-end . many mistakes in final output stem from subtle segmenter perturbations .
Approach: They propose a phonetic segmentation system that trains on raw sound files alone . it can modulate computational exactness and reduce acoustic model size, they argue .
Outcome: The proposed method reduces the size of the acoustic model and training epochs.
KIT-Multi: A Translation-Oriented Multilingual Embedding Corpus (L18-1)

Copied to clipboard

Challenge: Cross-lingual word embeddings are representations of words across languages in a shared continuous vector space.
Approach: They propose a multilingual word embedding corpus which is acquired by neural machine translation and is based on monolingual data.
Outcome: The proposed method is competitive with existing methods but on the cross-lingual document classification task, it obtains the best figures.
Typological Features for Multilingual Delexicalised Dependency Parsing (N19-1)

Copied to clipboard

Challenge: Existing universal models to describe the syntax of languages are debated for decades . a new study examines the plausibility of universal grammars in dependency parsing .
Approach: They propose to use typological features to describe the syntax of languages to train a multilingual dependency parser.
Outcome: The proposed model can be trained on 40 languages with the help of typological features.
From Phonology to Syntax: Unsupervised Linguistic Typology at Different Levels with Language Embeddings (N18-1)

Copied to clipboard

Challenge: linguistic typology is the classification of languages according to their linguistic properties.
Approach: They learn distributed language representations which can be used to predict typological properties on a massively multilingual scale.
Outcome: The proposed model can predict typological properties on a massively multilingual scale.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations