FonBund: A Library for Combining Cross-lingual Phonological Segment Data (L18-1)
Copied to clipboard
| Challenge: | Speech and language technology is currently only available for a tiny fraction of the world's languages. |
| Approach: | They propose to map phonetic segments in International Phonetic Alphabet into multiple articulatory feature representations using a free-source library. |
| Outcome: | The proposed library can be easily modified to support new phonological segment inventories. |
Similar Papers
VoxCommunis: A Corpus for Cross-linguistic Phonetic Analysis (2022.lrec-1)
Copied to clipboard
| Challenge: | Until recently, the movement towards large-scale cross-linguistic phonetic research has been limited. |
| Approach: | They propose to use the VoxCommunis Corpus to facilitate cross-linguistic phonetic research . corpus contains acoustic models, pronunciation lexicons, word- and phone-level alignments . |
| Outcome: | The VoxCommunis Corpus contains acoustic models, pronunciation lexicons, word- and phone-level alignments . the corpus is free to download and use under a CC0 license . |
Towards Universal Segmentations: UniSegments 1.0 (2022.lrec-1)
Copied to clipboard
Zdeněk Žabokrtský, Niyati Bafna, Jan Bodnár, Lukáš Kyjánek, Emil Svoboda, Magda Ševčíková, Jonáš Vidra
| Challenge: | Existing data resources for morphological segmentation are limited to 32 languages . a large number of word forms exist, with some sub-parts being "recycled" many times . |
| Approach: | They propose a multilingual data resource for morphological segmentation in 32 languages . they analyze diversity of how individual linguistic phenomena are captured across them . |
| Outcome: | The proposed scheme is based on 17 existing data resources relevant for segmentation in 32 languages. |
PWESuite: Phonetic Word Embeddings and Tasks They Facilitate (2024.lrec-main)
Copied to clipboard
Vilém Zouhar, Kalvin Chang, Chenxuan Cui, Nate B. Carlson, Nathaniel Romney Robinson, Mrinmaya Sachan, David R. Mortensen
| Challenge: | Existing word embedding methods overlook phonetic information that is crucial for many tasks. |
| Approach: | They propose three methods that use articulatory features to build phonetically informed word embeddings. |
| Outcome: | The proposed methods improve word retrieval and correlation with sound similarity and on rhyme and cognate detection tasks. |
A Corpus for Large-Scale Phonetic Typology (2020.acl-main)
Copied to clipboard
Elizabeth Salesky, Eleanor Chodroff, Tiago Pimentel, Matthew Wiesner, Ryan Cotterell, Alan W Black, Jason Eisner
| Challenge: | Existing multilingual speech corpora have limited data in many languages . existing corpus is limited to a small number of languages with available data . |
| Approach: | They propose a large-scale phonetic typology corpus with phoneme-level labels and phoneme alignments in 690 readings spanning 635 languages. |
| Outcome: | The proposed corpus covers 635 languages and includes acoustic-phonetic measures of vowels and sibilants. |
SegBo: A Database of Borrowed Sounds in the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | phonological segment borrowing is a process through which languages acquire new contrastive speech sounds as the result of borrowing words from other languages. |
| Approach: | They propose to use a database to aggregate borrowed phonological segments from languages to create a new contrastive sound. |
| Outcome: | The proposed database is based on a cross-linguistic database of borrowed phonological segments. |
Optical Character Recognition for the International Phonetic Alphabet (2026.eacl-short)
Copied to clipboard
| Challenge: | Grammar books are increasingly used as additional reference resources for low-resource languages . a significant portion of these documents come from scans and require an OCR tool . |
| Approach: | They compare two neural OCR frameworks and a large vision-language model with a synthetic dataset based on Wiktionary to study the International Phonetic Alphabet (IPA). |
| Outcome: | The proposed model improves on the International Phonetic Alphabet (IPA) character set. |
Phonotomizer: A Compact, Unsupervised, Online Training Approach to Real-Time, Multilingual Phonetic Segmentation (2025.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to phonetic segmentation are hierarchical and end-to-end . many mistakes in final output stem from subtle segmenter perturbations . |
| Approach: | They propose a phonetic segmentation system that trains on raw sound files alone . it can modulate computational exactness and reduce acoustic model size, they argue . |
| Outcome: | The proposed method reduces the size of the acoustic model and training epochs. |
KIT-Multi: A Translation-Oriented Multilingual Embedding Corpus (L18-1)
Copied to clipboard
| Challenge: | Cross-lingual word embeddings are representations of words across languages in a shared continuous vector space. |
| Approach: | They propose a multilingual word embedding corpus which is acquired by neural machine translation and is based on monolingual data. |
| Outcome: | The proposed method is competitive with existing methods but on the cross-lingual document classification task, it obtains the best figures. |
Typological Features for Multilingual Delexicalised Dependency Parsing (N19-1)
Copied to clipboard
| Challenge: | Existing universal models to describe the syntax of languages are debated for decades . a new study examines the plausibility of universal grammars in dependency parsing . |
| Approach: | They propose to use typological features to describe the syntax of languages to train a multilingual dependency parser. |
| Outcome: | The proposed model can be trained on 40 languages with the help of typological features. |
From Phonology to Syntax: Unsupervised Linguistic Typology at Different Levels with Language Embeddings (N18-1)
Copied to clipboard
| Challenge: | linguistic typology is the classification of languages according to their linguistic properties. |
| Approach: | They learn distributed language representations which can be used to predict typological properties on a massively multilingual scale. |
| Outcome: | The proposed model can predict typological properties on a massively multilingual scale. |