Massively Multilingual Token-Based Typology Using the Parallel Bible Corpus (2024.lrec-main)
Copied to clipboard
| Challenge: | linguistic typology data from the parallel Bible corpus is limited and not available for annotated corpora and automatic parsing tools. |
| Approach: | They analyze word order statistics extracted from the Bible corpus from two angles: stability across different translations in the same language and comparability with Universal Dependencies corpora and typological database classifications from URIEL and Grambank. |
| Outcome: | The results show that word order statistics extracted from the Bible corpus are reliable and generalisable across different translations in the same language. |
Similar Papers
The Johns Hopkins University Bible Corpus: 1600+ Tongues for Typological Exploration (2020.lrec-1)
Copied to clipboard
Arya D. McCarthy, Rachel Wicks, Dylan Lewis, Aaron Mueller, Winston Wu, Oliver Adams, Garrett Nicolai, Matt Post, David Yarowsky
| Challenge: | Our corpus spans 1611 diverse written languages, with constituents of more than 90 language families. |
| Approach: | They propose to scrape and merge online resources and merge them with existing corpora to create a verse-parallel scheme for all translations. |
| Outcome: | The results show that the Bible provides high coverage of core vocabulary. |
MaSS: A Large and Clean Multilingual Corpus of Sentence-aligned Spoken Utterances Extracted from the Bible (2020.lrec-1)
Copied to clipboard
| Challenge: | The Bible is the same for all the languages, thus constituting a multilingual and comparable 2 spoken corpus, is not exploited to date. |
| Approach: | They propose to add multilingual links between small speech segments in different languages . they use a large dataset of 8,130 parallel spoken utterances across 8 languages - maSS . |
| Outcome: | The proposed model can build automatic speech recognition models for 700 languages. |
ParCourE: A Parallel Corpus Explorer for a Massively Multilingual Corpus (2021.acl-demo)
Copied to clipboard
| Challenge: | 7000 languages worldwide are spoken, but most research is focused on English . multilinguality is essential for multilingual research, and is a key component of the process. |
| Approach: | They propose a wordaligned parallel corpus that can be browsed using an online tool . they use the word alignment tools SimAlign and BabelNet to find the alignments . |
| Outcome: | The proposed tool can be set up for any parallel corpus and explores its quality and properties. |
Fine-grained Morphosyntactic Analysis and Generation Tools for More Than One Thousand Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Using morphosyntactic tools, we train and distribute tools for approximately one thousand languages. |
| Approach: | They train and distribute morphosyntactic tools for approximately one thousand languages. |
| Outcome: | The results show that the tools generalize well across rare and common forms alike. |
An Analysis of Massively Multilingual Neural Machine Translation for Low-Resource Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | In this study, we explore massively multilingual low-resource neural machine translation. |
| Approach: | They propose to use Bible translations to train models with up to 1,107 source languages and create multilingual corpora varying the number and relatedness of source languages. |
| Outcome: | The proposed approach is highly language-specific and can be tailored to the source language and its typology. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
Taxi1500: A Dataset for Multilingual Text Classification in 1500 Languages (2025.naacl-short)
Copied to clipboard
| Challenge: | a large-scale text classification dataset encompassing 1504 languages is needed to address this gap . low-resource languages are often overlooked due to the scarcity of evaluation datasets. |
| Approach: | They propose to use translations of the Bible to construct a large-scale text classification dataset that covers 1504 languages and annotate them using crowdsourcing. |
| Outcome: | The proposed dataset covers 1504 languages and is available to the public. |
Morphology Matters: A Multilingual Language Modeling Analysis (2021.tacl-1)
Copied to clipboard
| Challenge: | Existing studies on inflectional morphology disagree on whether or not it makes languages harder to model. |
| Approach: | They propose to use a corpus of 145 Bible translations in 92 languages to investigate whether inflectional morphology makes languages harder to model. |
| Outcome: | The proposed model trains with linguistically motivated subword segmentation strategies and reduces the impact of morphology on language modeling. |
Creating a Translation Matrix of the Bible’s Names Across 591 Languages (L18-1)
Copied to clipboard
| Challenge: | In low-resource languages, the Bible is the only significant bilingual, or even monolingual, text available . standard word alignment tools can be noisy, making downstream tasks difficult . a novel resource of 1129 aligned Bible person and place names is developed . |
| Approach: | They propose to use Bible person and place names as a tool for translation and transliteration . they use weighted edit distance, machine translation-based transliterations and affixal induction and transformation models to improve the Bible's output. |
| Outcome: | The proposed model outperforms a widely used word aligner on 97% of test words on multilingual named-entity alignment and translation across 591 languages. |
Universal Dependencies and Quantitative Typological Trends. A Case Study on Word Order (L18-1)
Copied to clipboard
| Challenge: | a new method is proposed to acquire typological evidence from "gold" treebanks for different languages. |
| Approach: | They propose a method for acquiring typological evidence from "gold" treebanks for different languages. |
| Outcome: | The proposed method can shed light on key issues of the linguistic typological literature. |