Challenge: linguistic typology data from the parallel Bible corpus is limited and not available for annotated corpora and automatic parsing tools.
Approach: They analyze word order statistics extracted from the Bible corpus from two angles: stability across different translations in the same language and comparability with Universal Dependencies corpora and typological database classifications from URIEL and Grambank.
Outcome: The results show that word order statistics extracted from the Bible corpus are reliable and generalisable across different translations in the same language.

Similar Papers

The Johns Hopkins University Bible Corpus: 1600+ Tongues for Typological Exploration (2020.lrec-1)

Copied to clipboard

Challenge: Our corpus spans 1611 diverse written languages, with constituents of more than 90 language families.
Approach: They propose to scrape and merge online resources and merge them with existing corpora to create a verse-parallel scheme for all translations.
Outcome: The results show that the Bible provides high coverage of core vocabulary.
MaSS: A Large and Clean Multilingual Corpus of Sentence-aligned Spoken Utterances Extracted from the Bible (2020.lrec-1)

Copied to clipboard

Challenge: The Bible is the same for all the languages, thus constituting a multilingual and comparable 2 spoken corpus, is not exploited to date.
Approach: They propose to add multilingual links between small speech segments in different languages . they use a large dataset of 8,130 parallel spoken utterances across 8 languages - maSS .
Outcome: The proposed model can build automatic speech recognition models for 700 languages.
ParCourE: A Parallel Corpus Explorer for a Massively Multilingual Corpus (2021.acl-demo)

Copied to clipboard

Challenge: 7000 languages worldwide are spoken, but most research is focused on English . multilinguality is essential for multilingual research, and is a key component of the process.
Approach: They propose a wordaligned parallel corpus that can be browsed using an online tool . they use the word alignment tools SimAlign and BabelNet to find the alignments .
Outcome: The proposed tool can be set up for any parallel corpus and explores its quality and properties.
Fine-grained Morphosyntactic Analysis and Generation Tools for More Than One Thousand Languages (2020.lrec-1)

Copied to clipboard

Challenge: Using morphosyntactic tools, we train and distribute tools for approximately one thousand languages.
Approach: They train and distribute morphosyntactic tools for approximately one thousand languages.
Outcome: The results show that the tools generalize well across rare and common forms alike.
An Analysis of Massively Multilingual Neural Machine Translation for Low-Resource Languages (2020.lrec-1)

Copied to clipboard

Challenge: In this study, we explore massively multilingual low-resource neural machine translation.
Approach: They propose to use Bible translations to train models with up to 1,107 source languages and create multilingual corpora varying the number and relatedness of source languages.
Outcome: The proposed approach is highly language-specific and can be tailored to the source language and its typology.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
Taxi1500: A Dataset for Multilingual Text Classification in 1500 Languages (2025.naacl-short)

Copied to clipboard

Challenge: a large-scale text classification dataset encompassing 1504 languages is needed to address this gap . low-resource languages are often overlooked due to the scarcity of evaluation datasets.
Approach: They propose to use translations of the Bible to construct a large-scale text classification dataset that covers 1504 languages and annotate them using crowdsourcing.
Outcome: The proposed dataset covers 1504 languages and is available to the public.
Morphology Matters: A Multilingual Language Modeling Analysis (2021.tacl-1)

Copied to clipboard

Challenge: Existing studies on inflectional morphology disagree on whether or not it makes languages harder to model.
Approach: They propose to use a corpus of 145 Bible translations in 92 languages to investigate whether inflectional morphology makes languages harder to model.
Outcome: The proposed model trains with linguistically motivated subword segmentation strategies and reduces the impact of morphology on language modeling.
Creating a Translation Matrix of the Bible’s Names Across 591 Languages (L18-1)

Copied to clipboard

Challenge: In low-resource languages, the Bible is the only significant bilingual, or even monolingual, text available . standard word alignment tools can be noisy, making downstream tasks difficult . a novel resource of 1129 aligned Bible person and place names is developed .
Approach: They propose to use Bible person and place names as a tool for translation and transliteration . they use weighted edit distance, machine translation-based transliterations and affixal induction and transformation models to improve the Bible's output.
Outcome: The proposed model outperforms a widely used word aligner on 97% of test words on multilingual named-entity alignment and translation across 591 languages.
Universal Dependencies and Quantitative Typological Trends. A Case Study on Word Order (L18-1)

Copied to clipboard

Challenge: a new method is proposed to acquire typological evidence from "gold" treebanks for different languages.
Approach: They propose a method for acquiring typological evidence from "gold" treebanks for different languages.
Outcome: The proposed method can shed light on key issues of the linguistic typological literature.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations