Challenge: In low-resource languages, the Bible is the only significant bilingual, or even monolingual, text available . standard word alignment tools can be noisy, making downstream tasks difficult . a novel resource of 1129 aligned Bible person and place names is developed .
Approach: They propose to use Bible person and place names as a tool for translation and transliteration . they use weighted edit distance, machine translation-based transliterations and affixal induction and transformation models to improve the Bible's output.
Outcome: The proposed model outperforms a widely used word aligner on 97% of test words on multilingual named-entity alignment and translation across 591 languages.

Similar Papers

An Analysis of Massively Multilingual Neural Machine Translation for Low-Resource Languages (2020.lrec-1)

Copied to clipboard

Challenge: In this study, we explore massively multilingual low-resource neural machine translation.
Approach: They propose to use Bible translations to train models with up to 1,107 source languages and create multilingual corpora varying the number and relatedness of source languages.
Outcome: The proposed approach is highly language-specific and can be tailored to the source language and its typology.
Towards a Broad Coverage Named Entity Resource: A Data-Efficient Approach for Many Diverse Languages (2022.lrec-1)

Copied to clipboard

Challenge: Existing methods to extract named entity datasets from parallel corpora require large monolingual corporata or word aligners that are unavailable or perform poorly for underresourced languages.
Approach: They propose a method for creating a multilingual named entity resource from parallel corpora and apply it to the Parallel Bible Corpus, a corpus of more than 1000 languages.
Outcome: The proposed method outperforms existing methods in two tasks.
A Comparative Study of Extremely Low-Resource Transliteration of the World’s Languages (L18-1)

Copied to clipboard

Challenge: a phrase-based MT system performs better than other methods for transliterating Bible names . combining data and training a single neural system yields significant gains .
Approach: They compare several machine translation methods for transliterating Bible names . they find a phrase-based MT system performs better than other methods .
Outcome: The phrase-based MT system performs better than other methods, the study finds . but, the single-language system outperforms the phrase-backed MT systems .
Fine-grained Morphosyntactic Analysis and Generation Tools for More Than One Thousand Languages (2020.lrec-1)

Copied to clipboard

Challenge: Using morphosyntactic tools, we train and distribute tools for approximately one thousand languages.
Approach: They train and distribute morphosyntactic tools for approximately one thousand languages.
Outcome: The results show that the tools generalize well across rare and common forms alike.
Massively Multilingual Token-Based Typology Using the Parallel Bible Corpus (2024.lrec-main)

Copied to clipboard

Challenge: linguistic typology data from the parallel Bible corpus is limited and not available for annotated corpora and automatic parsing tools.
Approach: They analyze word order statistics extracted from the Bible corpus from two angles: stability across different translations in the same language and comparability with Universal Dependencies corpora and typological database classifications from URIEL and Grambank.
Outcome: The results show that word order statistics extracted from the Bible corpus are reliable and generalisable across different translations in the same language.
Learning Morphosyntactic Analyzers from the Bible via Iterative Annotation Projection across 26 Languages (P19-1)

Copied to clipboard

Challenge: Currently, computational tools for low-resource languages are limited by a lack of supervised training data.
Approach: They propose to use English taggers and parsers to project morphological information onto translations of the Bible in 26 different test languages.
Outcome: The proposed method reduces lemmatization and morphological analysis over a strong initial system.
ParaNames 1.0: Creating an Entity Name Corpus for 400+ Languages Using Wikidata (2024.lrec-main)

Copied to clipboard

Challenge: ParaNames is a massively multilingual parallel name resource . it provides names for 16.8 million entities in over 400 languages .
Approach: They propose a massively multilingual parallel name resource with 140 million names . they use Wikidata to standardize the data and perform canonical name translation .
Outcome: The proposed resource is the largest of its type to date and performs well on 10 languages.
JWSign: A Highly Multilingual Corpus of Bible Translations for more Diversity in Sign Language Processing (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing sign language datasets are limited and skewed towards high-income sign languages, mainly those from high-risk countries.
Approach: They propose a large and highly multilingual dataset for sign language translation: JWSign.
Outcome: The proposed dataset consists of 2,530 hours of Bible translations in 98 sign languages, featuring more than 1,500 individual signers.
MaSS: A Large and Clean Multilingual Corpus of Sentence-aligned Spoken Utterances Extracted from the Bible (2020.lrec-1)

Copied to clipboard

Challenge: The Bible is the same for all the languages, thus constituting a multilingual and comparable 2 spoken corpus, is not exploited to date.
Approach: They propose to add multilingual links between small speech segments in different languages . they use a large dataset of 8,130 parallel spoken utterances across 8 languages - maSS .
Outcome: The proposed model can build automatic speech recognition models for 700 languages.
Taxi1500: A Dataset for Multilingual Text Classification in 1500 Languages (2025.naacl-short)

Copied to clipboard

Challenge: a large-scale text classification dataset encompassing 1504 languages is needed to address this gap . low-resource languages are often overlooked due to the scarcity of evaluation datasets.
Approach: They propose to use translations of the Bible to construct a large-scale text classification dataset that covers 1504 languages and annotate them using crowdsourcing.
Outcome: The proposed dataset covers 1504 languages and is available to the public.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations