| Challenge: | In low-resource languages, the Bible is the only significant bilingual, or even monolingual, text available . standard word alignment tools can be noisy, making downstream tasks difficult . a novel resource of 1129 aligned Bible person and place names is developed . |
| Approach: | They propose to use Bible person and place names as a tool for translation and transliteration . they use weighted edit distance, machine translation-based transliterations and affixal induction and transformation models to improve the Bible's output. |
| Outcome: | The proposed model outperforms a widely used word aligner on 97% of test words on multilingual named-entity alignment and translation across 591 languages. |
Similar Papers
An Analysis of Massively Multilingual Neural Machine Translation for Low-Resource Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | In this study, we explore massively multilingual low-resource neural machine translation. |
| Approach: | They propose to use Bible translations to train models with up to 1,107 source languages and create multilingual corpora varying the number and relatedness of source languages. |
| Outcome: | The proposed approach is highly language-specific and can be tailored to the source language and its typology. |
Towards a Broad Coverage Named Entity Resource: A Data-Efficient Approach for Many Diverse Languages (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing methods to extract named entity datasets from parallel corpora require large monolingual corporata or word aligners that are unavailable or perform poorly for underresourced languages. |
| Approach: | They propose a method for creating a multilingual named entity resource from parallel corpora and apply it to the Parallel Bible Corpus, a corpus of more than 1000 languages. |
| Outcome: | The proposed method outperforms existing methods in two tasks. |
A Comparative Study of Extremely Low-Resource Transliteration of the World’s Languages (L18-1)
Copied to clipboard
| Challenge: | a phrase-based MT system performs better than other methods for transliterating Bible names . combining data and training a single neural system yields significant gains . |
| Approach: | They compare several machine translation methods for transliterating Bible names . they find a phrase-based MT system performs better than other methods . |
| Outcome: | The phrase-based MT system performs better than other methods, the study finds . but, the single-language system outperforms the phrase-backed MT systems . |
Fine-grained Morphosyntactic Analysis and Generation Tools for More Than One Thousand Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Using morphosyntactic tools, we train and distribute tools for approximately one thousand languages. |
| Approach: | They train and distribute morphosyntactic tools for approximately one thousand languages. |
| Outcome: | The results show that the tools generalize well across rare and common forms alike. |
Massively Multilingual Token-Based Typology Using the Parallel Bible Corpus (2024.lrec-main)
Copied to clipboard
| Challenge: | linguistic typology data from the parallel Bible corpus is limited and not available for annotated corpora and automatic parsing tools. |
| Approach: | They analyze word order statistics extracted from the Bible corpus from two angles: stability across different translations in the same language and comparability with Universal Dependencies corpora and typological database classifications from URIEL and Grambank. |
| Outcome: | The results show that word order statistics extracted from the Bible corpus are reliable and generalisable across different translations in the same language. |
Learning Morphosyntactic Analyzers from the Bible via Iterative Annotation Projection across 26 Languages (P19-1)
Copied to clipboard
| Challenge: | Currently, computational tools for low-resource languages are limited by a lack of supervised training data. |
| Approach: | They propose to use English taggers and parsers to project morphological information onto translations of the Bible in 26 different test languages. |
| Outcome: | The proposed method reduces lemmatization and morphological analysis over a strong initial system. |
ParaNames 1.0: Creating an Entity Name Corpus for 400+ Languages Using Wikidata (2024.lrec-main)
Copied to clipboard
| Challenge: | ParaNames is a massively multilingual parallel name resource . it provides names for 16.8 million entities in over 400 languages . |
| Approach: | They propose a massively multilingual parallel name resource with 140 million names . they use Wikidata to standardize the data and perform canonical name translation . |
| Outcome: | The proposed resource is the largest of its type to date and performs well on 10 languages. |
JWSign: A Highly Multilingual Corpus of Bible Translations for more Diversity in Sign Language Processing (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing sign language datasets are limited and skewed towards high-income sign languages, mainly those from high-risk countries. |
| Approach: | They propose a large and highly multilingual dataset for sign language translation: JWSign. |
| Outcome: | The proposed dataset consists of 2,530 hours of Bible translations in 98 sign languages, featuring more than 1,500 individual signers. |
MaSS: A Large and Clean Multilingual Corpus of Sentence-aligned Spoken Utterances Extracted from the Bible (2020.lrec-1)
Copied to clipboard
| Challenge: | The Bible is the same for all the languages, thus constituting a multilingual and comparable 2 spoken corpus, is not exploited to date. |
| Approach: | They propose to add multilingual links between small speech segments in different languages . they use a large dataset of 8,130 parallel spoken utterances across 8 languages - maSS . |
| Outcome: | The proposed model can build automatic speech recognition models for 700 languages. |
Taxi1500: A Dataset for Multilingual Text Classification in 1500 Languages (2025.naacl-short)
Copied to clipboard
| Challenge: | a large-scale text classification dataset encompassing 1504 languages is needed to address this gap . low-resource languages are often overlooked due to the scarcity of evaluation datasets. |
| Approach: | They propose to use translations of the Bible to construct a large-scale text classification dataset that covers 1504 languages and annotate them using crowdsourcing. |
| Outcome: | The proposed dataset covers 1504 languages and is available to the public. |