Parallel Corpora in Mboshi (Bantu C25, Congo-Brazzaville) (L18-1)

Copied to clipboard

Challenge: BULB project aims to provide tools to language documentation and description for unwritten languages . language-based technologies are needed to support the collection of data and to provide linguistic documentation for the languages.
Approach: This paper presents multimodal and parallel data collections in Mboshi, as part of the French-German BULB project.
Outcome: The proposed data collection includes pictures and videos documenting social practices, agriculture, wildlife and plants.

Similar Papers

BULBasaa: A Bilingual Basaa-French Speech Corpus for the Evaluation of Language Documentation Tools (L18-1)

Copied to clipboard

Challenge: Approximately 50 hours of Bàsàá speech were collected and then carefully re-spoken and orally translated into French .
Approach: They propose to provide an automatic phonetic transcription using a set of derived phone-like units.
Outcome: The proposed method provides an automatic phonetic transcription using a set of derived phone-like units.
A New Massive Multilingual Dataset for High-Performance Language Technologies (2024.lrec-main)

Copied to clipboard

Challenge: a new massive multilingual dataset is available for language modeling and machine translation training.
Approach: They present a massive multilingual dataset using web crawls from the Internet Archive and CommonCrawl . they use open-source software tools and high-performance computing to acquire, manage and process large corpora .
Outcome: The HPLT language resources is a massive multilingual dataset . it includes monolingual and bilingual corpora extracted from CommonCrawl and the Internet Archive . the results are published online at the journal journal cense4 .
A Multilingual Dataset for Evaluating Parallel Sentence Extraction from Comparable Corpora (L18-1)

Copied to clipboard

Challenge: BUCC Shared Task aims to extract parallel sentences from comparable corporad . resulting corpus contains about 3.5 million distinct sentences in english, french, german, Russian, and Chinese .
Approach: They present challenges faced to build a parallel sentences dataset from comparable corporad . they emphasize issues faced to include Chinese as one of the languages .
Outcome: The 2017 BUCC Shared Task was a first for this task . the dataset contains 3.5 million sentences in English, French, German, Russian, and Chinese .
A Multilingual Parallel Corpora Collection Effort for Indian Languages (2020.lrec-1)

Copied to clipboard

Challenge: Currently, neural network based approaches for machine translation are data hungry and sentence-level aligned parallel pairs are the currency.
Approach: They propose to build sentence aligned parallel corpora across 10 Indian languages using online sources which have content shared across languages.
Outcome: The proposed corpora significantly extends existing resources that are either not large enough or are restricted to a specific domain (such as health).
A Very Low Resource Language Speech Corpus for Computational Language Documentation Experiments (L18-1)

Copied to clipboard

Challenge: a new study aims to document endangered languages using a speech corpus . linguistic documentation is limited to the phonetic, lexical and syntactic levels .
Approach: They propose to use a speech corpus to document endangered languages in field . they propose to collect 5k speech utterances aligned to French text translations .
Outcome: The proposed language corpus is used to document endangered languages in field linguists . it is multilingual and contains 5k speech utterances aligned to french text translations - the authors show it can be used in a zero-resource task .
A Recipe of Parallel Corpora Exploitation for Multilingual Large Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Recent studies have highlighted the potential of exploiting parallel corpora to enhance multilingual large language models.
Approach: They investigate the impact of parallel corpora quality and quantity, training objectives, and model size on performance of multilingual large language models enhanced with parallel corporeal.
Outcome: The proposed approach improves performance in bilingual and general-purpose tasks.
Parallel Corpora for the Biomedical Domain (L18-1)

Copied to clipboard

Challenge: Existing corpora of parallel corporata are being used in the biomedical domain . MT is known to support readers' access to textual documents in a language other than their native language .
Approach: They propose to leverage parallel corpora to implement cross-lingual information retrieval or machine translation tools.
Outcome: The proposed corpus is being used in the biomedical task at the conference on machine translation (WMT'16 and WMT'17) it can be leveraged to provide access to health information in languages other than English.
Augmenting Librispeech with French Translations: A Multimodal Corpus for Direct Speech Translation Evaluation (L18-1)

Copied to clipboard

Challenge: Recent work in spoken language translation (SLT) has attempted to build end-to-end speech-totext translation without using source language transcription during learning or decoding.
Approach: They propose to augment an existing (monolingual) corpus: LibriSpeech.
Outcome: The proposed corpus is derived from read audiobooks from the LibriVox project and has been carefully segmented and aligned.
SwissSLi: The Multi-parallel Sign Language Corpus for Switzerland (2024.lrec-main)

Copied to clipboard

Challenge: Using a CC BY-NC-SA 4.0 license, this corpus contains parallel sign language videos and spoken language subtitles.
Approach: They introduce SwissSLi, the first sign language corpus that contains parallel data of all three Swiss sign languages.
Outcome: The proposed corpus contains parallel sign language videos and spoken language subtitles.
Towards Building an Automatic Transcription System for Language Documentation: Experiences from Muyu (2020.lrec-1)

Copied to clipboard

Challenge: Language documentation is a rapidly growing field due to its urgency.
Approach: They propose to use phoneme recognition to automatically recognize spoken languages and translate them to global languages.
Outcome: The proposed tool performs better than existing methods with American English, Austrian German and Slovenian as source and target languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations