Annie Rialland, Martine Adda-Decker, Guy-Noël Kouarata, Gilles Adda, Laurent Besacier, Lori Lamel, Elodie Gauthier, Pierre Godard, Jamison Cooper-Leavitt
| Challenge: | BULB project aims to provide tools to language documentation and description for unwritten languages . language-based technologies are needed to support the collection of data and to provide linguistic documentation for the languages. |
| Approach: | This paper presents multimodal and parallel data collections in Mboshi, as part of the French-German BULB project. |
| Outcome: | The proposed data collection includes pictures and videos documenting social practices, agriculture, wildlife and plants. |
Similar Papers
BULBasaa: A Bilingual Basaa-French Speech Corpus for the Evaluation of Language Documentation Tools (L18-1)
Copied to clipboard
Fatima Hamlaoui, Emmanuel-Moselly Makasso, Markus Müller, Jonas Engelmann, Gilles Adda, Alex Waibel, Sebastian Stüker
| Challenge: | Approximately 50 hours of Bàsàá speech were collected and then carefully re-spoken and orally translated into French . |
| Approach: | They propose to provide an automatic phonetic transcription using a set of derived phone-like units. |
| Outcome: | The proposed method provides an automatic phonetic transcription using a set of derived phone-like units. |
A New Massive Multilingual Dataset for High-Performance Language Technologies (2024.lrec-main)
Copied to clipboard
Ona de Gibert, Graeme Nail, Nikolay Arefyev, Marta Bañón, Jelmer van der Linde, Shaoxiong Ji, Jaume Zaragoza-Bernabeu, Mikko Aulamo, Gema Ramírez-Sánchez, Andrey Kutuzov, Sampo Pyysalo, Stephan Oepen, Jörg Tiedemann
| Challenge: | a new massive multilingual dataset is available for language modeling and machine translation training. |
| Approach: | They present a massive multilingual dataset using web crawls from the Internet Archive and CommonCrawl . they use open-source software tools and high-performance computing to acquire, manage and process large corpora . |
| Outcome: | The HPLT language resources is a massive multilingual dataset . it includes monolingual and bilingual corpora extracted from CommonCrawl and the Internet Archive . the results are published online at the journal journal cense4 . |
A Multilingual Dataset for Evaluating Parallel Sentence Extraction from Comparable Corpora (L18-1)
Copied to clipboard
| Challenge: | BUCC Shared Task aims to extract parallel sentences from comparable corporad . resulting corpus contains about 3.5 million distinct sentences in english, french, german, Russian, and Chinese . |
| Approach: | They present challenges faced to build a parallel sentences dataset from comparable corporad . they emphasize issues faced to include Chinese as one of the languages . |
| Outcome: | The 2017 BUCC Shared Task was a first for this task . the dataset contains 3.5 million sentences in English, French, German, Russian, and Chinese . |
A Multilingual Parallel Corpora Collection Effort for Indian Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Currently, neural network based approaches for machine translation are data hungry and sentence-level aligned parallel pairs are the currency. |
| Approach: | They propose to build sentence aligned parallel corpora across 10 Indian languages using online sources which have content shared across languages. |
| Outcome: | The proposed corpora significantly extends existing resources that are either not large enough or are restricted to a specific domain (such as health). |
A Very Low Resource Language Speech Corpus for Computational Language Documentation Experiments (L18-1)
Copied to clipboard
Pierre Godard, Gilles Adda, Martine Adda-Decker, Juan Benjumea, Laurent Besacier, Jamison Cooper-Leavitt, Guy-Noel Kouarata, Lori Lamel, Hélène Maynard, Markus Mueller, Annie Rialland, Sebastian Stueker, François Yvon, Marcely Zanon-Boito
| Challenge: | a new study aims to document endangered languages using a speech corpus . linguistic documentation is limited to the phonetic, lexical and syntactic levels . |
| Approach: | They propose to use a speech corpus to document endangered languages in field . they propose to collect 5k speech utterances aligned to French text translations . |
| Outcome: | The proposed language corpus is used to document endangered languages in field linguists . it is multilingual and contains 5k speech utterances aligned to french text translations - the authors show it can be used in a zero-resource task . |
A Recipe of Parallel Corpora Exploitation for Multilingual Large Language Models (2025.findings-naacl)
Copied to clipboard
| Challenge: | Recent studies have highlighted the potential of exploiting parallel corpora to enhance multilingual large language models. |
| Approach: | They investigate the impact of parallel corpora quality and quantity, training objectives, and model size on performance of multilingual large language models enhanced with parallel corporeal. |
| Outcome: | The proposed approach improves performance in bilingual and general-purpose tasks. |
Parallel Corpora for the Biomedical Domain (L18-1)
Copied to clipboard
| Challenge: | Existing corpora of parallel corporata are being used in the biomedical domain . MT is known to support readers' access to textual documents in a language other than their native language . |
| Approach: | They propose to leverage parallel corpora to implement cross-lingual information retrieval or machine translation tools. |
| Outcome: | The proposed corpus is being used in the biomedical task at the conference on machine translation (WMT'16 and WMT'17) it can be leveraged to provide access to health information in languages other than English. |
Augmenting Librispeech with French Translations: A Multimodal Corpus for Direct Speech Translation Evaluation (L18-1)
Copied to clipboard
| Challenge: | Recent work in spoken language translation (SLT) has attempted to build end-to-end speech-totext translation without using source language transcription during learning or decoding. |
| Approach: | They propose to augment an existing (monolingual) corpus: LibriSpeech. |
| Outcome: | The proposed corpus is derived from read audiobooks from the LibriVox project and has been carefully segmented and aligned. |
SwissSLi: The Multi-parallel Sign Language Corpus for Switzerland (2024.lrec-main)
Copied to clipboard
| Challenge: | Using a CC BY-NC-SA 4.0 license, this corpus contains parallel sign language videos and spoken language subtitles. |
| Approach: | They introduce SwissSLi, the first sign language corpus that contains parallel data of all three Swiss sign languages. |
| Outcome: | The proposed corpus contains parallel sign language videos and spoken language subtitles. |
Towards Building an Automatic Transcription System for Language Documentation: Experiences from Muyu (2020.lrec-1)
Copied to clipboard
| Challenge: | Language documentation is a rapidly growing field due to its urgency. |
| Approach: | They propose to use phoneme recognition to automatically recognize spoken languages and translate them to global languages. |
| Outcome: | The proposed tool performs better than existing methods with American English, Austrian German and Slovenian as source and target languages. |