Generating Monolingual Dataset for Low Resource Language Bodo from old books using Google Keep (2022.lrec-1)
Copied to clipboard
Sanjib Narzary, Maharaj Brahma, Mwnthai Narzary, Gwmsrang Muchahary, Pranav Kumar Singh, Apurbalal Senapati, Sukumar Nandi, Bidisha Som
| Challenge: | Bodo is a scheduled Indian language spoken largely by the Boda community in Assam and other northeastern Indian states. |
| Approach: | They propose to generate a monolingual Bodo corpus from different books using Google Keep for OCR. |
| Outcome: | The proposed method generates a monolingual Bodo corpus from different books using free, accessible, and daily-usable applications. |
Similar Papers
No Data to Crawl? Monolingual Corpus Creation from PDF Files of Truly low-Resource Languages in Peru (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing methods for extracting text from PDF files are expensive and limited by the absence of web content of endangered languages. |
| Approach: | They propose a method for creating monolingual corpora for four endangered languages . they use a PDF file format with multilingual sentences and noisy pages . |
| Outcome: | The proposed method allows the creation of clean corpora for the four languages, a key resource for natural language processing tasks nowadays. |
Researching Less-Resourced Languages – the DigiSami Corpus (L18-1)
Copied to clipboard
| Challenge: | DigiSami project aims to support research on endangered languages . it uses spoken corpus and speech technology for the Fenno-Ugric language North Sami . |
| Approach: | They describe the DigiSami project and its research results for the Fenno-Ugric language North Sami . they discuss ethical and privacy issues related to data collection for less-resourced languages and indigenous communities . |
| Outcome: | The DigiSami project focuses on spoken corpus collection and speech technology for the Fenno-Ugric language North Sami. |
A description and demonstration of SAFAR framework (2021.eacl-demos)
Copied to clipboard
Karim Bouzoubaa, Younes Jaafar, Driss Namly, Ridouane Tachicart, Rachida Tajmout, Hakima Khamar, Hamid Jaafar, Lhoussain Aouragh, Abdellah Yousfi
| Challenge: | Existing NLP infrastructures are naming them "toolkit", "platform" and "framework" authors present a monolingual framework dedicated to Arabic language . |
| Approach: | They propose a monolingual framework dedicated to Arabic language . they propose namings for existing infrastructures: "toolkit", "platform" and "framework" |
| Outcome: | The proposed framework is dedicated to Arabic language, especially the modern standard Arabic and Moroccan dialect. |
Language Model Priors and Data Augmentation Strategies for Low-resource Machine Translation: A Case Study Using Finnish to Northern Sámi (2024.findings-acl)
Copied to clipboard
| Challenge: | a new study examines the use of monolingual data for improving low-resource machine translation. |
| Approach: | They investigate ways of using monolingual data for improving low-resource machine translation. |
| Outcome: | The proposed model can perform better on the target-side data without augmentation of parallel data. |
GlotLID: Language Identification for Low-Resource Languages (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing web-mined datasets for low-resource languages have been useful for low resource NLP. |
| Approach: | They propose a model that identifies 1665 low-resource languages and a new model that is rigorously evaluated and reliable. |
| Outcome: | The proposed model outperforms baselines when balancing F1 and false positive rate (FPR). |
Building a User-Generated Content North-African Arabizi Treebank: Tackling Hell (2020.acl-main)
Copied to clipboard
Djamé Seddah, Farah Essaidi, Amal Fethi, Matthieu Futeral, Benjamin Muller, Pedro Javier Ortiz Suárez, Benoît Sagot, Abhishek Srivastava
| Challenge: | a treebank for a north-African Arabic dialect known for code-switching is made freely available . authors: geopolitical events are a factor highlighting a language deficiency in terms of natural language processing resources . |
| Approach: | They propose to make a treebank for a romanized user-generated content variety of Algerian . they supplement it with 50k unlabeled sentences from common crawl and web-crawled data . |
| Outcome: | The proposed treebank is made of 1500 sentences, fully annotated in morpho-syntax and universal dependency syntax, with full translation at both the word and sentence levels. |
OCR Improves Machine Translation for Low-Resource Languages (2022.findings-acl)
Copied to clipboard
| Challenge: | Despite many recent successes, Machine Translation still lacks support or fails to achieve good performance for most low-resource languages. |
| Approach: | They propose a benchmark to evaluate OCR systems on low-resource languages and low- resource scripts. |
| Outcome: | The proposed benchmark evaluates state-of-the-art OCR systems on low-resource languages and low-rural scripts. |
BalsuTalka.lv - Boosting the Common Voice Corpus for Low-Resource Languages (2024.lrec-main)
Copied to clipboard
Roberts Dargis, Arturs Znotins, Ilze Auzina, Baiba Saulite, Sanita Reinsone, Raivis Dejus, Antra Klavinska, Normunds Gruzitis
| Challenge: | Latvian is a low-resource language for many NLP tasks, but most speech corpora are closed data . a crowdsourcing campaign to create a relatively large, diverse and open speech corpus for Latvian has been launched . |
| Approach: | a crowdsourcing campaign is helping to create an open speech corpus for Latvian . the goal is to enlarge the datasets and make them more diverse . authors use the opensource Mozilla Common Voice platform to validate speech samples . |
| Outcome: | a crowdsourcing initiative has increased the size and speaker diversity of the Latvian Common Voice 17.0 dataset by more than tenfold in less than a year. |
Multilingual Dependency Parsing for Low-Resource Languages: Case Studies on North Saami and Komi-Zyrian (L18-1)
Copied to clipboard
| Challenge: | Developing systems for low-resource languages is a crucial issue for Natural Language Processing (NLP). |
| Approach: | They propose a method for parsing low-resource languages with very small training corpora using multilingual word embeddings and annotated corporata of larger languages. |
| Outcome: | The proposed method improves dependency parsing for low-resource languages with very small training corpora compared to previous work . it also explores whether contemporary contact languages or genetically related languages would be the most fruitful starting point for multilingual parsers. |
BDPROTO: A Database of Phonological Inventories from Ancient and Reconstructed Languages (L18-1)
Copied to clipboard
| Challenge: | BDPROTO is a database of phonological inventory data from 137 ancient and reconstructed languages. |
| Approach: | They present a database of phonological inventory data from 137 ancient and reconstructed languages. |
| Outcome: | The BDPROTO database is a publicly available, unicode-compliant resource . it contains phonological inventory data from 137 ancient and reconstructed languages . |