Challenge: Bodo is a scheduled Indian language spoken largely by the Boda community in Assam and other northeastern Indian states.
Approach: They propose to generate a monolingual Bodo corpus from different books using Google Keep for OCR.
Outcome: The proposed method generates a monolingual Bodo corpus from different books using free, accessible, and daily-usable applications.

Similar Papers

No Data to Crawl? Monolingual Corpus Creation from PDF Files of Truly low-Resource Languages in Peru (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods for extracting text from PDF files are expensive and limited by the absence of web content of endangered languages.
Approach: They propose a method for creating monolingual corpora for four endangered languages . they use a PDF file format with multilingual sentences and noisy pages .
Outcome: The proposed method allows the creation of clean corpora for the four languages, a key resource for natural language processing tasks nowadays.
Researching Less-Resourced Languages – the DigiSami Corpus (L18-1)

Copied to clipboard

Challenge: DigiSami project aims to support research on endangered languages . it uses spoken corpus and speech technology for the Fenno-Ugric language North Sami .
Approach: They describe the DigiSami project and its research results for the Fenno-Ugric language North Sami . they discuss ethical and privacy issues related to data collection for less-resourced languages and indigenous communities .
Outcome: The DigiSami project focuses on spoken corpus collection and speech technology for the Fenno-Ugric language North Sami.
A description and demonstration of SAFAR framework (2021.eacl-demos)

Copied to clipboard

Challenge: Existing NLP infrastructures are naming them "toolkit", "platform" and "framework" authors present a monolingual framework dedicated to Arabic language .
Approach: They propose a monolingual framework dedicated to Arabic language . they propose namings for existing infrastructures: "toolkit", "platform" and "framework"
Outcome: The proposed framework is dedicated to Arabic language, especially the modern standard Arabic and Moroccan dialect.
Language Model Priors and Data Augmentation Strategies for Low-resource Machine Translation: A Case Study Using Finnish to Northern Sámi (2024.findings-acl)

Copied to clipboard

Challenge: a new study examines the use of monolingual data for improving low-resource machine translation.
Approach: They investigate ways of using monolingual data for improving low-resource machine translation.
Outcome: The proposed model can perform better on the target-side data without augmentation of parallel data.
GlotLID: Language Identification for Low-Resource Languages (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing web-mined datasets for low-resource languages have been useful for low resource NLP.
Approach: They propose a model that identifies 1665 low-resource languages and a new model that is rigorously evaluated and reliable.
Outcome: The proposed model outperforms baselines when balancing F1 and false positive rate (FPR).
Building a User-Generated Content North-African Arabizi Treebank: Tackling Hell (2020.acl-main)

Copied to clipboard

Challenge: a treebank for a north-African Arabic dialect known for code-switching is made freely available . authors: geopolitical events are a factor highlighting a language deficiency in terms of natural language processing resources .
Approach: They propose to make a treebank for a romanized user-generated content variety of Algerian . they supplement it with 50k unlabeled sentences from common crawl and web-crawled data .
Outcome: The proposed treebank is made of 1500 sentences, fully annotated in morpho-syntax and universal dependency syntax, with full translation at both the word and sentence levels.
OCR Improves Machine Translation for Low-Resource Languages (2022.findings-acl)

Copied to clipboard

Challenge: Despite many recent successes, Machine Translation still lacks support or fails to achieve good performance for most low-resource languages.
Approach: They propose a benchmark to evaluate OCR systems on low-resource languages and low- resource scripts.
Outcome: The proposed benchmark evaluates state-of-the-art OCR systems on low-resource languages and low-rural scripts.
BalsuTalka.lv - Boosting the Common Voice Corpus for Low-Resource Languages (2024.lrec-main)

Copied to clipboard

Challenge: Latvian is a low-resource language for many NLP tasks, but most speech corpora are closed data . a crowdsourcing campaign to create a relatively large, diverse and open speech corpus for Latvian has been launched .
Approach: a crowdsourcing campaign is helping to create an open speech corpus for Latvian . the goal is to enlarge the datasets and make them more diverse . authors use the opensource Mozilla Common Voice platform to validate speech samples .
Outcome: a crowdsourcing initiative has increased the size and speaker diversity of the Latvian Common Voice 17.0 dataset by more than tenfold in less than a year.
Multilingual Dependency Parsing for Low-Resource Languages: Case Studies on North Saami and Komi-Zyrian (L18-1)

Copied to clipboard

Challenge: Developing systems for low-resource languages is a crucial issue for Natural Language Processing (NLP).
Approach: They propose a method for parsing low-resource languages with very small training corpora using multilingual word embeddings and annotated corporata of larger languages.
Outcome: The proposed method improves dependency parsing for low-resource languages with very small training corpora compared to previous work . it also explores whether contemporary contact languages or genetically related languages would be the most fruitful starting point for multilingual parsers.
BDPROTO: A Database of Phonological Inventories from Ancient and Reconstructed Languages (L18-1)

Copied to clipboard

Challenge: BDPROTO is a database of phonological inventory data from 137 ancient and reconstructed languages.
Approach: They present a database of phonological inventory data from 137 ancient and reconstructed languages.
Outcome: The BDPROTO database is a publicly available, unicode-compliant resource . it contains phonological inventory data from 137 ancient and reconstructed languages .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations