Challenge: Approximately 50 hours of Bàsàá speech were collected and then carefully re-spoken and orally translated into French .
Approach: They propose to provide an automatic phonetic transcription using a set of derived phone-like units.
Outcome: The proposed method provides an automatic phonetic transcription using a set of derived phone-like units.

Similar Papers

Parallel Corpora in Mboshi (Bantu C25, Congo-Brazzaville) (L18-1)

Copied to clipboard

Challenge: BULB project aims to provide tools to language documentation and description for unwritten languages . language-based technologies are needed to support the collection of data and to provide linguistic documentation for the languages.
Approach: This paper presents multimodal and parallel data collections in Mboshi, as part of the French-German BULB project.
Outcome: The proposed data collection includes pictures and videos documenting social practices, agriculture, wildlife and plants.
BasahaCorpus: An Expanded Linguistic Resource for Readability Assessment in Central Philippine Languages (2023.emnlp-main)

Copied to clipboard

Challenge: Current research on automatic readability assessment (ARA) has focused on improving the performance of models in high-resource languages such as English.
Approach: They propose a hierarchical cross-lingual modeling approach that takes advantage of a language’s placement in the family tree to increase the amount of available training data.
Outcome: The proposed model improves the performance of models in high-resource languages such as English and Hiligaynon, minasbate, Karay-a, and Rinconada.
ZIPA: A family of efficient models for multilingual phone recognition (2025.acl-long)

Copied to clipboard

Challenge: IPA transcriptions capture major articulatory contrasts in speech sounds, including the voicing status, place of articulation, manner of voicing, and tongue positions.
Approach: They present ZIPA, a family of efficient speech models that advances the state-of-the-art performance of crosslinguistic phone recognition.
Outcome: The proposed model outperforms existing phone recognition systems on 17,000+ hours of normalized phone transcriptions and a novel evaluation set capturing unseen languages and sociophonetic variation.
Beyond Atomic Characters: Glyph-Aware Sub-character Alignment for Low-Resource Multilingual OCR (2026.acl-long)

Copied to clipboard

Challenge: Low-resource multilingual OCR models struggle with complex script structures and data scarcity.
Approach: They propose a framework for multilingual character recognition that integrates visual and linguistic backbones with a novel glyph-aware interface.
Outcome: The proposed framework improves on high-resolution visual and language backbones with glyph-aware interface.
A Very Low Resource Language Speech Corpus for Computational Language Documentation Experiments (L18-1)

Copied to clipboard

Challenge: a new study aims to document endangered languages using a speech corpus . linguistic documentation is limited to the phonetic, lexical and syntactic levels .
Approach: They propose to use a speech corpus to document endangered languages in field . they propose to collect 5k speech utterances aligned to French text translations .
Outcome: The proposed language corpus is used to document endangered languages in field linguists . it is multilingual and contains 5k speech utterances aligned to french text translations - the authors show it can be used in a zero-resource task .
ViNLI: A Vietnamese Corpus for Studies on Open-Domain Natural Language Inference (2022.coling-1)

Copied to clipboard

Challenge: a large-scale corpus is needed for studies on natural language inference (NLI) for Vietnamese, which can be considered a low-resource language.
Approach: They propose a corpus for evaluating Vietnamese natural language inference models . they use a human-annotated corpus extracted from more than 800 online news articles .
Outcome: The ViNLI corpus is created and evaluated with a strict process of quality control . the best system performance is still far from human performance (a 14.20% gap in accuracy).
Batayan: A Filipino NLP benchmark for evaluating Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have demonstrated remarkable capabilities on widely benchmarked high-resource languages.
Approach: They propose a benchmark that systematically evaluates LLMs across three key natural language processing competencies: understanding, reasoning, and generation.
Outcome: The proposed benchmark covers eight tasks covering Tagalog and code-switched Taglish utterances.
Unmasking the Myth of Effortless Big Data - Making an Open Source Multi-lingual Infrastructure and Building Language Resources from Scratch (2022.lrec-1)

Copied to clipboard

Challenge: During the last two decades, machine learning approaches have dominated the field of natural language processing (NLP) weak literary traditions give rise to corpora too unreliable to function as a model for NLP tools.
Approach: They propose an alternative to corpus-based language technology that can provide language technology solutions for minority languages.
Outcome: The proposed approach can provide language technology solutions for minority languages outside the reach of corpus-based language technology.
BaSCo: An Annotated Basque-Spanish Code-Switching Corpus for Natural Language Understanding (2022.lrec-1)

Copied to clipboard

Challenge: Basque-Spanish code-switching is a widespread phenomenon among bilingual speakers in the Basque Country.
Approach: They propose to use annotated utterances to train bilingual chatbots in Basque and Spanish to cover the phenomenon of code-switching.
Outcome: The proposed corpus is the first with annotated linguistic resources encompassing Basque-Spanish code-switching.
Spoken Language Treebanks in Universal Dependencies: an Overview (2022.lrec-1)

Copied to clipboard

Challenge: spoken language treebanks have divergent annotation schemes limiting cross-resource explorations . many spoken language trees have no written form, but many of the world languages have no spoken form at all.
Approach: They propose to use the Universal Dependencies annotation scheme to annotate spoken language treebanks using a morphosyntactic annotation scheme.
Outcome: The proposed treebanks differ significantly with respect to the inventory and format of transcribed phenomena and the principles adopted in their morphosyntactic annotation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations