BULBasaa: A Bilingual Basaa-French Speech Corpus for the Evaluation of Language Documentation Tools (L18-1)
Copied to clipboard
Fatima Hamlaoui, Emmanuel-Moselly Makasso, Markus Müller, Jonas Engelmann, Gilles Adda, Alex Waibel, Sebastian Stüker
| Challenge: | Approximately 50 hours of Bàsàá speech were collected and then carefully re-spoken and orally translated into French . |
| Approach: | They propose to provide an automatic phonetic transcription using a set of derived phone-like units. |
| Outcome: | The proposed method provides an automatic phonetic transcription using a set of derived phone-like units. |
Similar Papers
Parallel Corpora in Mboshi (Bantu C25, Congo-Brazzaville) (L18-1)
Copied to clipboard
Annie Rialland, Martine Adda-Decker, Guy-Noël Kouarata, Gilles Adda, Laurent Besacier, Lori Lamel, Elodie Gauthier, Pierre Godard, Jamison Cooper-Leavitt
| Challenge: | BULB project aims to provide tools to language documentation and description for unwritten languages . language-based technologies are needed to support the collection of data and to provide linguistic documentation for the languages. |
| Approach: | This paper presents multimodal and parallel data collections in Mboshi, as part of the French-German BULB project. |
| Outcome: | The proposed data collection includes pictures and videos documenting social practices, agriculture, wildlife and plants. |
BasahaCorpus: An Expanded Linguistic Resource for Readability Assessment in Central Philippine Languages (2023.emnlp-main)
Copied to clipboard
| Challenge: | Current research on automatic readability assessment (ARA) has focused on improving the performance of models in high-resource languages such as English. |
| Approach: | They propose a hierarchical cross-lingual modeling approach that takes advantage of a language’s placement in the family tree to increase the amount of available training data. |
| Outcome: | The proposed model improves the performance of models in high-resource languages such as English and Hiligaynon, minasbate, Karay-a, and Rinconada. |
ZIPA: A family of efficient models for multilingual phone recognition (2025.acl-long)
Copied to clipboard
| Challenge: | IPA transcriptions capture major articulatory contrasts in speech sounds, including the voicing status, place of articulation, manner of voicing, and tongue positions. |
| Approach: | They present ZIPA, a family of efficient speech models that advances the state-of-the-art performance of crosslinguistic phone recognition. |
| Outcome: | The proposed model outperforms existing phone recognition systems on 17,000+ hours of normalized phone transcriptions and a novel evaluation set capturing unseen languages and sociophonetic variation. |
Beyond Atomic Characters: Glyph-Aware Sub-character Alignment for Low-Resource Multilingual OCR (2026.acl-long)
Copied to clipboard
| Challenge: | Low-resource multilingual OCR models struggle with complex script structures and data scarcity. |
| Approach: | They propose a framework for multilingual character recognition that integrates visual and linguistic backbones with a novel glyph-aware interface. |
| Outcome: | The proposed framework improves on high-resolution visual and language backbones with glyph-aware interface. |
A Very Low Resource Language Speech Corpus for Computational Language Documentation Experiments (L18-1)
Copied to clipboard
Pierre Godard, Gilles Adda, Martine Adda-Decker, Juan Benjumea, Laurent Besacier, Jamison Cooper-Leavitt, Guy-Noel Kouarata, Lori Lamel, Hélène Maynard, Markus Mueller, Annie Rialland, Sebastian Stueker, François Yvon, Marcely Zanon-Boito
| Challenge: | a new study aims to document endangered languages using a speech corpus . linguistic documentation is limited to the phonetic, lexical and syntactic levels . |
| Approach: | They propose to use a speech corpus to document endangered languages in field . they propose to collect 5k speech utterances aligned to French text translations . |
| Outcome: | The proposed language corpus is used to document endangered languages in field linguists . it is multilingual and contains 5k speech utterances aligned to french text translations - the authors show it can be used in a zero-resource task . |
ViNLI: A Vietnamese Corpus for Studies on Open-Domain Natural Language Inference (2022.coling-1)
Copied to clipboard
| Challenge: | a large-scale corpus is needed for studies on natural language inference (NLI) for Vietnamese, which can be considered a low-resource language. |
| Approach: | They propose a corpus for evaluating Vietnamese natural language inference models . they use a human-annotated corpus extracted from more than 800 online news articles . |
| Outcome: | The ViNLI corpus is created and evaluated with a strict process of quality control . the best system performance is still far from human performance (a 14.20% gap in accuracy). |
Batayan: A Filipino NLP benchmark for evaluating Large Language Models (2025.acl-long)
Copied to clipboard
Jann Railey Montalan, Jimson Paulo Layacan, David Demitri Africa, Richell Isaiah S. Flores, Michael T. Lopez Ii, Theresa Denise Magsajo, Anjanette Cayabyab, William Chandra Tjhi
| Challenge: | Recent advances in large language models (LLMs) have demonstrated remarkable capabilities on widely benchmarked high-resource languages. |
| Approach: | They propose a benchmark that systematically evaluates LLMs across three key natural language processing competencies: understanding, reasoning, and generation. |
| Outcome: | The proposed benchmark covers eight tasks covering Tagalog and code-switched Taglish utterances. |
Unmasking the Myth of Effortless Big Data - Making an Open Source Multi-lingual Infrastructure and Building Language Resources from Scratch (2022.lrec-1)
Copied to clipboard
Linda Wiechetek, Katri Hiovain-Asikainen, Inga Lill Sigga Mikkelsen, Sjur Moshagen, Flammie Pirinen, Trond Trosterud, Børre Gaup
| Challenge: | During the last two decades, machine learning approaches have dominated the field of natural language processing (NLP) weak literary traditions give rise to corpora too unreliable to function as a model for NLP tools. |
| Approach: | They propose an alternative to corpus-based language technology that can provide language technology solutions for minority languages. |
| Outcome: | The proposed approach can provide language technology solutions for minority languages outside the reach of corpus-based language technology. |
BaSCo: An Annotated Basque-Spanish Code-Switching Corpus for Natural Language Understanding (2022.lrec-1)
Copied to clipboard
| Challenge: | Basque-Spanish code-switching is a widespread phenomenon among bilingual speakers in the Basque Country. |
| Approach: | They propose to use annotated utterances to train bilingual chatbots in Basque and Spanish to cover the phenomenon of code-switching. |
| Outcome: | The proposed corpus is the first with annotated linguistic resources encompassing Basque-Spanish code-switching. |
Spoken Language Treebanks in Universal Dependencies: an Overview (2022.lrec-1)
Copied to clipboard
| Challenge: | spoken language treebanks have divergent annotation schemes limiting cross-resource explorations . many spoken language trees have no written form, but many of the world languages have no spoken form at all. |
| Approach: | They propose to use the Universal Dependencies annotation scheme to annotate spoken language treebanks using a morphosyntactic annotation scheme. |
| Outcome: | The proposed treebanks differ significantly with respect to the inventory and format of transcribed phenomena and the principles adopted in their morphosyntactic annotation. |