Challenge: lexicographical data is often difficult to find for less resourced languages . Xhosa is a popular language in south africa, but it is often suboptimal for many languages despite its multilingual nature .
Approach: They propose a new source of lexicographical data for Xhosa, a language spoken by 8 million speakers.
Outcome: The proposed model can be used in multilingual and federated environments and is extensible to other languages.

Similar Papers

LinguaMeta: Unified Metadata for Thousands of Languages (2024.lrec-main)

Copied to clipboard

Challenge: LinguaMeta is a unified repository of language metadata for thousands of languages.
Approach: They introduce LinguaMeta, a unified resource for language metadata for thousands of languages.
Outcome: The proposed resource is intended for use by researchers and organizations who aim to extend technology to thousands of languages.
TaTA: A Multilingual Table-to-Text Dataset for African Languages (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing data-to-text generation datasets are limited to English and a small number of other languages.
Approach: They create the first large multilingual table-to-text dataset with a focus on African languages.
Outcome: The proposed dataset includes 8,700 examples in nine languages including four African languages and a zero-shot test language.
AfriCLIRMatrix: Enabling Cross-Lingual Information Retrieval for African Languages (2022.emnlp-main)

Copied to clipboard

Challenge: Existing datasets for cross-lingual information retrieval are limited in many languages, especially those spoken in Africa.
Approach: They propose to build a test collection for cross-lingual information retrieval in 15 diverse African languages.
Outcome: AfriCLIRMatrix contains 6 million queries in English and 23 million relevance judgments automatically mined from Wikipedia inter-language links, covering many more African languages than any existing information retrieval test collection.
MultiLexBATS: Multilingual Dataset of Lexical Semantic Relations (2024.lrec-main)

Copied to clipboard

Challenge: Prior work has focused on analysing lexical semantic relations in word embeddings or probing pretrained language models (PLMs) with some exceptions.
Approach: They propose to use a multilingual parallel dataset of lexical semantic relations adapted from BATS in 15 languages including low-resource languages such as Bambara, Lithuanian, and Albanian as an experiment on cross-lingual transfer of relational knowledge.
Outcome: The proposed dataset is adapted from a BATS-based dataset in 15 languages including low-resource languages such as Bambara, Lithuanian, and Albanian.
GlossLM: A Massively Multilingual Corpus and Pretrained Model for Interlinear Glossed Text (2024.emnlp-main)

Copied to clipboard

Challenge: Existing resources for standardized, easily accessible IGT data limit their applicability to linguistic research.
Approach: They compile the largest existing corpus of interlinear glossed text data from a variety of sources and use it to generate annotated text.
Outcome: The proposed model outperforms SOTA models on monolingual corpora by 6.6%.
Writing System and Speaker Metadata for 2,800+ Language Varieties (2022.lrec-1)

Copied to clipboard

Challenge: Currently, language technologies are easily available in only a small minority of the world's 7,000+ language varieties.
Approach: They propose to use an open-source dataset to provide the writing system(s) for each of the 2,800+ languages used in the world today and an estimated speaker count for each.
Outcome: The dataset provides the attested writing system(s) for each of these 2,800+ varieties, as well as an estimated speaker count for each variety.
SERENGETI: Massively Multilingual Language Models for Africa (2023.findings-acl)

Copied to clipboard

Challenge: Pretrained models acquire valuable, generalizable linguistic information during pretraining and have advanced the state of the art on task-specific finetuning.
Approach: They develop a set of massively multilingual language models that covers 517 African languages and language varieties.
Outcome: The proposed models outperform 4 models that cover 4-23 African languages on eight natural language understanding tasks, achieving 82.27 average F_1.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
Charting the Landscape of African NLP: Mapping Progress and Shaping the Road Ahead (2025.emnlp-main)

Copied to clipboard

Challenge: African languages are often left behind in state-of-the-art natural language processing systems and large language models.
Approach: They analyze 884 research papers on NLP for African languages published over past five years . they identify key trends shaping the field and outline promising directions .
Outcome: The findings identify key trends shaping the field and outline promising directions . the authors analyze 884 research papers on NLP for African languages published over the past five years .
The IgboAPI Dataset: Empowering Igbo Language Technologies through Multi-dialectal Enrichment (2024.lrec-main)

Copied to clipboard

Challenge: UNESCO projects that the Igbo language will be endangered by 2025 . primary obstacle in developing dialectal-aware language technologies is lack of comprehensive dialectal datasets.
Approach: They propose to use a multi-dialectal Igbo-English dictionary dataset to enhance the representation of Igbe dialects.
Outcome: The proposed dataset enables machine translation systems to handle dialect variations in sentences.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations