Preparation and Usage of Xhosa Lexicographical Data for a Multilingual, Federated Environment (L18-1)
Copied to clipboard
| Challenge: | lexicographical data is often difficult to find for less resourced languages . Xhosa is a popular language in south africa, but it is often suboptimal for many languages despite its multilingual nature . |
| Approach: | They propose a new source of lexicographical data for Xhosa, a language spoken by 8 million speakers. |
| Outcome: | The proposed model can be used in multilingual and federated environments and is extensible to other languages. |
Similar Papers
LinguaMeta: Unified Metadata for Thousands of Languages (2024.lrec-main)
Copied to clipboard
| Challenge: | LinguaMeta is a unified repository of language metadata for thousands of languages. |
| Approach: | They introduce LinguaMeta, a unified resource for language metadata for thousands of languages. |
| Outcome: | The proposed resource is intended for use by researchers and organizations who aim to extend technology to thousands of languages. |
TaTA: A Multilingual Table-to-Text Dataset for African Languages (2023.findings-emnlp)
Copied to clipboard
Sebastian Gehrmann, Sebastian Ruder, Vitaly Nikolaev, Jan Botha, Michael Chavinda, Ankur Parikh, Clara Rivera
| Challenge: | Existing data-to-text generation datasets are limited to English and a small number of other languages. |
| Approach: | They create the first large multilingual table-to-text dataset with a focus on African languages. |
| Outcome: | The proposed dataset includes 8,700 examples in nine languages including four African languages and a zero-shot test language. |
AfriCLIRMatrix: Enabling Cross-Lingual Information Retrieval for African Languages (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing datasets for cross-lingual information retrieval are limited in many languages, especially those spoken in Africa. |
| Approach: | They propose to build a test collection for cross-lingual information retrieval in 15 diverse African languages. |
| Outcome: | AfriCLIRMatrix contains 6 million queries in English and 23 million relevance judgments automatically mined from Wikipedia inter-language links, covering many more African languages than any existing information retrieval test collection. |
MultiLexBATS: Multilingual Dataset of Lexical Semantic Relations (2024.lrec-main)
Copied to clipboard
Dagmar Gromann, Hugo Goncalo Oliveira, Lucia Pitarch, Elena-Simona Apostol, Jordi Bernad, Eliot Bytyçi, Chiara Cantone, Sara Carvalho, Francesca Frontini, Radovan Garabik, Jorge Gracia, Letizia Granata, Fahad Khan, Timotej Knez, Penny Labropoulou, Chaya Liebeskind, Maria Pia Di Buono, Ana Ostroški Anić, Sigita Rackevičienė, Ricardo Rodrigues, Gilles Sérasset, Linas Selmistraitis, Mahammadou Sidibé, Purificação Silvano, Blerina Spahiu, Enriketa Sogutlu, Ranka Stanković, Ciprian-Octavian Truică, Giedre Valunaite Oleskeviciene, Slavko Zitnik, Katerina Zdravkova
| Challenge: | Prior work has focused on analysing lexical semantic relations in word embeddings or probing pretrained language models (PLMs) with some exceptions. |
| Approach: | They propose to use a multilingual parallel dataset of lexical semantic relations adapted from BATS in 15 languages including low-resource languages such as Bambara, Lithuanian, and Albanian as an experiment on cross-lingual transfer of relational knowledge. |
| Outcome: | The proposed dataset is adapted from a BATS-based dataset in 15 languages including low-resource languages such as Bambara, Lithuanian, and Albanian. |
GlossLM: A Massively Multilingual Corpus and Pretrained Model for Interlinear Glossed Text (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing resources for standardized, easily accessible IGT data limit their applicability to linguistic research. |
| Approach: | They compile the largest existing corpus of interlinear glossed text data from a variety of sources and use it to generate annotated text. |
| Outcome: | The proposed model outperforms SOTA models on monolingual corpora by 6.6%. |
Writing System and Speaker Metadata for 2,800+ Language Varieties (2022.lrec-1)
Copied to clipboard
| Challenge: | Currently, language technologies are easily available in only a small minority of the world's 7,000+ language varieties. |
| Approach: | They propose to use an open-source dataset to provide the writing system(s) for each of the 2,800+ languages used in the world today and an estimated speaker count for each. |
| Outcome: | The dataset provides the attested writing system(s) for each of these 2,800+ varieties, as well as an estimated speaker count for each variety. |
SERENGETI: Massively Multilingual Language Models for Africa (2023.findings-acl)
Copied to clipboard
| Challenge: | Pretrained models acquire valuable, generalizable linguistic information during pretraining and have advanced the state of the art on task-specific finetuning. |
| Approach: | They develop a set of massively multilingual language models that covers 517 African languages and language varieties. |
| Outcome: | The proposed models outperform 4 models that cover 4-23 African languages on eight natural language understanding tasks, achieving 82.27 average F_1. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
Charting the Landscape of African NLP: Mapping Progress and Shaping the Road Ahead (2025.emnlp-main)
Copied to clipboard
| Challenge: | African languages are often left behind in state-of-the-art natural language processing systems and large language models. |
| Approach: | They analyze 884 research papers on NLP for African languages published over past five years . they identify key trends shaping the field and outline promising directions . |
| Outcome: | The findings identify key trends shaping the field and outline promising directions . the authors analyze 884 research papers on NLP for African languages published over the past five years . |
The IgboAPI Dataset: Empowering Igbo Language Technologies through Multi-dialectal Enrichment (2024.lrec-main)
Copied to clipboard
Chris Chinenye Emezue, Ifeoma Okoh, Chinedu Emmanuel Mbonu, Chiamaka Chukwuneke, Daisy Monika Lal, Ignatius Ezeani, Paul Rayson, Ijemma Onwuzulike, Chukwuma Onyebuchi Okeke, Gerald Okey Nweya, Bright Ikechukwu Ogbonna, Chukwuebuka Uchenna Oraegbunam, Esther Chidinma Awo-Ndubuisi, Akudo Amarachukwu Osuagwu
| Challenge: | UNESCO projects that the Igbo language will be endangered by 2025 . primary obstacle in developing dialectal-aware language technologies is lack of comprehensive dialectal datasets. |
| Approach: | They propose to use a multi-dialectal Igbo-English dictionary dataset to enhance the representation of Igbe dialects. |
| Outcome: | The proposed dataset enables machine translation systems to handle dialect variations in sentences. |