Papers by Davis David

5 papers
MasakhaNER: Named Entity Recognition for African Languages (2021.tacl-1)

Copied to clipboard

Challenge: (2020) African languages are underrepresented in existing natural language processing datasets, research, and tools due to lack of datasets and reproducible results.
Approach: They propose to create a dataset for named entity recognition (NER) in ten African languages.
Outcome: The results of the first large dataset for named entity recognition (NER) in ten African languages are released to inform future research on African NLP.
AFRIDOC-MT: Document-level MT Corpus for African Languages (2025.emnlp-main)

Copied to clipboard

Challenge: AFRIDOC-MT is a document-level multi-parallel translation dataset covering five languages . AFRITIC-MT models perform better on sentences than general-purpose LLMs .
Approach: They propose a document-level multi-parallel translation dataset covering English and five African languages.
Outcome: The proposed dataset covers 334 health and 271 information technology news documents . it shows that NLLB-200 achieves the best average performance among standard models .
AfriSenti: A Twitter Sentiment Analysis Benchmark for African Languages (2023.emnlp-main)

Copied to clipboard

Challenge: Africa has the highest linguistic diversity among all continents.
Approach: They introduce a sentiment analysis benchmark that contains >110,000 tweets in 14 African languages . they describe the data collection methodology, annotation process, and challenges .
Outcome: The proposed dataset contains >110,000 tweets in 14 African languages . the tweets were annotated by native speakers and used in the shared task .
AfriMMT-EA: Multi-domain Machine Translation for Low-Resource East African Languages (2026.findings-eacl)

Copied to clipboard

Challenge: Recent advances in open-source large language models have demonstrated strong multilingual capabilities through data-efficient adaptation strategies.
Approach: They propose to use AfriMMT-EA to refine two multilingual versions of Gemma-3 to better understand the region's linguistic and cultural diversity.
Outcome: The proposed datasets comprise 54 local languages across five East African countries.
The Indigenous Languages Technology project at NRC Canada: An empowerment-oriented approach to developing language software (2020.coling-main)

Copied to clipboard

Challenge: This paper describes the first, three-year phase of a project at the National Research Council of Canada that is developing software to assist Indigenous communities in preserving their languages and extending their use.
Approach: They describe the first phase of a project at the National Research Council of Canada that is developing software to assist Indigenous communities in preserving their languages.
Outcome: The proposed software will help Indigenous communities preserve and revitalize their languages and extend their use.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations