Papers by Jade Abbott

4 papers
MasakhaNER: Named Entity Recognition for African Languages (2021.tacl-1)

Copied to clipboard

Challenge: (2020) African languages are underrepresented in existing natural language processing datasets, research, and tools due to lack of datasets and reproducible results.
Approach: They propose to create a dataset for named entity recognition (NER) in ten African languages.
Outcome: The results of the first large dataset for named entity recognition (NER) in ten African languages are released to inform future research on African NLP.
The Esethu Framework: Reimagining Sustainable Dataset Governance and Curation for Low-Resource Languages (2025.acl-long)

Copied to clipboard

Challenge: Esethu Framework is a community-centric data license that empowers local communities and ensures equitable benefit-sharing from their linguistic resource.
Approach: They propose a community-centric data license to empower local communities and ensure equitable benefit-sharing from their linguistic resource.
Outcome: The proposed dataset contains read speech from native isiXhosa speakers enriched with demographic and linguistic metadata.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations