Papers by Ignatius Ezeani

10 papers
MasakhaNER: Named Entity Recognition for African Languages (2021.tacl-1)

Copied to clipboard

Challenge: (2020) African languages are underrepresented in existing natural language processing datasets, research, and tools due to lack of datasets and reproducible results.
Approach: They propose to create a dataset for named entity recognition (NER) in ten African languages.
Outcome: The results of the first large dataset for named entity recognition (NER) in ten African languages are released to inform future research on African NLP.
The IgboAPI Dataset: Empowering Igbo Language Technologies through Multi-dialectal Enrichment (2024.lrec-main)

Copied to clipboard

Challenge: UNESCO projects that the Igbo language will be endangered by 2025 . primary obstacle in developing dialectal-aware language technologies is lack of comprehensive dialectal datasets.
Approach: They propose to use a multi-dialectal Igbo-English dictionary dataset to enhance the representation of Igbe dialects.
Outcome: The proposed dataset enables machine translation systems to handle dialect variations in sentences.
Introducing the Welsh Text Summarisation Dataset and Baseline Systems (2022.lrec-1)

Copied to clipboard

Challenge: Welsh is an official language in Wales and is spoken by an estimated 884,300 people . historically, the language has been in decline and represents a minority language in the country despite having official status .
Approach: They introduce the first Welsh summarisation dataset which is available to researchers as a free resource.
Outcome: The proposed summarisation system will be used as a benchmark for summarisers in other minority language contexts.
Infrastructure for Semantic Annotation in the Genomics Domain (2020.lrec-1)

Copied to clipboard

Challenge: a novel infrastructure for biomedical text mining combines NLP and corpus linguistics methods to provide a comprehensive corpus for literature-based discovery.
Approach: They propose a novel pipeline for the collection, annotation, storage, retrieval and analysis of biomedical and life sciences literature . it uses an updatable Gene Ontology Semantic Tagger and a NLP pipeline scheduler to collect and process the corpus.
Outcome: The proposed infrastructure allows for extreme-scale research on the open access PubMed Central archive.
IgboBERT Models: Building and Training Transformer Models for the Igbo Language (2022.lrec-1)

Copied to clipboard

Challenge: This paper focuses on building resources for named entity recognition for Igbo, a language mainly spoken in the south eastern part of Nigeria.
Approach: They present a standard Igbo named entity recognition dataset and results from fine-tuning transformer IgbeNER models.
Outcome: The proposed dataset and model improves on the IgboNER task while training and fine-tuning a transformer model with comparatively little Igbe text data.
LENS: Learning Entities from Narratives of Skin Cancer (2025.coling-demos)

Copied to clipboard

Challenge: Learning entities from narratives of skin cancer (LENS) is an automatic entity recognition system built on colloquial writings from skin cancer-related forums.
Approach: They propose to use reddit forums to create an automatic entity recognition system that can be used to predict skin cancer outcomes.
Outcome: LENS achieves an overall entity-level F1 score of 0.561 . other notable results include “CANC_T” (0.747), “STG” (0.888), “POB” (0.914), “GENDER” (0.750), “A/G” (00.646), “EMO” (0.619), and “MHD” (0.503).
Igbo Diacritic Restoration using Embedding Models (N18-4)

Copied to clipboard

Challenge: Igbo is a low-resource language spoken by approximately 30 million people worldwide.
Approach: They propose to use word embeddings to restore diacritics in Igbo by using a pre-processing task that replaces missing diacrittics on words from which they have been removed.
Outcome: The embedding models performed better than n-gram models on the diacritic restoration task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations