Papers by Mahmoud El-Haj

7 papers
Introducing the Welsh Text Summarisation Dataset and Baseline Systems (2022.lrec-1)

Copied to clipboard

Challenge: Welsh is an official language in Wales and is spoken by an estimated 884,300 people . historically, the language has been in decline and represents a minority language in the country despite having official status .
Approach: They introduce the first Welsh summarisation dataset which is available to researchers as a free resource.
Outcome: The proposed summarisation system will be used as a benchmark for summarisers in other minority language contexts.
Profiling Medical Journal Articles Using a Gene Ontology Semantic Tagger (L18-1)

Copied to clipboard

Challenge: a growing number of scientific publications are based on sub-divisions and sub-communities of expertise becoming disconnected from each other.
Approach: They propose to examine corpora derived from bodies of genetics literature and use it to make comparisons and improve retrieval methods.
Outcome: The proposed methods will help to make comparisons and improve retrieval methods using domain knowledge via an existing gene ontology.
Infrastructure for Semantic Annotation in the Genomics Domain (2020.lrec-1)

Copied to clipboard

Challenge: a novel infrastructure for biomedical text mining combines NLP and corpus linguistics methods to provide a comprehensive corpus for literature-based discovery.
Approach: They propose a novel pipeline for the collection, annotation, storage, retrieval and analysis of biomedical and life sciences literature . it uses an updatable Gene Ontology Semantic Tagger and a NLP pipeline scheduler to collect and process the corpus.
Outcome: The proposed infrastructure allows for extreme-scale research on the open access PubMed Central archive.
IgboBERT Models: Building and Training Transformer Models for the Igbo Language (2022.lrec-1)

Copied to clipboard

Challenge: This paper focuses on building resources for named entity recognition for Igbo, a language mainly spoken in the south eastern part of Nigeria.
Approach: They present a standard Igbo named entity recognition dataset and results from fine-tuning transformer IgbeNER models.
Outcome: The proposed dataset and model improves on the IgboNER task while training and fine-tuning a transformer model with comparatively little Igbe text data.
Habibi - a multi Dialect multi National Arabic Song Lyrics Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Unlike western music, Arabic songs are poorly classified and the majority of the songs available online are classified under Modern Arabic Pop genre or what is now known as Franco-Arabic .
Approach: They introduce Habibi the first Arabic Song Lyrics corpus for singers from 18 different Arabic countries.
Outcome: The proposed corpus contains more than 30,000 Arabic song lyrics in 6 Arabic dialects for singers from 18 different arab countries.
CoFiF Plus: A French Financial Narrative Summarisation Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Existing corpora for financial narrative summarisation exists in English but there is a significant lack of financial text resources in the French language.
Approach: They propose to use natural language processing to analyse financial documents to find the best summarisation methods.
Outcome: The proposed dataset is the first to provide a comprehensive set of financial text written in French.
Arabic Dialect Identification in the Context of Bivalency and Code-Switching (L18-1)

Copied to clipboard

Challenge: Existing methods for identifying Arabic dialects require significant amounts of annotated training data which is costly and time consuming to produce.
Approach: They propose a novel approach to Arabic dialect identification using language bivalency and written code-switching to identify Arabic dialects.
Outcome: The proposed method can reach more than 76% and score well (66%) when tested on unseen data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations