Papers by Mahmoud El-Haj
Introducing the Welsh Text Summarisation Dataset and Baseline Systems (2022.lrec-1)
Copied to clipboard
| Challenge: | Welsh is an official language in Wales and is spoken by an estimated 884,300 people . historically, the language has been in decline and represents a minority language in the country despite having official status . |
| Approach: | They introduce the first Welsh summarisation dataset which is available to researchers as a free resource. |
| Outcome: | The proposed summarisation system will be used as a benchmark for summarisers in other minority language contexts. |
Profiling Medical Journal Articles Using a Gene Ontology Semantic Tagger (L18-1)
Copied to clipboard
| Challenge: | a growing number of scientific publications are based on sub-divisions and sub-communities of expertise becoming disconnected from each other. |
| Approach: | They propose to examine corpora derived from bodies of genetics literature and use it to make comparisons and improve retrieval methods. |
| Outcome: | The proposed methods will help to make comparisons and improve retrieval methods using domain knowledge via an existing gene ontology. |
Infrastructure for Semantic Annotation in the Genomics Domain (2020.lrec-1)
Copied to clipboard
Mahmoud El-Haj, Nathan Rutherford, Matthew Coole, Ignatius Ezeani, Sheryl Prentice, Nancy Ide, Jo Knight, Scott Piao, John Mariani, Paul Rayson, Keith Suderman
| Challenge: | a novel infrastructure for biomedical text mining combines NLP and corpus linguistics methods to provide a comprehensive corpus for literature-based discovery. |
| Approach: | They propose a novel pipeline for the collection, annotation, storage, retrieval and analysis of biomedical and life sciences literature . it uses an updatable Gene Ontology Semantic Tagger and a NLP pipeline scheduler to collect and process the corpus. |
| Outcome: | The proposed infrastructure allows for extreme-scale research on the open access PubMed Central archive. |
IgboBERT Models: Building and Training Transformer Models for the Igbo Language (2022.lrec-1)
Copied to clipboard
| Challenge: | This paper focuses on building resources for named entity recognition for Igbo, a language mainly spoken in the south eastern part of Nigeria. |
| Approach: | They present a standard Igbo named entity recognition dataset and results from fine-tuning transformer IgbeNER models. |
| Outcome: | The proposed dataset and model improves on the IgboNER task while training and fine-tuning a transformer model with comparatively little Igbe text data. |
Habibi - a multi Dialect multi National Arabic Song Lyrics Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Unlike western music, Arabic songs are poorly classified and the majority of the songs available online are classified under Modern Arabic Pop genre or what is now known as Franco-Arabic . |
| Approach: | They introduce Habibi the first Arabic Song Lyrics corpus for singers from 18 different Arabic countries. |
| Outcome: | The proposed corpus contains more than 30,000 Arabic song lyrics in 6 Arabic dialects for singers from 18 different arab countries. |
CoFiF Plus: A French Financial Narrative Summarisation Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing corpora for financial narrative summarisation exists in English but there is a significant lack of financial text resources in the French language. |
| Approach: | They propose to use natural language processing to analyse financial documents to find the best summarisation methods. |
| Outcome: | The proposed dataset is the first to provide a comprehensive set of financial text written in French. |
Arabic Dialect Identification in the Context of Bivalency and Code-Switching (L18-1)
Copied to clipboard
| Challenge: | Existing methods for identifying Arabic dialects require significant amounts of annotated training data which is costly and time consuming to produce. |
| Approach: | They propose a novel approach to Arabic dialect identification using language bivalency and written code-switching to identify Arabic dialects. |
| Outcome: | The proposed method can reach more than 76% and score well (66%) when tested on unseen data. |