Papers by Eric Atwell
Challenging the Transformer-based models with a Classical Arabic dataset: Quran and Hadith (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing benchmark datasets have a low readability index which does not reflect real-world complex data. |
| Approach: | They constructed a dataset of Quran-verse and Hadith-teaching pairs by consulting sources of reputable religious experts. |
| Outcome: | The proposed models performed on a binary classification task to identify whether two pieces of CA text convey the same underlying message. |
Constructing a Bilingual Hadith Corpus Using a Segmentation Tool (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing studies on Hadith have focused on the Quran, leaving it relatively unexplored. |
| Approach: | They propose to gather and construct a bilingual parallel corpus of Islamic Hadith using a custom segmentation tool that annotates the two Hadithe components with 92% accuracy. |
| Outcome: | The proposed method minimises the costs of language resource creation and produces consistent results independently from previous knowledge and experiences that usually influence human annotators. |
Web-based Annotation Tool for Inflectional Language Resources (L18-1)
Copied to clipboard
| Challenge: | Wasim is a web-based tool for semi-automatic morphosyntactic annotation of inflectional languages. |
| Approach: | They present a web-based tool for semi-automatic morphosyntactic annotation of inflectional languages resources. |
| Outcome: | The tool has high flexibility in segmenting tokens, editing, diacritizing, labelling tokens and segments. |