Papers by Fahad Khan
Arabic Mini-ClimateGPT : A Climate Change and Sustainability Tailored Arabic LLM (2023.findings-emnlp)
Copied to clipboard
Sahal Mullappilly, Abdelrahman Shaker, Omkar Thawakar, Hisham Cholakkal, Rao Anwer, Salman Khan, Fahad Khan
| Challenge: | Recent large language models like ChatGPT and Bard excel in a wide variety of NLP tasks but are not specifically tailored for climate related domain specific information. |
| Approach: | They propose a lightweight Arabic Mini-ClimateGPT that is built on an open-source LLM and specifically fine-tuned on a conversational-style instruction tuning curated Arabic dataset Clima500-Instruct. |
| Outcome: | The proposed model surpasses the baseline LLM in 88.3% of cases during ChatGPT-based evaluation and human expert prefers it over other open-source models. |
Modelling Etymology in LMF/TEI: The Grande Dicionário Houaiss da Língua Portuguesa Dictionary as a Use Case (2020.lrec-1)
Copied to clipboard
| Challenge: | In this article, we will introduce two of the new parts of the Lexical Markup Framework (LMF) ISO standard . part 3 deals with etymological and diachronic data and part 4 consists of a TEI serialisation of all of the prior parts of TEIS model. |
| Approach: | They introduce two parts of the Lexical Markup Framework (LMF) ISO standard, part 3 dealing with etymological and diachronic data and part 4 containing TEI serialisation of all prior parts of a model. |
| Outcome: | The proposed models are based on examples taken from a Portuguese dictionary conversion and are then compared with TEI-XML models. |
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models (2024.acl-long)
Copied to clipboard
| Challenge: | a surge of deep learning applications for video understanding have led to major advancements in video-related tasks. |
| Approach: | They propose a multimodal video-based conversation model that merges a video-adapted visual encoder with an LLM and a dataset that is easily scalable and robust to label noise. |
| Outcome: | The proposed model can understand and generate detailed conversations about videos. |
A Multilingual Evaluation Dataset for Monolingual Word Sense Alignment (2020.lrec-1)
Copied to clipboard
Sina Ahmadi, John Philip McCrae, Sanni Nimb, Fahad Khan, Monica Monachini, Bolette Pedersen, Thierry Declerck, Tanja Wissik, Andrea Bellandi, Irene Pisani, Thomas Troelsgård, Sussi Olsen, Simon Krek, Veronika Lipp, Tamás Váradi, László Simon, András Gyorffy, Carole Tiberius, Tanneke Schoonheim, Yifat Ben Moshe, Maya Rudich, Raya Abu Ahmad, Dorielle Lonke, Kira Kovalenko, Margit Langemets, Jelena Kallas, Oksana Dereza, Theodorus Fransen, David Cillessen, David Lindemann, Mikel Alonso, Ana Salgado, José Luis Sancho, Rafael-J. Ureña-Ruiz, Jordi Porta Zamorano, Kiril Simov, Petya Osenova, Zara Kancheva, Ivaylo Radev, Ranka Stanković, Andrej Perdih, Dejan Gabrovsek
| Challenge: | a new dataset aims to align monolingual dictionaries with a single sense level for 15 languages . this dataset covers a wide range of languages and resources . |
| Approach: | They propose to manually align monolingual dictionaries with possible semantic relationships . they use 15 languages to create a new baseline for the task of monolingual word sense alignment . |
| Outcome: | The proposed dataset covers 15 languages and covers the more challenging task of linking general-purpose language. |
One Language to rule them all: modelling Morphological Patterns in a Large Scale Italian Lexicon with SWRL (L18-1)
Copied to clipboard
| Challenge: | Linked data (LD) is a popular way of publishing lexical resources, but technical limitations and potentialities of LD are not understood as they should be. |
| Approach: | They propose to use the Semantic Web Rule Language to encode morphological patterns for a lexicographic publication as linked open data. |
| Outcome: | The proposed language allows the automatic derivation of inflectional variants of entries in the lexicon. |
MultiLexBATS: Multilingual Dataset of Lexical Semantic Relations (2024.lrec-main)
Copied to clipboard
Dagmar Gromann, Hugo Goncalo Oliveira, Lucia Pitarch, Elena-Simona Apostol, Jordi Bernad, Eliot Bytyçi, Chiara Cantone, Sara Carvalho, Francesca Frontini, Radovan Garabik, Jorge Gracia, Letizia Granata, Fahad Khan, Timotej Knez, Penny Labropoulou, Chaya Liebeskind, Maria Pia Di Buono, Ana Ostroški Anić, Sigita Rackevičienė, Ricardo Rodrigues, Gilles Sérasset, Linas Selmistraitis, Mahammadou Sidibé, Purificação Silvano, Blerina Spahiu, Enriketa Sogutlu, Ranka Stanković, Ciprian-Octavian Truică, Giedre Valunaite Oleskeviciene, Slavko Zitnik, Katerina Zdravkova
| Challenge: | Prior work has focused on analysing lexical semantic relations in word embeddings or probing pretrained language models (PLMs) with some exceptions. |
| Approach: | They propose to use a multilingual parallel dataset of lexical semantic relations adapted from BATS in 15 languages including low-resource languages such as Bambara, Lithuanian, and Albanian as an experiment on cross-lingual transfer of relational knowledge. |
| Outcome: | The proposed dataset is adapted from a BATS-based dataset in 15 languages including low-resource languages such as Bambara, Lithuanian, and Albanian. |
Towards the Construction of a WordNet for Old English (2022.lrec-1)
Copied to clipboard
Fahad Khan, Francisco J. Minaya Gómez, Rafael Cruz González, Harry Diakoff, Javier E. Diaz Vera, John P. McCrae, Ciara O’Loughlin, William Michael Short, Sander Stolk
| Challenge: | In this paper we discuss our preliminary work towards the construction of a WordNet for Old English, taking our inspiration from other similar WN construction projects for ancient languages such as Ancient Greek, Latin and Sanskrit. |
| Approach: | They propose to use a legacy Old English dictionary to build a WordNet for Old English using a lexicographic resource and the naisc system to automatically compile a provisional version of the WordNet. |
| Outcome: | The proposed OldEWN will be based on lemmas and definitions extracted from a legacy Old English dictionary and will be automatically compile and enriched by experts using the naisc system. |
On Modelling Corpus Citations in Computational Lexical Resources (2024.lrec-main)
Copied to clipboard
| Challenge: | TEI and OntoLex deal with corpus citations in lexicons. |
| Approach: | They argue that TEI and OntoLex can be used to model corpus citations in lexicons . they also argue that they should be combined to achieve a more accurate encoding . |
| Outcome: | The proposed approach favours a combination of TEI and OntoLex . the proposed approach is based on a model of an example entry from a legacy dictionary . |
BiMediX: Bilingual Medical Mixture of Experts LLM (2024.findings-emnlp)
Copied to clipboard
Sara Pieri, Sahal Shaji Mullappilly, Fahad Khan, Rao Anwer, Salman Khan, Timothy Baldwin, Hisham Cholakkal
| Challenge: | a new bilingual medical mixture of experts LLM is designed for seamless interaction in both English and Arabic. |
| Approach: | They propose a semi-automated English-to-Arabic translation pipeline with human refinement to ensure high-quality translations. |
| Outcome: | The proposed model outperforms state-of-the-art medical LLMs in Arabic and Arabic . it outperformed the generic Arabic-English bilingual LLM, Jais-30B by 10% and 15% . |