Papers by Mohamed Abdalla
Collaboration or Corporate Capture? Quantifying NLP’s Reliance on Industry Artifacts and Contributions (2024.acl-long)
Copied to clipboard
| Challenge: | EMNLP 2022 citations are three times greater than expected for pre-trained models . industry participation in the Association of Computational Linguistics (ACL) anthology has increased 180% from 2017 to 2022. |
| Approach: | They surveyed 100 papers published at EMNLP 2022 to determine the ratio of their citations to industry models. |
| Outcome: | a new study shows that industry citations are three times greater than expected . the study aims to better understand whether industry collaboration is still collaboration . industry participation in the 2023 AI index report is the top takeaway . |
SemRel2024: A Collection of Semantic Textual Relatedness Datasets for 13 Languages (2024.findings-acl)
Copied to clipboard
Nedjma Ousidhoum, Shamsuddeen Muhammad, Mohamed Abdalla, Idris Abdulmumin, Ibrahim Ahmad, Sanchit Ahuja, Alham Aji, Vladimir Araujo, Abinew Ayele, Pavan Baswani, Meriem Beloucif, Chris Biemann, Sofia Bourhim, Christine Kock, Genet Dekebo, Oumaima Hourrane, Gopichand Kanumolu, Lokesh Madasu, Samuel Rutunda, Manish Shrivastava, Thamar Solorio, Nirmal Surange, Hailegnaw Tilaye, Krishnapriya Vishnubhotla, Genta Winata, Seid Yimam, Saif Mohammad
| Challenge: | SemRel datasets are annotated by native speakers across 13 languages . they are used to characterise the relationship between two units of text . |
| Approach: | They propose to use a semantic relatedness dataset to measure the degree of semantic textual relatedness between sentences in Afrikaans, Algerian Arabic, Amharic, English, Hausa, Hindi, Indonesian, Kinyarwanda, Marathi, Moroccan Arabic, Modern Standard Arabic, Spanish, and Telugu. |
| Outcome: | The proposed datasets are annotated by native speakers across 13 languages and represent the semantic relatedness of 13 languages. |
Not What the Doctor Ordered: Surveying LLM-based De-identification and Quantifying Clinical Information Loss (2025.emnlp-main)
Copied to clipboard
Kiana Aghakasiri, Noopur Zambare, JoAnn Thai, Carrie Ye, Mayur Mehta, J Ross Mitchell, Mohamed Abdalla
| Challenge: | De-identification is an application of NLP where automated algorithms remove identifying information of patients and providers. |
| Approach: | They propose to use generative large language models to de-identify patients and providers . they propose to validate existing metrics to quantify extent of inappropriate removal . |
| Outcome: | The proposed method is based on a survey of LLM-based de-identification research . it shows that the models perform poorly in identifying clinically relevant changes . |
Citation Amnesia: On The Recency Bias of NLP and Other Academic Fields (2025.coling-main)
Copied to clipboard
| Challenge: | citation age is a key factor in determining whether older works are cited in scientific journals or not. |
| Approach: | They examine the tendency of NLP to cite older work across 20 fields of study over 43 years (1980–2023) . they put NLP’s propensity to citation older work in the context of these 20 other fields to see whether differences can be observed . |
| Outcome: | The trend is strongest in NLP and ML research (-12.8% and -5.5% in citation age from previous peaks) |
Towards Fair and Efficient De-identification: Quantifying the Efficiency and Generalizability of De-identification Approaches (2026.findings-eacl)
Copied to clipboard
| Challenge: | a recent study has not examined their generalizability between formats, cultures, and genders. |
| Approach: | They evaluate large language models (LLMs) and small LLMs at clinical de-identification . they show that smaller models achieve comparable performance while substantially reducing inference cost . |
| Outcome: | The proposed models outperform larger models in de-identification tasks with limited data . the models can be fine-tuned with limited datasets to outperformed larger models . |
We are Who We Cite: Bridges of Influence Between Natural Language Processing and Other Academic Fields (2023.emnlp-main)
Copied to clipboard
| Challenge: | In this paper, we quantify the degree of influence between 23 fields of study and NLP (on each other) |
| Approach: | They quantify the degree of influence between 23 fields of study and NLP on each other . they find that cross-field engagement of NLP has declined from 0.58 in 1980 to 0.31 in 2022 . |
| Outcome: | The proposed Citation Field Diversity Index (CFDI) has declined from 0.58 in 1980 to 0.31 in 2022, the authors show . |
What Makes Sentences Semantically Related? A Textual Relatedness Dataset and Empirical Study (2023.eacl-main)
Copied to clipboard
| Challenge: | Existing work on semantic relatedness has focused on semantic similarity because of a lack of relatedness datasets. |
| Approach: | They propose a dataset for semantic relatedness that has 5,500 English sentence pairs manually annotated using a comparative annotation framework. |
| Outcome: | The proposed dataset has 5,500 English sentence pairs manually annotated using a comparative annotation framework. |
The Elephant in the Room: Analyzing the Presence of Big Tech in Natural Language Processing Research (2023.acl-long)
Copied to clipboard
Mohamed Abdalla, Jan Philip Wahle, Terry Ruas, Aurélie Névéol, Fanny Ducel, Saif Mohammad, Karen Fort
| Challenge: | Recent advances in deep learning methods for natural language processing (NLP) have created new business opportunities and made NLP research critical for industry development. |
| Approach: | They examine industry presence in the field since the early 90s and characterize it using a corpus of 78,187 NLP publications and 701 resumes of NLP publication authors. |
| Outcome: | The authors find that industry presence among NLP authors has been steady before a steep increase over the past five years (180% growth from 2017 to 2022). |