Papers by Mohamed Abdalla

8 papers
Collaboration or Corporate Capture? Quantifying NLP’s Reliance on Industry Artifacts and Contributions (2024.acl-long)

Copied to clipboard

Challenge: EMNLP 2022 citations are three times greater than expected for pre-trained models . industry participation in the Association of Computational Linguistics (ACL) anthology has increased 180% from 2017 to 2022.
Approach: They surveyed 100 papers published at EMNLP 2022 to determine the ratio of their citations to industry models.
Outcome: a new study shows that industry citations are three times greater than expected . the study aims to better understand whether industry collaboration is still collaboration . industry participation in the 2023 AI index report is the top takeaway .
SemRel2024: A Collection of Semantic Textual Relatedness Datasets for 13 Languages (2024.findings-acl)

Copied to clipboard

Challenge: SemRel datasets are annotated by native speakers across 13 languages . they are used to characterise the relationship between two units of text .
Approach: They propose to use a semantic relatedness dataset to measure the degree of semantic textual relatedness between sentences in Afrikaans, Algerian Arabic, Amharic, English, Hausa, Hindi, Indonesian, Kinyarwanda, Marathi, Moroccan Arabic, Modern Standard Arabic, Spanish, and Telugu.
Outcome: The proposed datasets are annotated by native speakers across 13 languages and represent the semantic relatedness of 13 languages.
Not What the Doctor Ordered: Surveying LLM-based De-identification and Quantifying Clinical Information Loss (2025.emnlp-main)

Copied to clipboard

Challenge: De-identification is an application of NLP where automated algorithms remove identifying information of patients and providers.
Approach: They propose to use generative large language models to de-identify patients and providers . they propose to validate existing metrics to quantify extent of inappropriate removal .
Outcome: The proposed method is based on a survey of LLM-based de-identification research . it shows that the models perform poorly in identifying clinically relevant changes .
Citation Amnesia: On The Recency Bias of NLP and Other Academic Fields (2025.coling-main)

Copied to clipboard

Challenge: citation age is a key factor in determining whether older works are cited in scientific journals or not.
Approach: They examine the tendency of NLP to cite older work across 20 fields of study over 43 years (1980–2023) . they put NLP’s propensity to citation older work in the context of these 20 other fields to see whether differences can be observed .
Outcome: The trend is strongest in NLP and ML research (-12.8% and -5.5% in citation age from previous peaks)
Towards Fair and Efficient De-identification: Quantifying the Efficiency and Generalizability of De-identification Approaches (2026.findings-eacl)

Copied to clipboard

Challenge: a recent study has not examined their generalizability between formats, cultures, and genders.
Approach: They evaluate large language models (LLMs) and small LLMs at clinical de-identification . they show that smaller models achieve comparable performance while substantially reducing inference cost .
Outcome: The proposed models outperform larger models in de-identification tasks with limited data . the models can be fine-tuned with limited datasets to outperformed larger models .
We are Who We Cite: Bridges of Influence Between Natural Language Processing and Other Academic Fields (2023.emnlp-main)

Copied to clipboard

Challenge: In this paper, we quantify the degree of influence between 23 fields of study and NLP (on each other)
Approach: They quantify the degree of influence between 23 fields of study and NLP on each other . they find that cross-field engagement of NLP has declined from 0.58 in 1980 to 0.31 in 2022 .
Outcome: The proposed Citation Field Diversity Index (CFDI) has declined from 0.58 in 1980 to 0.31 in 2022, the authors show .
What Makes Sentences Semantically Related? A Textual Relatedness Dataset and Empirical Study (2023.eacl-main)

Copied to clipboard

Challenge: Existing work on semantic relatedness has focused on semantic similarity because of a lack of relatedness datasets.
Approach: They propose a dataset for semantic relatedness that has 5,500 English sentence pairs manually annotated using a comparative annotation framework.
Outcome: The proposed dataset has 5,500 English sentence pairs manually annotated using a comparative annotation framework.
The Elephant in the Room: Analyzing the Presence of Big Tech in Natural Language Processing Research (2023.acl-long)

Copied to clipboard

Challenge: Recent advances in deep learning methods for natural language processing (NLP) have created new business opportunities and made NLP research critical for industry development.
Approach: They examine industry presence in the field since the early 90s and characterize it using a corpus of 78,187 NLP publications and 701 resumes of NLP publication authors.
Outcome: The authors find that industry presence among NLP authors has been steady before a steep increase over the past five years (180% growth from 2017 to 2022).

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations