Papers by Firoj Alam
A Survey on Multimodal Disinformation Detection (2022.coling-1)
Copied to clipboard
Firoj Alam, Stefano Cresci, Tanmoy Chakraborty, Fabrizio Silvestri, Dimiter Dimitrov, Giovanni Da San Martino, Shaden Shaar, Hamed Firooz, Preslav Nakov
| Challenge: | Recent years have witnessed the proliferation of offensive content online such as fake news, propaganda, misinformation, and disinformation. |
| Approach: | They propose to tackle online multimodal offensive content using different modalities and combinations thereof. |
| Outcome: | The proposed approach combines factuality and harmfulness in a framework that can be used for multiple modalities and combinations of modality. |
TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking (2025.findings-acl)
Copied to clipboard
Shahriar Kabir Nahin, Rabindra Nath Nandi, Sagor Sarker, Quazi Sarwar Muhtaseem, Md Kowsher, Apu Chandraw Shill, Md Ibrahim, Mehadi Hasan Menon, Tareq Al Muntasir, Firoj Alam
| Challenge: | Existing benchmarking datasets for Bangla LLMs are not available for all languages. |
| Approach: | They present TituLLMs, the first large pretrained Bangla LLMs, available in 1b and 3b parameter sizes. |
| Outcome: | The proposed model outperforms existing models in Bangla, but not always in the first place. |
POLAR: A Benchmark for Multilingual, Multicultural, and Multi-Event Online Polarization (2026.findings-acl)
Copied to clipboard
Usman Naseem, Robert Geislinger, Juan Ren, Sarah Kohail, Rudy Alexandro Garrido Veliz, P Sam Sahil, Yiran Zhang, Idris Abdulmumin, Marco Antonio Stranisci, Özge Alacam, Cengiz Acarturk, Aisha Jabr, Saba Anwar, Abinew Ali Ayele, Simona Frenda, Alessandra Teresa Cignarella, Elena Tutubalina, Oleg Rogov, Aung Kyaw Htet, Xintong Wang, Surendrabikram Thapa, Kritesh Rauniyar, Tanmoy Chakraborty, MD Arfeen Zeeshan, Dheeraj Kodati, Satya Keerthi, Sahar Moradizeyveh, Firoj Alam, Md Arid Hasan, Syed Ishtiaque Ahmed, Ye Kyaw Thu, Shantipriya Parida, Ihsan Ayyub Qazi, Lilian Diana Awuor Wanzare, Nelson Odhiambo Onyango, Clemencia Siro, Jane Wanjiru Kimani, Ibrahim Said Ahmad, Adem Chanie Ali, Martin Semmann, Chris Biemann, Shamsuddeen Hassan Muhammad, Seid Muhie Yimam
| Challenge: | polarization is a pervasive threat to democratic institutions, civil discourse, and social cohesion worldwide . most existing datasets focus on English or high-resource languages, reflecting a widespread trend across NLP tasks . |
| Approach: | They propose a multilingual, multicultural, and multi-event dataset with over 110K instances in 22 languages drawn from diverse online platforms and real-world events. |
| Outcome: | The proposed dataset analyzes polarization detection, type, and manifestation using a variety of annotation platforms adapted to each cultural context. |
From RAG to Agentic RAG for Faithful Islamic Question Answering (2026.findings-acl)
Copied to clipboard
Gagan Bhatia, Hamdy Mubarak, Mustafa Jarrar, George Mikros, Fadi Zaraket, Mahmoud Alhirthani, Mutaz al-Khatib, Logan Cochrane, Kareem Mohamed Darwish, Rashid Yahiaoui, Firoj Alam
| Challenge: | Large Language Models (LLMs) are increasingly used for Islamic question answering, where ungrounded responses may carry serious religious consequences. |
| Approach: | They propose a bilingual, bilingual, Arabic/English benchmark with atomic single-gold answers that measures hallucination and abstention. |
| Outcome: | The proposed model improves accuracy and robustness even with a small model. |
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs (2025.coling-main)
Copied to clipboard
Basel Mousi, Nadir Durrani, Fatema Ahmad, Md. Arid Hasan, Maram Hasanain, Tameem Kabbani, Fahim Dalvi, Shammur Absar Chowdhury, Firoj Alam
| Challenge: | a recent study has found that Arabic is underrepresented in Large Language Models, especially in dialectal variations. |
| Approach: | They propose a benchmark for Arabic Dialect and Cultural Evaluation that evaluates Arabic dialect comprehension and generation. |
| Outcome: | The proposed model outperforms multilingual models on dialect comprehension and generation, but significant challenges persist in dialect identification, generation, and translation. |
BnTTS: Few-Shot Speaker Adaptation in Low-Resource Setting (2025.findings-naacl)
Copied to clipboard
Mohammad Jahid Ibna Basher, Md Kowsher, Md Saiful Islam, Rabindra Nath Nandi, Nusrat Jahan Prottasha, Mehadi Hasan Menon, Tareq Al Muntasir, Shammur Absar Chowdhury, Firoj Alam, Niloofar Yousefi, Ozlem Garibay
| Challenge: | Empirical evaluations in few-shot settings show that BnTTS significantly improves the naturalness, intelligibility, and speaker fidelity of synthesized Bangla speech. |
| Approach: | They propose to integrate Bangla into a multilingual TTS pipeline with modifications to account for the phonetic and linguistic characteristics of the language. |
| Outcome: | The proposed framework improves the naturalness, intelligibility, and speaker fidelity of synthesized Bangla speech compared to state-of-the-art systems. |
NativQA: Multilingual Culturally-Aligned Natural Query for LLMs (2025.findings-acl)
Copied to clipboard
Md. Arid Hasan, Maram Hasanain, Fatema Ahmad, Sahinur Rahman Laskar, Sunaya Upadhyay, Vrunda N Sukhadia, Mucahid Kutlu, Shammur Absar Chowdhury, Firoj Alam
| Challenge: | Existing frameworks for QA datasets lack regional specificity and cultural specificity. |
| Approach: | They propose a framework to quench native language QA datasets in native languages for LLM evaluation and tuning. |
| Outcome: | The proposed framework is scalable, language-independent and can be used to build culturally and regionally aligned QA datasets in native languages. |
ArCovidVac: Analyzing Arabic Tweets About COVID-19 Vaccination (2022.lrec-1)
Copied to clipboard
| Challenge: | Social media are integrated with our daily life and are used to circulate information. |
| Approach: | They develop and publicly release the first largest manually annotated Arabic tweet dataset for COVID-19 vaccination campaign. |
| Outcome: | The proposed dataset is the largest manually annotated Arabic tweet dataset for COVID-19 vaccination campaign, covering many countries in the Arab region. |
CritiSense: Critical Digital Literacy and Resilience Against Misinformation (2026.acl-demo)
Copied to clipboard
Firoj Alam, Fatema Ahmad, Ali Ezzat Shahroor, Mohamed Bayan Kmainasi, Elisa Sartori, Giovanni Da San Martino, Abul Hasnat, Raian Ali
| Challenge: | a recent study found that social media misinformation is reactive and claim-specific, and can degrade under temporal and cross-lingual/domain shift. |
| Approach: | They present a mobile media-literacy app that builds digital literacy skills through short, interactive challenges with instant feedback. |
| Outcome: | The app is the first multilingual and modular platform to improve digital literacy skills. |
Zero- and Few-Shot Prompting with LLMs: A Comparative Study with Fine-tuned Models for Bangla Sentiment Analysis (2024.lrec-main)
Copied to clipboard
Md. Arid Hasan, Shudipta Das, Afiyat Anjum, Firoj Alam, Anika Anjum, Avijit Sarker, Sheak Rashed Haider Noori
| Challenge: | Recent performance of Large Language Models (LLMs) in low-resource languages is under-researched due to resource constraints. |
| Approach: | They present a manually annotated dataset encompassing 33,606 Bangla tweets and Facebook comments. |
| Outcome: | The proposed model outperforms other models even in zero and few-shot scenarios. |
LLMs for Low Resource Languages in Multilingual, Multimodal and Dialectal Settings (2024.eacl-tutorials)
Copied to clipboard
| Challenge: | Recent advances in AI can be attributed to the remarkable performance of Large Language Models (LLMs) success of LLMs depends on specific training techniques, such as instruction tuning and prompting . |
| Approach: | They explore the capabilities of Large Language Models (LLMs) in various tasks and languages . they also examine their performance, fine-tuning, instructions tuning, and close vs. open models . |
| Outcome: | The proposed model can be used for speech and multimodal tasks across modalities, languages, and dialects. |
LlamaLens: Specialized Multilingual LLM for Analyzing News and Social Media Content (2025.findings-naacl)
Copied to clipboard
Mohamed Bayan Kmainasi, Ali Ezzat Shahroor, Maram Hasanain, Sahinur Rahman Laskar, Naeemul Hassan, Firoj Alam
| Challenge: | Large Language Models (LLMs) have demonstrated remarkable success as general-purpose task solvers across various fields. |
| Approach: | They propose to develop a specialized LLM for analyzing news and social media content in a multilingual context. |
| Outcome: | The proposed model outperforms the current state-of-the-art on 23 testing sets and achieves comparable performance on 8 sets. |
Fanar-Sadiq: A Multi-Agent Architecture for Grounded Islamic QA (2026.acl-industry)
Copied to clipboard
Ummar Abbas, Mourad Ouzzani, Mohamed Y. Eltabakh, Omar Sinan, Gagan Bhatia, Hamdy Mubarak, Majd Hawasly, Mohammed Qusay Hashim, Kareem Mohamed Darwish, Firoj Alam
| Challenge: | Large language models (LLMs) can answer religious knowledge queries fluently, but they often hallucinate and misattribute sources. |
| Approach: | They propose a bilingual Arabic-English Islamic QA system that uses a multi-agent, tool-augmented architecture to route Islamic queries to specialized modules. |
| Outcome: | The proposed system is based on a multi-agent, tool-augmented architecture and has received over 1.9M accesses in less than a year. |
Once Correct, Still Wrong: Counterfactual Hallucination in Multilingual Vision-Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing hallucination benchmarks rarely test this failure mode outside Western contexts and English. |
| Approach: | They propose a multimodal benchmark built from images spanning 17 MENA countries . they use a CFHR-based test to measure hallucination beyond raw accuracy . |
| Outcome: | The proposed model is based on images from 17 MENA countries . it measures counterfactual acceptance conditioned on correctly answering the true statement. |
LAraBench: Benchmarking Arabic AI with Large Language Models (2024.eacl-long)
Copied to clipboard
Ahmed Abdelali, Hamdy Mubarak, Shammur Chowdhury, Maram Hasanain, Basel Mousi, Sabri Boughorbel, Samir Abdaljalil, Yassine El Kheir, Daniel Izham, Fahim Dalvi, Majd Hawasly, Nizi Nazar, Youssef Elshahawy, Ahmed Ali, Nadir Durrani, Natasa Milic-Frayling, Firoj Alam
| Challenge: | Recent advances in Large Language Models (LLMs) have significantly influenced the landscape of language and speech research. |
| Approach: | They used GPT-3.5-turbo, GPT-4, BLOOMZ, Jais-13b-chat, Whisper, and USM to tackle 33 distinct tasks across 61 datasets. |
| Outcome: | The proposed model outperforms SOTA models in zero-shot learning, with a few exceptions. |
Detecting Propaganda Techniques in Memes (2021.acl-long)
Copied to clipboard
Dimitar Dimitrov, Bishr Bin Ali, Shaden Shaar, Firoj Alam, Fabrizio Silvestri, Hamed Firooz, Preslav Nakov, Giovanni Da San Martino
| Challenge: | Propaganda can be defined as a form of communication that aims to influence opinions or the actions of people towards a specific goal. |
| Approach: | They propose to detect the type of propaganda techniques used in memes by annotating them with 22 techniques. |
| Outcome: | The proposed model identifies 22 propaganda techniques in memes, which can appear in text, image or both . |
MemeIntel: Explainable Detection of Propagandistic and Hateful Memes (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for label detection and explanation generation have been limited in understanding complex issues . identifying propaganda and hate in memes is essential for combating misinformation and minimizing harm . |
| Approach: | They propose an explanation-enhanced dataset for propaganda memes in Arabic and hateful memes on English to solve these tasks. |
| Outcome: | The proposed model outperforms the current state-of-the-art in label detection and explanation generation. |
Annotating the Annotators: Analysis, Insights and Modelling from an Annotation Campaign on Persuasion Techniques Detection (2025.findings-acl)
Copied to clipboard
Davide Bassi, Dimitar Iliyanov Dimitrov, Bernardo D’Auria, Firoj Alam, Maram Hasanain, Christian Moro, Luisa Orrù, Gian Piero Turchi, Preslav Nakov, Giovanni Da San Martino
| Challenge: | Existing annotation campaigns based on heuristic guidelines have not been thoroughly discussed. |
| Approach: | They propose a probabilistic model for optimizing intervention scheduling to reduce the cost of an expert oversight in annotation tasks. |
| Outcome: | The proposed model advocates for an expert oversight in annotation tasks and periodic quality audits to reduce costs. |
The Role of Context in Detecting Previously Fact-Checked Claims (2022.findings-naacl)
Copied to clipboard
| Challenge: | Recent years have seen the proliferation of disinformation and fake news online. |
| Approach: | They propose to model the context of a political debate and the contexts of the document describing the fact-checked claim. |
| Outcome: | The proposed model improves on the state-of-the-art model by modeling the context of the claim . the experimental results show that the model can provide 10+ points of improvement over the state of the art model . |
Assisting the Human Fact-Checkers: Detecting All Previously Fact-Checked Claims in a Document (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Recent years have brought us a proliferation of false claims online, which spread fast . fact-checkers have been using automated fact-finding to verify claims . |
| Approach: | They propose a system that can detect claims that can be fact-checked by a given database . they create a manually annotated document dataset and propose evaluation measures . |
| Outcome: | The proposed system achieves sizable performance gains over strong baselines. |
Effect of Post-processing on Contextualized Word Representations (2022.coling-1)
Copied to clipboard
| Challenge: | Post-processing of static embeddings has been shown to improve their performance on both lexical and sequence-level tasks. |
| Approach: | They standardize individual neuron activations using z-score, min-max normalization, and remove top principal components using the all-but-the-top method. |
| Outcome: | The proposed method unwraps vital information present in the representations for both lexical and sequence classification tasks. |
Domain Adaptation with Adversarial Training and Graph Embeddings (P18-1)
Copied to clipboard
| Challenge: | Existing models for deep neural networks can handle data distributions between source and target domains, but they must deal with data distribution drifts. |
| Approach: | They propose a model that leverages unlabeled and labeled data from a related domain to deal with distribution drifts. |
| Outcome: | The proposed model improves over baselines on two real-world disaster datasets. |
Large Language Models for Propaganda Span Annotation (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Using propagandistic techniques to manipulate online audiences is increasing in recent years. |
| Approach: | They investigate whether Large Language Models (LLMs) such as GPT-4 can extract propagandistic spans and the potential of employing them to collect more cost-effective annotations. |
| Outcome: | The proposed model provides labels that have higher agreement with expert annotators and lead to specialized models that achieve state-of-the-art over an unseen Arabic testing set. |
ArMeme: Propagandistic Content in Arabic Memes (2024.emnlp-main)
Copied to clipboard
| Challenge: | a lack of media literacy is a major factor contributing to the spread of misleading information on social media. |
| Approach: | They analyze a dataset of 6K Arabic memes with manual annotations . they propose to develop computational tools for their detection . |
| Outcome: | The proposed dataset is a first resource for Arabic multimodal research. |
Can GPT-4 Identify Propaganda? Annotation and Detection of Propaganda Spans in News Articles (2024.lrec-main)
Copied to clipboard
| Challenge: | Using large language models (LLMs) to detect propaganda from text is a challenge for the development of sophisticated models. |
| Approach: | They propose to use a large propaganda dataset to identify propagandistic content in text, visual, or multimodal languages to improve their models. |
| Outcome: | The proposed model performs better on a large propaganda dataset than the existing models on skewed datasets. |
Analyzing Encoded Concepts in Transformer Language Models (2022.naacl-main)
Copied to clipboard
| Challenge: | a new framework to analyze how latent concepts are encoded in representations learned in pre-trained lan-guage models is proposed . conceptX uses clustering to discover the encoded concepts and align them with a large set of human-defined concepts. |
| Approach: | They propose a framework to analyze how latent concepts are encoded in representations learned within pre-trained lan-guage models. |
| Outcome: | The proposed framework explains encoded concepts by aligning with human-defined concepts. |
On the Transformation of Latent Space in Fine-Tuned NLP Models (2022.emnlp-main)
Copied to clipboard
| Challenge: | a large body of work analyzed the knowledge learned within representations of pre-trained models. |
| Approach: | They use hierarchical clustering to discover latent concepts in representational space . they compare pre-trained and fine-tuned models and perform a thorough analysis . |
| Outcome: | The results show that the model space evolves towards task-specific concepts whereas the lower layers retain generic concepts acquired in the pre-trained model. |
PropXplain: Can LLMs Enable Explainable Propaganda Detection? (2025.findings-emnlp)
Copied to clipboard
Maram Hasanain, Md Arid Hasan, Mohamed Bayan Kmainasi, Elisa Sartori, Ali Ezzat Shahroor, Giovanni Da San Martino, Firoj Alam
| Challenge: | Currently, propagandistic content detection studies focus on detection, with little attention given to explanations justifying the predicted label. |
| Approach: | They propose a multilingual explanation-enhanced dataset and an explanation-based LLM to address this issue. |
| Outcome: | The proposed model performs comparably while also generating explanations. |
LLM-Based Multi-Task Bangla Hate Speech Detection: Type, Severity, and Target (2026.acl-long)
Copied to clipboard
| Challenge: | Existing work on social media platforms is limited in its ability to detect hate speech . a lack of reliable and scalable automated hate speech detection systems is a challenge for low-resource languages like Bangla. |
| Approach: | They propose to use a single-task, single-targeted, single language dataset to identify hate speech in Bangla. |
| Outcome: | The proposed dataset is the largest manually annotated Bangla hate-speech dataset to date. |
Fighting the COVID-19 Infodemic: Modeling the Perspective of Journalists, Fact-Checkers, Social Media Platforms, Policy Makers, and the Society (2021.findings-emnlp)
Copied to clipboard
Firoj Alam, Shaden Shaar, Fahim Dalvi, Hassan Sajjad, Alex Nikolov, Hamdy Mubarak, Giovanni Da San Martino, Ahmed Abdelali, Nadir Durrani, Kareem Darwish, Abdulaziz Al-Homaid, Wajdi Zaghouani, Tommaso Caselli, Gijs Danoe, Friso Stolk, Britt Bruntink, Preslav Nakov
| Challenge: | a dataset of 16K manually annotated tweets is used to analyze disinformation . the democratic nature of social media has raised questions about the quality and the factuality of the information that is shared on these platforms. |
| Approach: | They use a dataset of manually annotated tweets to analyze COVID-19 disinformation . they show that tweets contain fake cures, rumors, conspiracy theories and xenophobia . |
| Outcome: | The proposed dataset shows that it is useful in monolingual vs. multilingual settings. |