Papers by Mohammadamin Shafiei
TruthTrap: A Bilingual Benchmark for Evaluating Factually Correct Yet Misleading Information in Question Answering (2026.findings-eacl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly used to answer factual, information-seeking questions (ISQs). |
| Approach: | They propose to use a dataset to evaluate large language models to generate human-like text on ISQs in two languages, English and Farsi, and then use it to evaluate nine LLMs. |
| Outcome: | The proposed dataset shows that accuracy drops by 25% when models encounter misleading yet factual hints. |
MultiHoax: A Dataset of Multi-hop False-premise questions (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks focus on single-hop FPQs, but real-world reasoning often requires multi-hop inference . state-of-the-art LLMs struggle to detect false premises across different countries, knowledge categories, and multi-step reasoning types. |
| Approach: | They propose a benchmark to evaluate Large Language Models' ability to handle false premises in complex, multi-step reasoning tasks. |
| Outcome: | The proposed tests show that state-of-the-art LLMs struggle to detect false premises across different countries, knowledge categories, and multi-hop reasoning types. |
Can I Introduce My Boyfriend to My Grandmother? Evaluating Large Language Models Capabilities on Iranian Social Norm Classification (2025.findings-naacl)
Copied to clipboard
| Challenge: | Introducing the Iranian Social Norms dataset, a collection of 1,699 social norms, with Farsi adding linguistic complexity. |
| Approach: | They propose a collection of Iranian social norms with English translations and a novel Iranian dataset. |
| Outcome: | The Iranian Social Norms dataset is the first to be used in the Farsi language . it includes 1,699 social norms including environments, demographic features, and scope annotation, alongside English translations. |
Beyond Hate Speech: NLP’s Challenges and Opportunities in Uncovering Dehumanizing Language (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing hate speech datasets rarely contain enough instances of dehumanizing content, and current models struggle to distinguish such language from more benign forms of hate or offense. |
| Approach: | They evaluate four state-of-the-art large language models for dehumanization detection. |
| Outcome: | The proposed models perform only moderately under an optimized configuration, while others over-predict dehumanization for some identities, while under-identifying it for others. |