Papers by Mohammadamin Shafiei

4 papers
TruthTrap: A Bilingual Benchmark for Evaluating Factually Correct Yet Misleading Information in Question Answering (2026.findings-eacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used to answer factual, information-seeking questions (ISQs).
Approach: They propose to use a dataset to evaluate large language models to generate human-like text on ISQs in two languages, English and Farsi, and then use it to evaluate nine LLMs.
Outcome: The proposed dataset shows that accuracy drops by 25% when models encounter misleading yet factual hints.
MultiHoax: A Dataset of Multi-hop False-premise questions (2025.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks focus on single-hop FPQs, but real-world reasoning often requires multi-hop inference . state-of-the-art LLMs struggle to detect false premises across different countries, knowledge categories, and multi-step reasoning types.
Approach: They propose a benchmark to evaluate Large Language Models' ability to handle false premises in complex, multi-step reasoning tasks.
Outcome: The proposed tests show that state-of-the-art LLMs struggle to detect false premises across different countries, knowledge categories, and multi-hop reasoning types.
Can I Introduce My Boyfriend to My Grandmother? Evaluating Large Language Models Capabilities on Iranian Social Norm Classification (2025.findings-naacl)

Copied to clipboard

Challenge: Introducing the Iranian Social Norms dataset, a collection of 1,699 social norms, with Farsi adding linguistic complexity.
Approach: They propose a collection of Iranian social norms with English translations and a novel Iranian dataset.
Outcome: The Iranian Social Norms dataset is the first to be used in the Farsi language . it includes 1,699 social norms including environments, demographic features, and scope annotation, alongside English translations.
Beyond Hate Speech: NLP’s Challenges and Opportunities in Uncovering Dehumanizing Language (2025.emnlp-main)

Copied to clipboard

Challenge: Existing hate speech datasets rarely contain enough instances of dehumanizing content, and current models struggle to distinguish such language from more benign forms of hate or offense.
Approach: They evaluate four state-of-the-art large language models for dehumanization detection.
Outcome: The proposed models perform only moderately under an optimized configuration, while others over-predict dehumanization for some identities, while under-identifying it for others.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations