Papers by Ehsaneddin Asgari
Emo3D: Metric and Benchmarking Dataset for 3D Facial Expression Generation from Emotion Description (2025.findings-naacl)
Copied to clipboard
Mahshid Dehghani, Amirahmad Shafiee, Ali Shafiei, Neda Fallah, Farahmand Alizadeh, Mohammad Mehdi Gholinejad, Hamid Behroozi, Jafar Habibi, Ehsaneddin Asgari
| Challenge: | Existing 3D facial emotion modeling models are constrained by limited emotion classes and insufficient datasets. |
| Approach: | They propose a 3D facial emotion modeling dataset that spans a wide spectrum of human emotions . they use large language models to generate a diverse array of textual descriptions . |
| Outcome: | Emo3D is an extensive dataset that spans human emotions with images and 3D blendshapes. |
Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation (2025.findings-acl)
Copied to clipboard
Mohammad Mahdi Abootorabi, Amirhosein Zobeiri, Mahdi Dehghani, Mohammadali Mohammadkhani, Bardia Mohammadi, Omid Ghahroodi, Mahdieh Soleymani Baghshah, Ehsaneddin Asgari
| Challenge: | Large Language Models (LLMs) suffer from hallucinations and outdated knowledge due to their reliance on static training data. |
| Approach: | They review training strategies, robustness enhancements, loss functions, and agent-based approaches and outline open challenges and future directions to guide research in this evolving field. |
| Outcome: | The proposed model improves accuracy and accuracy while integrating external dynamic information for improved factual grounding. |
UniSent: Universal Adaptable Sentiment Lexica for 1000+ Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Sentiment lexica are vital for sentiment analysis in absence of document-level annotations . linguistic resources are limited for at least a few hundred languages, putting them at risk of extinction . |
| Approach: | They introduce UniSent universal sentiment lexica for 1000+ languages . they use a Bible corpus to project sentiment information from English to other languages based on Twitter data . |
| Outcome: | The proposed method mitigates domain mismatch between Bible and Twitter by using embeddings . it compares to other sentiment seeding methods in a subset of languages with ground truth available . |
MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment (2026.findings-eacl)
Copied to clipboard
Omid Ghahroodi, Arshia Hemmat, Marzia Nouri, Seyed Mohammad Hadi Hosseini, Doratossadat Dastgheib, Mohammad Vali Sanian, Alireza Sahebi, Reihaneh Zohrabi, Mohammad Hossein Rohban, Ehsaneddin Asgari, Mahdieh Soleymani Baghshah
| Challenge: | Recent advances in large vision-language models have primarily focused on English, with limited attention given to other languages. |
| Approach: | They propose a dataset to evaluate Persian VLMs across scientific, reasoning, and human-level understanding tasks. |
| Outcome: | The proposed model performs well across scientific reasoning, reasoning, and human-level understanding tasks in Persian and English. |
MorphBPE: Morphology-Aware Tokenization for Efficient LLM Training (2026.findings-acl)
Copied to clipboard
| Challenge: | Tokenization is a key design choice in modern NLP systems and a critical bottleneck for multilingual Large Language Models. |
| Approach: | They propose a tokenization extension that constrains merge operations to respect morpheme boundaries while preserving inference. |
| Outcome: | The proposed tokenization improves morphological coherence and language model cross-entropy in four languages. |
KnowMAN: Weakly Supervised Multinomial Adversarial Networks (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to weakly supervised training lack labeled data . weakly-supervised training can result in heuristic but noisy labels . |
| Approach: | They propose a scheme that allows to control influence of signals associated with specific labeling functions. |
| Outcome: | The proposed scheme improves results compared to weakly supervised learning with a pre-trained transformer language model and a feature-based baseline. |
TuringQ: Benchmarking AI Comprehension in Theory of Computation (2024.findings-emnlp)
Copied to clipboard
| Challenge: | TuringQ is the first benchmark designed to evaluate the reasoning capabilities of large language models (LLMs) in the theory of computation. |
| Approach: | They propose a benchmark to evaluate the reasoning capabilities of large language models in the theory of computation. |
| Outcome: | The proposed system shows competitive accuracy when compared to human evaluation. |
Detecting Subtle Biases: An Ethical Lens on Underexplored Areas in AI Language Models Biases (2026.eacl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly embedded in the daily lives of individuals across diverse social classes. |
| Approach: | They propose to analyze LLMs' responses to 1,016 scenarios categorized into ethical, unethical, and neutral types. |
| Outcome: | The proposed model analyzed 1,016 scenarios categorized into ethical, unethical, and neutral types. |
Transformers for Bridging Persian Dialects: Transliteration Model for Tajiki and Iranian Scripts (2024.lrec-main)
Copied to clipboard
| Challenge: | Despite its profound linguistic and cultural significance, Tajiki Persian remains a low-resource language with scant digitized datasets for computational applications. |
| Approach: | They propose to use Shahnameh, a seminal Persian epic poem, to train and assess Tajiki Persian transliteration models using two prominent sequence-to-sequence architectures: GRU with attention and transformer. |
| Outcome: | The proposed model outperforms pre-trained models with attention and transformer. |
Taxi1500: A Dataset for Multilingual Text Classification in 1500 Languages (2025.naacl-short)
Copied to clipboard
| Challenge: | a large-scale text classification dataset encompassing 1504 languages is needed to address this gap . low-resource languages are often overlooked due to the scarcity of evaluation datasets. |
| Approach: | They propose to use translations of the Bible to construct a large-scale text classification dataset that covers 1504 languages and annotate them using crowdsourcing. |
| Outcome: | The proposed dataset covers 1504 languages and is available to the public. |
Almieyar-Oryx-BloomBench: A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language Models (2026.findings-acl)
Copied to clipboard
Mohammad Mahdi Abootorabi, Omid Ghahroodi, Anas Madkoor, Marzia Nouri, Doratossadat Dastgheib, Ehsaneddin Asgari
| Challenge: | Existing evaluations focus on piecemeal or disconnected tasks, obscuring critical cognitive weaknesses and providing little insight for targeted improvement. |
| Approach: | They propose a bilingual, cognitively human-grounded multimodal benchmark for VLMs that evaluates six levels of cognition through carefully designed image–question–answer tasks. |
| Outcome: | The proposed framework ensures scalability, cultural inclusivity, and linguistic fidelity. |
The Touché23-ValueEval Dataset for Identifying Human Values behind Arguments (2024.lrec-main)
Copied to clipboard
Nailia Mirzakhmedova, Johannes Kiesel, Milad Alshomary, Maximilian Heinrich, Nicolas Handke, Xiaoni Cai, Valentin Barriere, Doratossadat Dastgheib, Omid Ghahroodi, MohammadAli SadraeiJavaheri, Ehsaneddin Asgari, Lea Kawaletz, Henning Wachsmuth, Benno Stein
| Challenge: | Cultural norms can influence the prioritization of values, leading to distinct perspectives on debatable topics. |
| Approach: | They present a Touché23-ValueEval dataset that annotates 4780 new arguments and annotated 54 human values. |
| Outcome: | The Touché23-ValueEval dataset doubles the original Webis-ArgValués-22 dataset to 9324 arguments. |
HarfoSokhan: A Comprehensive Parallel Dataset for Transitions between Persian Colloquial and Formal Variations (2026.eacl-long)
Copied to clipboard
Hamid Jahad Sarvestani, Vida Ramezanian, Saee Saadat, Neda Taghizadeh Serajeh, Maryam Sadat Razavi Taheri, Shohreh Kasaei, MohammadAmin Fazli, Ehsaneddin Asgari
| Challenge: | A wide array of NLP/NLU models have been developed for the Persian language but performance drops when applied to the colloquial form of Persian. |
| Approach: | They propose to use a large-scale colloquial to formal Persian parallel dataset to train a GPT2 model that exhibited remarkable proficiency in colloqual to informal text style transfer. |
| Outcome: | The proposed dataset outperforms OpenAI’s GPT-3.5-turbo model and a leading rule-based system in colloquial to formal Persian conversion. |
Hengam: An Adversarially Trained Transformer for Persian Temporal Tagging (2022.aacl-main)
Copied to clipboard
| Challenge: | A wide array of natural language processing (NLP) applications relies on accurately identifying events and their respective occurrence times. |
| Approach: | They propose an adversarially trained transformer for Persian temporal tagging that can generalize over the HengamTagger’s rules. |
| Outcome: | The proposed tool outperforms state-of-the-art methods on a diverse and manually created dataset. |