Papers by Omkar Thawakar
Time Travel: A Comprehensive Benchmark to Evaluate LMMs on Historical and Cultural Artifacts (2025.findings-acl)
Copied to clipboard
Sara Ghaboura, Ketan Pravin More, Ritesh Thawkar, Wafa Al Ghallabi, Omkar Thawakar, Fahad Shahbaz Khan, Hisham Cholakkal, Salman Khan, Rao Muhammad Anwer
| Challenge: | TimeTravel is a benchmark of 10,250 expert-verified historical artifact samples spanning 266 distinct cultures across 10 major historical regions. |
| Approach: | They evaluate contemporary AI models on TimeTravel, highlighting their strengths and identifying areas for improvement. |
| Outcome: | The timeTravel benchmark covers 266 cultures and 10 major historical regions and aims to establish AI as reliable partner in preserving cultural heritage. |
LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs (2025.findings-acl)
Copied to clipboard
Omkar Thawakar, Dinura Dissanayake, Ketan Pravin More, Ritesh Thawkar, Ahmed Heakl, Noor Ahsan, Yuhao Li, Ilmuz Zaman Mohammed Zumri, Jean Lahoud, Rao Muhammad Anwer, Hisham Cholakkal, Ivan Laptev, Mubarak Shah, Fahad Shahbaz Khan, Salman Khan
| Challenge: | Existing approaches do not emphasize step-wise problem-solving. |
| Approach: | They propose a visual reasoning chain benchmark and a fine-grained reasoning metric that evaluates correctness and logical coherence at each step. |
| Outcome: | The proposed framework outperforms existing models in six benchmarks and is 5x faster during inference scaling. |
Arabic Mini-ClimateGPT : A Climate Change and Sustainability Tailored Arabic LLM (2023.findings-emnlp)
Copied to clipboard
Sahal Mullappilly, Abdelrahman Shaker, Omkar Thawakar, Hisham Cholakkal, Rao Anwer, Salman Khan, Fahad Khan
| Challenge: | Recent large language models like ChatGPT and Bard excel in a wide variety of NLP tasks but are not specifically tailored for climate related domain specific information. |
| Approach: | They propose a lightweight Arabic Mini-ClimateGPT that is built on an open-source LLM and specifically fine-tuned on a conversational-style instruction tuning curated Arabic dataset Clima500-Instruct. |
| Outcome: | The proposed model surpasses the baseline LLM in 88.3% of cases during ChatGPT-based evaluation and human expert prefers it over other open-source models. |
CAMEL-Bench: A Comprehensive Arabic LMM Benchmark (2025.findings-naacl)
Copied to clipboard
Sara Ghaboura, Ahmed Heakl, Omkar Thawakar, Ali Husain Salem Abdulla Alharthi, Ines Riahi, Abduljalil Radman, Jorma Laaksonen, Fahad Shahbaz Khan, Salman Khan, Rao Muhammad Anwer
| Challenge: | Recent years have witnessed a significant interest in developing large multimodal models capable of performing various visual reasoning and understanding tasks. |
| Approach: | They propose to use Arabic as a language to evaluate large multi-modal models capable of performing visual reasoning and understanding tasks. |
| Outcome: | The proposed benchmark comprises eight diverse domains and 38 sub-domains to represent a large population of over 400 million speakers. |
Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs (2025.emnlp-main)
Copied to clipboard
Wafa Al Ghallabi, Ritesh Thawkar, Sara Ghaboura, Ketan Pravin More, Omkar Thawakar, Hisham Cholakkal, Salman Khan, Rao Muhammad Anwer
| Challenge: | a benchmark is designed to assess the comprehension of Arabic poetry by large language models in 12 historical eras. |
| Approach: | They propose a benchmark to assess the comprehension of Arabic poetry by large language models in 12 historical eras. |
| Outcome: | The benchmark assesses the comprehension of Arabic poetry by large language models in 12 historical eras. |
DuwatBench: Bridging Language and Visual Heritage through an Arabic Calligraphy Benchmark for Multimodal Understanding (2026.eacl-long)
Copied to clipboard
Shubham Patle, Sara Ghaboura, Hania Tariq, Mohammad Usman Khan, Omkar Thawakar, Rao Muhammad Anwer, Salman Khan
| Challenge: | a benchmark of 1,272 samples containing about 1,475 unique words is available for Arabic calligraphy . the dataset reflects real-world challenges in Arabic writing, such as calligraphic variation and artistic distortions . |
| Approach: | They evaluated 13 leading Arabic and multilingual multimodal models and paired them with sentence-level annotations to evaluate their calligraphy models. |
| Outcome: | The benchmark evaluates 13 leading Arabic and multilingual multimodal models . it shows they struggle with calligraphic variation, artistic distortions, and precise visual–text alignment. |