Papers by Omkar Thawakar

6 papers
Time Travel: A Comprehensive Benchmark to Evaluate LMMs on Historical and Cultural Artifacts (2025.findings-acl)

Copied to clipboard

Challenge: TimeTravel is a benchmark of 10,250 expert-verified historical artifact samples spanning 266 distinct cultures across 10 major historical regions.
Approach: They evaluate contemporary AI models on TimeTravel, highlighting their strengths and identifying areas for improvement.
Outcome: The timeTravel benchmark covers 266 cultures and 10 major historical regions and aims to establish AI as reliable partner in preserving cultural heritage.
LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs (2025.findings-acl)

Copied to clipboard

Challenge: Existing approaches do not emphasize step-wise problem-solving.
Approach: They propose a visual reasoning chain benchmark and a fine-grained reasoning metric that evaluates correctness and logical coherence at each step.
Outcome: The proposed framework outperforms existing models in six benchmarks and is 5x faster during inference scaling.
Arabic Mini-ClimateGPT : A Climate Change and Sustainability Tailored Arabic LLM (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent large language models like ChatGPT and Bard excel in a wide variety of NLP tasks but are not specifically tailored for climate related domain specific information.
Approach: They propose a lightweight Arabic Mini-ClimateGPT that is built on an open-source LLM and specifically fine-tuned on a conversational-style instruction tuning curated Arabic dataset Clima500-Instruct.
Outcome: The proposed model surpasses the baseline LLM in 88.3% of cases during ChatGPT-based evaluation and human expert prefers it over other open-source models.
CAMEL-Bench: A Comprehensive Arabic LMM Benchmark (2025.findings-naacl)

Copied to clipboard

Challenge: Recent years have witnessed a significant interest in developing large multimodal models capable of performing various visual reasoning and understanding tasks.
Approach: They propose to use Arabic as a language to evaluate large multi-modal models capable of performing visual reasoning and understanding tasks.
Outcome: The proposed benchmark comprises eight diverse domains and 38 sub-domains to represent a large population of over 400 million speakers.
Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: a benchmark is designed to assess the comprehension of Arabic poetry by large language models in 12 historical eras.
Approach: They propose a benchmark to assess the comprehension of Arabic poetry by large language models in 12 historical eras.
Outcome: The benchmark assesses the comprehension of Arabic poetry by large language models in 12 historical eras.
DuwatBench: Bridging Language and Visual Heritage through an Arabic Calligraphy Benchmark for Multimodal Understanding (2026.eacl-long)

Copied to clipboard

Challenge: a benchmark of 1,272 samples containing about 1,475 unique words is available for Arabic calligraphy . the dataset reflects real-world challenges in Arabic writing, such as calligraphic variation and artistic distortions .
Approach: They evaluated 13 leading Arabic and multilingual multimodal models and paired them with sentence-level annotations to evaluate their calligraphy models.
Outcome: The benchmark evaluates 13 leading Arabic and multilingual multimodal models . it shows they struggle with calligraphic variation, artistic distortions, and precise visual–text alignment.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations