Papers by Ahmed Heakl
LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs (2025.findings-acl)
Copied to clipboard
Omkar Thawakar, Dinura Dissanayake, Ketan Pravin More, Ritesh Thawkar, Ahmed Heakl, Noor Ahsan, Yuhao Li, Ilmuz Zaman Mohammed Zumri, Jean Lahoud, Rao Muhammad Anwer, Hisham Cholakkal, Ivan Laptev, Mubarak Shah, Fahad Shahbaz Khan, Salman Khan
| Challenge: | Existing approaches do not emphasize step-wise problem-solving. |
| Approach: | They propose a visual reasoning chain benchmark and a fine-grained reasoning metric that evaluates correctness and logical coherence at each step. |
| Outcome: | The proposed framework outperforms existing models in six benchmarks and is 5x faster during inference scaling. |
Guaranteed Guess: A Language Modeling Approach for CISC-to-RISC Transpilation with Testing Guarantees (2025.findings-emnlp)
Copied to clipboard
| Challenge: | ISA-centric transpilation pipelines are used to translate low-level programs between ISAs . GG provides high code coverage across unit tests and better energy efficiency . |
| Approach: | They propose a ISA-centric transpilation pipeline that embeds large language models into software testing frameworks to ensure accuracy. |
| Outcome: | The proposed method achieves high code coverage across unit tests and functional/semantic correctness of 99% on HumanEval and 49% on BringupBench programs. |
CAMEL-Bench: A Comprehensive Arabic LMM Benchmark (2025.findings-naacl)
Copied to clipboard
Sara Ghaboura, Ahmed Heakl, Omkar Thawakar, Ali Husain Salem Abdulla Alharthi, Ines Riahi, Abduljalil Radman, Jorma Laaksonen, Fahad Shahbaz Khan, Salman Khan, Rao Muhammad Anwer
| Challenge: | Recent years have witnessed a significant interest in developing large multimodal models capable of performing various visual reasoning and understanding tasks. |
| Approach: | They propose to use Arabic as a language to evaluate large multi-modal models capable of performing visual reasoning and understanding tasks. |
| Outcome: | The proposed benchmark comprises eight diverse domains and 38 sub-domains to represent a large population of over 400 million speakers. |
SAHM: A Benchmark for Arabic Financial and Shari’ah-Compliant Reasoning (2026.acl-long)
Copied to clipboard
Rania Elbadry, Sarfraz Ahmad, Ahmed Heakl, Dani Bouch, Momina Ahsan, Muhra AlMahri, Marwa Elsaid Khalil, Yuxia Wang, Salem Lahlou, Sophia Ananiadou, Veselin Stoyanov, Jimin Huang, Xueqing Peng, Preslav Nakov, Zhuohan Xie
| Challenge: | English financial NLP has progressed rapidly through benchmarks for sentiment, document understanding, and financial question answering. |
| Approach: | They propose a document-grounded benchmark and instruction-tuning dataset for Arabic financial NLP and Shari’ah-compliant reasoning. |
| Outcome: | The proposed dataset contains 14,380 expert-verified instances spanning seven tasks . it includes financial sentiment analysis, extractive summarization, and event–cause reasoning . |
KITAB-Bench: A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding (2025.findings-acl)
Copied to clipboard
Ahmed Heakl, Muhammad Abdullah Sohail, Mukul Ranjan, Rania Elbadry, Ghazi Shazan Ahmad, Mohamed El-Geish, Omar Maher, Zhiqiang Shen, Fahad Shahbaz Khan, Salman Khan
| Challenge: | Optical Character Recognition (OCR) is a key component of document processing . Arabic text recognition has complex typographic and calligraphic features . |
| Approach: | They propose a comprehensive Arabic OCR benchmark that fills the gaps in evaluation systems. |
| Outcome: | The proposed benchmark outperforms existing models in Arabic by 60% in the character error rate . the best model achieves only 65% accuracy in PDF-to-Markdown conversion . |
CASS: Nvidia to AMD Transpilation with Data, Models, and Benchmark (2026.acl-long)
Copied to clipboard
Ahmed Heakl, Gustavo Bertolo Stahl, Sarim Hashmi, Seung Hun Eddie Han, Mukul Ranjan, Arina Kharlamova, Salman Khan, Abdulrahman Mahmoud
| Challenge: | Cross-architecture GPU code translation is essential for unlocking low-level hardware portability, yet no scalable solution exists. |
| Approach: | They propose a dataset and model suite for source- and assembly-level GPU code translation that trains domain-specific translation models that achieve 88.2% accuracy on CUDA HIP and 69.1% on SASS RDNA3 . |
| Outcome: | The proposed model achieves 88.2% accuracy on CUDA HIP and 69.1% on SASS RDNA3 outperforming commercial baselines including GPT-5.1, Claude-4.5, and Hipify by wide margins. |
MASEval: Extending Multi-Agent Evaluation from Models to Systems (2026.acl-demo)
Copied to clipboard
Cornelius Emde, Alexander Rubinstein, Anmol Goel, Ahmed Heakl, Sangdoo Yun, Seong Joon Oh, Martin Gubri
| Challenge: | MASEval provides a framework-agnostic, system-level comparison across any agent framework and benchmark. |
| Approach: | They propose a Python library that treats the entire agentic system as the unit of analysis. |
| Outcome: | The proposed framework treats the entire agentic system as the unit of analysis. |