Papers by Mukul Ranjan

3 papers
KITAB-Bench: A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding (2025.findings-acl)

Copied to clipboard

Challenge: Optical Character Recognition (OCR) is a key component of document processing . Arabic text recognition has complex typographic and calligraphic features .
Approach: They propose a comprehensive Arabic OCR benchmark that fills the gaps in evaluation systems.
Outcome: The proposed benchmark outperforms existing models in Arabic by 60% in the character error rate . the best model achieves only 65% accuracy in PDF-to-Markdown conversion .
CASS: Nvidia to AMD Transpilation with Data, Models, and Benchmark (2026.acl-long)

Copied to clipboard

Challenge: Cross-architecture GPU code translation is essential for unlocking low-level hardware portability, yet no scalable solution exists.
Approach: They propose a dataset and model suite for source- and assembly-level GPU code translation that trains domain-specific translation models that achieve 88.2% accuracy on CUDA HIP and 69.1% on SASS RDNA3 .
Outcome: The proposed model achieves 88.2% accuracy on CUDA HIP and 69.1% on SASS RDNA3 outperforming commercial baselines including GPT-5.1, Claude-4.5, and Hipify by wide margins.
On the Cultural Anachronism and Temporal Reasoning in Vision Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Vision-Language Models (VLMs) are increasingly applied to cultural heritage materials.
Approach: They propose a temporal anachronism benchmark to evaluate temporal reasoning on 1,600 Indian cultural artifacts.
Outcome: The proposed model achieves only 58.7% accuracy on the best model, which is a significant performance gap across architectures and scales.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations