Papers by Arijit Maji

2 papers
DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models’ Understanding on Indian Culture (2025.emnlp-main)

Copied to clipboard

Challenge: DRISHTIKON is a first-of-its-kind multimodal and multilingual benchmark centered exclusively on Indian culture.
Approach: They evaluate a wide range of vision-language models across zero-shot and chain-of-thought settings and use them to evaluate cultural understanding of generative AI systems.
Outcome: The DRISHTIKON dataset covers 15 languages, all states and union territories, and incorporating over 64,000 aligned text-image pairs.
SANSKRITI: A Comprehensive Benchmark for Evaluating Language Models’ Knowledge of Indian Culture (2025.findings-acl)

Copied to clipboard

Challenge: Language models excel in syntactic and semantic analysis, while small language models struggle in region-specific contexts.
Approach: They evaluate SANSKRITI on leading Large Language Models, Indic Language Model, and Small Language Model (SLM) it covers 16 key attributes of Indian culture including rituals and ceremonies, history, tourism, cuisine, dance and music, costume, language, art, festivals, religion, medicine, transport, sports, nightlife and personalities.
Outcome: The SANSKRITI dataset covers 16 attributes of Indian culture . it reveals that many models struggle in region-specific contexts .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations