Papers by Akash Ghosh

15 papers
DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models’ Understanding on Indian Culture (2025.emnlp-main)

Copied to clipboard

Challenge: DRISHTIKON is a first-of-its-kind multimodal and multilingual benchmark centered exclusively on Indian culture.
Approach: They evaluate a wide range of vision-language models across zero-shot and chain-of-thought settings and use them to evaluate cultural understanding of generative AI systems.
Outcome: The DRISHTIKON dataset covers 15 languages, all states and union territories, and incorporating over 64,000 aligned text-image pairs.
A Survey of Multilingual Reasoning in Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: This survey provides the first in-depth review of multilingual reasoning in Language Models.
Approach: This survey provides the first in-depth review of multilingual reasoning in LMs.
Outcome: The present study provides the first in-depth review of multilingual reasoning in LMs.
Parameter-Efficient Instruction Tuning of Large Language Models For Extreme Financial Numeral Labelling (2024.naacl-long)

Copied to clipboard

Challenge: Existing methods to automatically annotate relevant numerals (GAAP metrics) occurring in financial documents are not cost-effective nor scalable.
Approach: They propose a generative paradigm for annotating GAAP metrics with XBRL tags using metric metadata and a parameter efficient model using LoRA.
Outcome: The proposed model outperforms baseline models on two financial numeric labeling datasets and outperformed several strong baseline models.
Infogen: Generating Complex Statistical Infographics from Documents (2025.acl-long)

Copied to clipboard

Challenge: Existing efforts to generate simple charts have focused on generating simple infographics from text-heavy documents.
Approach: They propose to generate statistical infographics composed of multiple sub-charts that are contextually accurate, insightful, and visually aligned.
Outcome: The proposed framework outperforms both open-source and closed LLMs in text-to-statistical infographic generation.
M3Retrieve: Benchmarking Multimodal Retrieval for Medicine (2025.emnlp-main)

Copied to clipboard

Challenge: Strong retrieval models are increasingly important in knowledge-intensive domains.
Approach: They propose a benchmark to evaluate multimodal retrieval models in medical settings . they examine 1.2 million text documents and 164K multimodal queries .
Outcome: The proposed model spans 5 domains,16 medical fields, and 4 distinct tasks with over 1.2 Million text documents and 164K multimodal queries.
How Robust Are the QA Models for Hybrid Scientific Tabular Data? A Study Using Customized Dataset (2024.lrec-main)

Copied to clipboard

Challenge: Existing tabular QA models are lacking in understanding their robustness on scientific information.
Approach: They propose a dataset to assess the robustness of tabular QA models on scientific hybrid tabular data.
Outcome: The proposed model performs well on scientific tables and text, while the best score is 0.462.
From Sights to Insights: Towards Summarization of Multimodal Clinical Documents (2024.acl-long)

Copied to clipboard

Challenge: a recent WHO report highlights a drastic doctor-to-patient ratio . telehealth is one of the most impactful sectors where AI advances can bring a significant revolution .
Approach: They propose an image-guided encoder-decoder model that uses contextual attention to create detailed visual-guides for multimodal documents.
Outcome: The proposed model outperforms state-of-the-art models on multimodal question and dialogue summarization tasks.
SANSKRITI: A Comprehensive Benchmark for Evaluating Language Models’ Knowledge of Indian Culture (2025.findings-acl)

Copied to clipboard

Challenge: Language models excel in syntactic and semantic analysis, while small language models struggle in region-specific contexts.
Approach: They evaluate SANSKRITI on leading Large Language Models, Indic Language Model, and Small Language Model (SLM) it covers 16 key attributes of Indian culture including rituals and ceremonies, history, tourism, cuisine, dance and music, costume, language, art, festivals, religion, medicine, transport, sports, nightlife and personalities.
Outcome: The SANSKRITI dataset covers 16 attributes of Indian culture . it reveals that many models struggle in region-specific contexts .
CURE-Med: Curriculum-Informed Reinforcement Learning for Multilingual Medical Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have produced strong performance in mathematical reasoning and code generation, but medical reasoning remains challenging because it requires domain knowledge.
Approach: They propose a multilingual medical reasoning dataset with open-ended reasoning queries with a single verifiable answer that spans thirteen languages.
Outcome: The proposed framework outperforms baselines and scales effectively across thirteen languages.
Let’s Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models’ Understanding of Sports (2025.emnlp-main)

Copied to clipboard

Challenge: Language Models (LMs) are primarily evaluated on globally popular sports, often overlooking regional and indigenous sporting traditions.
Approach: They propose to use multiple-choice questions (MCQs) to assess LMs' understanding of traditional sports across 60 countries and 6 continents.
Outcome: The new benchmark will be publicly available, fostering research in culturally aware AI systems.
ECTSum: A New Benchmark Dataset For Bullet Point Summarization of Long Earnings Call Transcripts (2022.emnlp-main)

Copied to clipboard

Challenge: ECTSum is a dataset for bullet-point summarization of earnings calls hosted by publicly traded companies.
Approach: They propose a dataset with transcripts of earnings calls and bullet point summaries derived from Reuters articles.
Outcome: The proposed dataset compares transcripts of earnings calls hosted by publicly traded companies with experts-written bullet point summaries derived from Reuters articles .
When Background Matters: Breaking Medical Vision Language Models by Transferable Attack (2026.acl-long)

Copied to clipboard

Challenge: Existing medical attacks focus on secondary objectives such as model stealing or adversarial fine-tuning, while transferable attacks from natural images introduce visible distortions that clinicians can easily detect. Existing transferable adversarials are less effective in the medical domain.
Approach: They propose a highly transferable black-box multimodal attack that induces incorrect yet clinically plausible diagnoses while keeping perturbations imperceptible.
Outcome: The proposed method induces incorrect yet clinically plausible diagnoses while keeping perturbations imperceptible.
RELIC: Enhancing Reward Model Generalization for Low-Resource Indic Languages with Few-Shot Examples (2025.findings-emnlp)

Copied to clipboard

Challenge: a new reward model for low-resource Indic languages is proposed . a preference-based training approach is prohibitively expensive, authors say .
Approach: a new in-context learning framework is proposed to train a retriever to select in-constext examples from low-resource Indic languages.
Outcome: a new in-context learning framework for reward modeling in low-resource Indic languages is developed . the proposed framework outperforms existing examples on three preference datasets .
HealthAlignSumm : Utilizing Alignment for Multimodal Summarization of Code-Mixed Healthcare Dialogues (2024.findings-emnlp)

Copied to clipboard

Challenge: Collaboration between doctors and AI scientists is leading to personalized models to stream-line healthcare tasks and improve productivity.
Approach: They propose to use alignment techniques to combine a doctor-patient dialogue with a visual component of the BART model.
Outcome: The proposed model in-tegrates visual components with the BART ar-chitecture.
A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models (2024.findings-emnlp)

Copied to clipboard

Challenge: a growing need to understand and alleviate FMs' propensity to produce hallucinated outputs, especially in high-stakes applications.
Approach: They propose a framework for detecting and mitigating hallucination in FMs . they synthesize recent advancements in detection and mitigation techniques .
Outcome: The proposed framework provides valuable insights for researchers, developers, and practitioners.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations