Papers by Mehar Bhatia

4 papers
From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Vision-Language Models (VLMs) have shown emerging capabilities through large-scale training that have made them gain popularity in recent years.
Approach: They propose to perform retrieval across universals and cultural visual grounding tasks to assess cultural diversity across universal and culture-specific local concepts.
Outcome: The proposed benchmarks show that the models perform significantly across cultures, underscoring the need for enhancing multicultural understanding in vision-language models.
CulturalBench: A Robust, Diverse and Challenging Benchmark for Measuring LMs’ Cultural Knowledge Through Human-AI Red-Teaming (2025.acl-long)

Copied to clipboard

Challenge: CulturalBench is a set of 1,696 human-written and human-verified questions to assess LMs’ cultural knowledge covering 45 global regions including underrepresented ones like Bangladesh, Zimbabwe, and Peru.
Approach: They construct a set of 1,696 human-written and human-verified questions to assess LMs' cultural knowledge, covering 45 global regions including underrepresented ones like Bangladesh, Zimbabwe, and Peru.
Outcome: The proposed model outperforms other models across cultures, while underperforming on questions related to North Africa, South America and Middle East.
CulturalFrames: Assessing Cultural Expectation Alignment in Text-to-Image Models and Evaluation Metrics (2025.findings-emnlp)

Copied to clipboard

Challenge: CulturalFrames is a benchmark designed for rigorous human evaluation of cultural representation in visual generations.
Approach: They propose to quantify the alignment of T2I models and evaluation metrics with respect to both explicit (stated) and implicit (unstated, implied by the prompt’s cultural context) cultural expectations.
Outcome: The proposed model is based on 983 prompts, 3637 images and 10k human annotations from 10 countries and 5 socio-cultural domains.
GD-COMET: A Geo-Diverse Commonsense Inference Model (2023.emnlp-main)

Copied to clipboard

Challenge: GD-COMET is a geo-diverse version of the COMET commonsense inference model . it captures and generates culturally nuanced commonsensense knowledge . lack of cultural awareness may lead to models perpetuating stereotypes and reinforcing societal inequalities for users from non-Western countries.
Approach: They propose a geo-diverse version of COMET commonsense reasoning model that generates inferences pertaining to a broad range of cultures.
Outcome: The proposed model generates inferences pertaining to a broad range of cultures and is culturally nuanced.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations