Papers with text-to-image
Lost in Translation? Translation Errors and Challenges for Fair Assessment of Text-to-Image Models on Multilingual Concepts (2024.naacl-short)
Copied to clipboard
| Challenge: | With growth in the popularity of text-to-image models has come interest in assessing their multilingual capabilities, including multilingual accessibility. |
| Approach: | They propose to correct translation errors in a concept list translated to seven languages and compare the outputs of the benchmark to those conditioned on the old. |
| Outcome: | The proposed benchmark contains translation errors in Spanish, Japanese, and Chinese. |
A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies have highlighted the existence of social biases within large vision and language models. |
| Approach: | They propose a framework for systematically evaluating gender, race, and age biases in vision-language models with respect to professions. |
| Outcome: | The proposed framework covers all supported inference modes of the recent vision-language models, including image-to-text, text-to image, and image- to-image. |
Mitigate One, Skew Another? Tackling Intersectional Biases in Text-to-Image Models (2025.findings-emnlp)
Copied to clipboard
Pushkar Shukla, Aditya Chinchure, Emily Diana, Alexander Tolbert, Kartik Hosanagar, Vineeth N. Balasubramanian, Leonid Sigal, Matthew A. Turk
| Challenge: | a new tool for analyzing and quantifying bias interactions in text-to-image models is being developed . a bias in text models can be deeply interrelated, but measuring such effects quantitatively remains a challenge. |
| Approach: | They propose a tool to quantify bias interactions in text-to-image models by analyzing and quantifying bias interactions along bias axes. |
| Outcome: | a new tool analyzes and quantifies bias interactions in text-to-image models . estimates show strong correlations with observed post-mitigation outcomes . |
Collaborative Generative AI: Integrating GPT-k for Efficient Editing in Text-to-Image Generation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Experimental results show that GPT-k models focus more on inserting modifiers than predicting spontaneous changes in the primary subject matter. |
| Approach: | They compare the common edits made by humans and GPT-k models to examine their performance in prompting T2I. |
| Outcome: | The proposed models improve the prompt editing process by 20-30%, the authors show . they show that humans tend to replace words and phrases with modifiers . |
Adversarial Attacks on Parts of Speech: An Empirical Study in Text-to-Image Generation (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Recent studies show that text-to-image models are vulnerable to adversarial perturbations . |
| Approach: | They investigate the impact of adversarial attacks on different POS tags within text prompts on T2I models. |
| Outcome: | The proposed model is vulnerable to adversarial perturbations with noun perturbations in text prompts. |
UniCM: A Unified Consistency Model For Efficient Multimodal Generation and Understanding (2026.findings-acl)
Copied to clipboard
| Challenge: | Consistency models (CMs) have shown promise in the efficient generation of both image and text. |
| Approach: | They propose to use a discrete token for both image and text generation to achieve a unified denoising perspective. |
| Outcome: | The proposed model outperforms SD3 on GenEval and Image Reward while being 1.5 faster at long-sequence generating speed. |
CulturalFrames: Assessing Cultural Expectation Alignment in Text-to-Image Models and Evaluation Metrics (2025.findings-emnlp)
Copied to clipboard
Shravan Nayak, Mehar Bhatia, Xiaofeng Zhang, Verena Rieser, Lisa Anne Hendricks, Sjoerd Van Steenkiste, Yash Goyal, Karolina Stanczak, Aishwarya Agrawal
| Challenge: | CulturalFrames is a benchmark designed for rigorous human evaluation of cultural representation in visual generations. |
| Approach: | They propose to quantify the alignment of T2I models and evaluation metrics with respect to both explicit (stated) and implicit (unstated, implied by the prompt’s cultural context) cultural expectations. |
| Outcome: | The proposed model is based on 983 prompts, 3637 images and 10k human annotations from 10 countries and 5 socio-cultural domains. |
Misalignment Attack on Text-to-Image Models via Text Embedding Optimization and Inversion (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Text embedding is a key component of modern NLP models but also poses additional risks. |
| Approach: | They propose a framework that optimizes embeddings and inverts them to obtain misaligned prompts. |
| Outcome: | The proposed framework exploits the continuity and distribution characteristics of text embeddings to obtain misaligned prompts of discrete tokens. |
Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures (2025.acl-long)
Copied to clipboard
| Challenge: | a dataset of 288 gesture-country pairs is used to evaluate AI systems' cultural awareness of offensive gestures and nonverbal signs. |
| Approach: | They use a dataset of 288 gesture-country pairs annotated for offensiveness, cultural significance, and contextual factors across 25 gestures and 85 countries. |
| Outcome: | The proposed dataset analyzes 288 gesture-country pairs across 25 gestures and 85 countries. |
T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts (2025.acl-long)
Copied to clipboard
Ziwei Huang, Wanggui He, Quanyu Long, Yandi Wang, Haoyuan Li, Zhelun Yu, Fangxun Shu, Weilong Dai, Hao Jiang, Fei Wu, Leilei Gan
| Challenge: | Existing studies on text-to-image (T2I) models focus on text alignment, image quality, and object composition capabilities. |
| Approach: | They propose a T2I-FactualBench benchmark to evaluate the factuality of knowledge-intensive concept generation. |
| Outcome: | The proposed framework evaluates the factuality of knowledge-intensive concept generation tasks. |
HealthCards: Exploring Text-to-Image Generation as Visual Aids for Healthcare Knowledge Democratizing and Education (2025.emnlp-main)
Copied to clipboard
| Challenge: | Text-to-image (T2I) generation has the potential to advance knowledge democratization and education. |
| Approach: | They explore ways to harness T2I models for generating health knowledge flashcards . they curated a high-quality healthcare knowledge flash card dataset . |
| Outcome: | The proposed models can generate health knowledge flashcards with appealing images . the results show that the open-source models can be fine tuned to generate health content . |
Multimodal Large Language Models for Multi-Subject In-Context Image Generation (2026.acl-long)
Copied to clipboard
| Challenge: | Recent advances in text-to-image (T2I) generation have enabled visually coherent image synthesis from descriptions, but generating images containing multiple given subjects remains challenging. |
| Approach: | They propose a model that enables automatic and scalable data generation without manual annotations to overcome the data scarcity. |
| Outcome: | The proposed model overcomes the data scarcity and lacks manual annotations. |
REVEALER: Reinforcement-Guided Visual Reasoning for Element-Level Text-Image Alignment Evaluation (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods for text-to-image alignment evaluation rely on coarse-grained metrics or static Question Answering pipelines that lack fine-grounded interpretability and struggle to reflect human preferences. |
| Approach: | They propose a reinforcement-guided visual reasoning framework for element-level text-to-image alignment evaluation. |
| Outcome: | The proposed framework achieves state-of-the-art results on four benchmarks and surpasses the strong proprietary Gemini 3 Pro and Training-based baselines. |