Papers with text-to-image

13 papers
Lost in Translation? Translation Errors and Challenges for Fair Assessment of Text-to-Image Models on Multilingual Concepts (2024.naacl-short)

Copied to clipboard

Challenge: With growth in the popularity of text-to-image models has come interest in assessing their multilingual capabilities, including multilingual accessibility.
Approach: They propose to correct translation errors in a concept list translated to seven languages and compare the outputs of the benchmark to those conditioned on the old.
Outcome: The proposed benchmark contains translation errors in Spanish, Japanese, and Chinese.
A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have highlighted the existence of social biases within large vision and language models.
Approach: They propose a framework for systematically evaluating gender, race, and age biases in vision-language models with respect to professions.
Outcome: The proposed framework covers all supported inference modes of the recent vision-language models, including image-to-text, text-to image, and image- to-image.
Mitigate One, Skew Another? Tackling Intersectional Biases in Text-to-Image Models (2025.findings-emnlp)

Copied to clipboard

Challenge: a new tool for analyzing and quantifying bias interactions in text-to-image models is being developed . a bias in text models can be deeply interrelated, but measuring such effects quantitatively remains a challenge.
Approach: They propose a tool to quantify bias interactions in text-to-image models by analyzing and quantifying bias interactions along bias axes.
Outcome: a new tool analyzes and quantifies bias interactions in text-to-image models . estimates show strong correlations with observed post-mitigation outcomes .
Collaborative Generative AI: Integrating GPT-k for Efficient Editing in Text-to-Image Generation (2023.emnlp-main)

Copied to clipboard

Challenge: Experimental results show that GPT-k models focus more on inserting modifiers than predicting spontaneous changes in the primary subject matter.
Approach: They compare the common edits made by humans and GPT-k models to examine their performance in prompting T2I.
Outcome: The proposed models improve the prompt editing process by 20-30%, the authors show . they show that humans tend to replace words and phrases with modifiers .
Adversarial Attacks on Parts of Speech: An Empirical Study in Text-to-Image Generation (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent studies show that text-to-image models are vulnerable to adversarial perturbations .
Approach: They investigate the impact of adversarial attacks on different POS tags within text prompts on T2I models.
Outcome: The proposed model is vulnerable to adversarial perturbations with noun perturbations in text prompts.
UniCM: A Unified Consistency Model For Efficient Multimodal Generation and Understanding (2026.findings-acl)

Copied to clipboard

Challenge: Consistency models (CMs) have shown promise in the efficient generation of both image and text.
Approach: They propose to use a discrete token for both image and text generation to achieve a unified denoising perspective.
Outcome: The proposed model outperforms SD3 on GenEval and Image Reward while being 1.5 faster at long-sequence generating speed.
CulturalFrames: Assessing Cultural Expectation Alignment in Text-to-Image Models and Evaluation Metrics (2025.findings-emnlp)

Copied to clipboard

Challenge: CulturalFrames is a benchmark designed for rigorous human evaluation of cultural representation in visual generations.
Approach: They propose to quantify the alignment of T2I models and evaluation metrics with respect to both explicit (stated) and implicit (unstated, implied by the prompt’s cultural context) cultural expectations.
Outcome: The proposed model is based on 983 prompts, 3637 images and 10k human annotations from 10 countries and 5 socio-cultural domains.
Misalignment Attack on Text-to-Image Models via Text Embedding Optimization and Inversion (2025.findings-emnlp)

Copied to clipboard

Challenge: Text embedding is a key component of modern NLP models but also poses additional risks.
Approach: They propose a framework that optimizes embeddings and inverts them to obtain misaligned prompts.
Outcome: The proposed framework exploits the continuity and distribution characteristics of text embeddings to obtain misaligned prompts of discrete tokens.
Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures (2025.acl-long)

Copied to clipboard

Challenge: a dataset of 288 gesture-country pairs is used to evaluate AI systems' cultural awareness of offensive gestures and nonverbal signs.
Approach: They use a dataset of 288 gesture-country pairs annotated for offensiveness, cultural significance, and contextual factors across 25 gestures and 85 countries.
Outcome: The proposed dataset analyzes 288 gesture-country pairs across 25 gestures and 85 countries.
T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts (2025.acl-long)

Copied to clipboard

Challenge: Existing studies on text-to-image (T2I) models focus on text alignment, image quality, and object composition capabilities.
Approach: They propose a T2I-FactualBench benchmark to evaluate the factuality of knowledge-intensive concept generation.
Outcome: The proposed framework evaluates the factuality of knowledge-intensive concept generation tasks.
HealthCards: Exploring Text-to-Image Generation as Visual Aids for Healthcare Knowledge Democratizing and Education (2025.emnlp-main)

Copied to clipboard

Challenge: Text-to-image (T2I) generation has the potential to advance knowledge democratization and education.
Approach: They explore ways to harness T2I models for generating health knowledge flashcards . they curated a high-quality healthcare knowledge flash card dataset .
Outcome: The proposed models can generate health knowledge flashcards with appealing images . the results show that the open-source models can be fine tuned to generate health content .
Multimodal Large Language Models for Multi-Subject In-Context Image Generation (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in text-to-image (T2I) generation have enabled visually coherent image synthesis from descriptions, but generating images containing multiple given subjects remains challenging.
Approach: They propose a model that enables automatic and scalable data generation without manual annotations to overcome the data scarcity.
Outcome: The proposed model overcomes the data scarcity and lacks manual annotations.
REVEALER: Reinforcement-Guided Visual Reasoning for Element-Level Text-Image Alignment Evaluation (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for text-to-image alignment evaluation rely on coarse-grained metrics or static Question Answering pipelines that lack fine-grounded interpretability and struggle to reflect human preferences.
Approach: They propose a reinforcement-guided visual reasoning framework for element-level text-to-image alignment evaluation.
Outcome: The proposed framework achieves state-of-the-art results on four benchmarks and surpasses the strong proprietary Gemini 3 Pro and Training-based baselines.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations