Culture-TRIP: Culturally-Aware Text-to-Image Generation with Iterative Prompt Refinement (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing text-to-image models fail to produce appropriate images for cultural concepts or objects not well known or underrepresented in western cultures, such as 'hangari' (a Korean utensil). |
| Approach: | They propose a method which iteratively refines the prompt to improve the alignment between the generated images and underrepresented cultural nouns in text-to-image models. |
| Outcome: | The proposed approach improves the alignment between the generated images and cultural nouns in text-to-image models. |
Similar Papers
Prompt Refinement with Image Pivot for Text-to-Image Generation (2024.acl-long)
Copied to clipboard
| Challenge: | Recent advances in text-to-image generation have markedly expanded the boundaries of digital artistry, enabling the creation of visually compelling images with unprecedented ease. |
| Approach: | They propose to decompose the prompt refinement process into two tasks: inferring user-preferred images from user languages and translating them into system languages. |
| Outcome: | Experiments show that PRIP outperforms baselines and transfers to unseen systems in a zero-shot manner. |
When Cultures Meet: Multicultural Text-to-Image Generation (2026.findings-acl)
Copied to clipboard
| Challenge: | a new task to evaluate text-to-image generation models for multicultural scenes is unexplored. |
| Approach: | They propose a benchmark task to evaluate text-to-image models in multicultural settings . they use a dataset of 9,000 images spanning five countries, three age groups, two genders, 25 historical landmarks, and five languages to analyze behavior . |
| Outcome: | The proposed benchmark analyzes the behavior of state-of-the-art models across multiple dimensions including alignment, image quality, aesthetics, knowledge, and fairness. |
Diffusion Models Through a Global Lens: Are They Culturally Inclusive? (2025.acl-long)
Copied to clipboard
Zahra Bayramli, Ayhan Suleymanzade, Na Min An, Huzama Ahmad, Eunsu Kim, Junyeong Park, James Thorne, Alice Oh
| Challenge: | Text-to-image diffusion models have produced compelling, detailed images from text prompts, but their ability to accurately represent cultural nuances remains an open question. |
| Approach: | They propose a benchmark to evaluate whether diffusion models can generate culturally specific images spanning ten countries. |
| Outcome: | The proposed model fails to generate culturally specific images spanning ten countries . it shows significant disparities in cultural relevance, description fidelity, and realism compared to real-world reference images. |
Prompt Expansion for Adaptive Text-to-Image Generation (2024.acl-long)
Copied to clipboard
| Challenge: | Text-to-image generation models are powerful but difficult to use. Users craft specific prompts to get better images, though the prompts can be repetitive. |
| Approach: | They propose a framework that takes a text query as input and outputs a set of expanded text prompts that are optimized to generate a wider variety of appealing images. |
| Outcome: | The proposed framework generates high-quality images from text prompts with less effort and is more aesthetically pleasing than baseline models. |
Iterative Prompt Refinement for Safer Text-to-Image Generation (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing safety methods for text-to-image models ignore the images produced . this can result in unsafe outputs or unnecessary changes to already safe prompts . |
| Approach: | They propose an iterative prompt refinement algorithm that uses Vision Language Models to analyze prompts and generated images. |
| Outcome: | The proposed method improves safety while maintaining user intent and reliability comparable to existing methods. |
DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models (2023.acl-long)
Copied to clipboard
| Challenge: | Recent advances in diffusion models have enabled high-quality image generation . generating images with desired details requires proper prompts . |
| Approach: | They analyze syntactic and semantic characteristics of diffusion models and their prompts . they pinpoint specific hyperparameter values and prompt styles that can lead to model errors . |
| Outcome: | The first large-scale text-to-image prompt dataset totals 6.5TB . it contains 14 million images generated by Stable Diffusion, 1.8 million unique prompts, and hyperparameters specified by real users. |
NeuroPrompts: An Adaptive Framework to Optimize Prompts for Text-to-Image Generation (2024.eacl-demo)
Copied to clipboard
| Challenge: | Recent advances in text-to-image diffusion models have made it difficult to obtain high-quality images. |
| Approach: | They propose an adaptive framework that automatically enhances a user's prompt to improve the quality of generation models. |
| Outcome: | The proposed framework generates prompts similar to those produced by human prompt engineers and provides user control over stylistic features via constraint set specification. |
BeautifulPrompt: Towards Automatic Prompt Engineering for Text-to-Image Synthesis (2023.emnlp-industry)
Copied to clipboard
| Challenge: | Recent text-to-image models require multiple passes of prompt engineering by humans to produce satisfactory results for real-world applications. |
| Approach: | They propose a deep generative model to generate high-quality prompts from raw descriptions using visual feedback. |
| Outcome: | The proposed model produces high-quality prompts from simple raw descriptions . it can be integrated to a cloud-native AI platform to provide better image generation service in the cloud. |
Follow the Flow: On Information Flow Across Textual Tokens in Text-to-Image Models (2026.acl-long)
Copied to clipboard
| Challenge: | Prior work focused on improving alignment by refining the diffusion process, ignoring the role of the text encoder, which guides the diffusion. |
| Approach: | They investigate how semantic information is distributed across token representations in text-to-image prompts by patching techniques to uncover encoding patterns. |
| Outcome: | The proposed model can improve alignment and generation quality by modifying the diffusion stage and the cross-attention mechanism. |
RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation (2025.findings-naacl)
Copied to clipboard
Viacheslav Vasilev, Julia Agafonova, Nikolai Gerasimenko, Alexander Kapitanov, Polina Mikhailova, Evelina Mironova, Denis Dimitrov
| Challenge: | Text-to-image generation models exhibit a strong bias toward English-speaking cultures, ignoring or misrepresenting the unique characteristics of other language groups, countries, and nationalities. |
| Approach: | They propose a RusCode benchmark to evaluate the quality of text-to-image generation containing elements of the Russian cultural code. |
| Outcome: | The proposed model is based on 1250 text prompts in Russian and their translations into English. |