Challenge: Existing text-to-image models fail to produce appropriate images for cultural concepts or objects not well known or underrepresented in western cultures, such as 'hangari' (a Korean utensil).
Approach: They propose a method which iteratively refines the prompt to improve the alignment between the generated images and underrepresented cultural nouns in text-to-image models.
Outcome: The proposed approach improves the alignment between the generated images and cultural nouns in text-to-image models.

Similar Papers

Prompt Refinement with Image Pivot for Text-to-Image Generation (2024.acl-long)

Copied to clipboard

Challenge: Recent advances in text-to-image generation have markedly expanded the boundaries of digital artistry, enabling the creation of visually compelling images with unprecedented ease.
Approach: They propose to decompose the prompt refinement process into two tasks: inferring user-preferred images from user languages and translating them into system languages.
Outcome: Experiments show that PRIP outperforms baselines and transfers to unseen systems in a zero-shot manner.
When Cultures Meet: Multicultural Text-to-Image Generation (2026.findings-acl)

Copied to clipboard

Challenge: a new task to evaluate text-to-image generation models for multicultural scenes is unexplored.
Approach: They propose a benchmark task to evaluate text-to-image models in multicultural settings . they use a dataset of 9,000 images spanning five countries, three age groups, two genders, 25 historical landmarks, and five languages to analyze behavior .
Outcome: The proposed benchmark analyzes the behavior of state-of-the-art models across multiple dimensions including alignment, image quality, aesthetics, knowledge, and fairness.
Diffusion Models Through a Global Lens: Are They Culturally Inclusive? (2025.acl-long)

Copied to clipboard

Challenge: Text-to-image diffusion models have produced compelling, detailed images from text prompts, but their ability to accurately represent cultural nuances remains an open question.
Approach: They propose a benchmark to evaluate whether diffusion models can generate culturally specific images spanning ten countries.
Outcome: The proposed model fails to generate culturally specific images spanning ten countries . it shows significant disparities in cultural relevance, description fidelity, and realism compared to real-world reference images.
Prompt Expansion for Adaptive Text-to-Image Generation (2024.acl-long)

Copied to clipboard

Challenge: Text-to-image generation models are powerful but difficult to use. Users craft specific prompts to get better images, though the prompts can be repetitive.
Approach: They propose a framework that takes a text query as input and outputs a set of expanded text prompts that are optimized to generate a wider variety of appealing images.
Outcome: The proposed framework generates high-quality images from text prompts with less effort and is more aesthetically pleasing than baseline models.
Iterative Prompt Refinement for Safer Text-to-Image Generation (2025.emnlp-main)

Copied to clipboard

Challenge: Existing safety methods for text-to-image models ignore the images produced . this can result in unsafe outputs or unnecessary changes to already safe prompts .
Approach: They propose an iterative prompt refinement algorithm that uses Vision Language Models to analyze prompts and generated images.
Outcome: The proposed method improves safety while maintaining user intent and reliability comparable to existing methods.
DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models (2023.acl-long)

Copied to clipboard

Challenge: Recent advances in diffusion models have enabled high-quality image generation . generating images with desired details requires proper prompts .
Approach: They analyze syntactic and semantic characteristics of diffusion models and their prompts . they pinpoint specific hyperparameter values and prompt styles that can lead to model errors .
Outcome: The first large-scale text-to-image prompt dataset totals 6.5TB . it contains 14 million images generated by Stable Diffusion, 1.8 million unique prompts, and hyperparameters specified by real users.
NeuroPrompts: An Adaptive Framework to Optimize Prompts for Text-to-Image Generation (2024.eacl-demo)

Copied to clipboard

Challenge: Recent advances in text-to-image diffusion models have made it difficult to obtain high-quality images.
Approach: They propose an adaptive framework that automatically enhances a user's prompt to improve the quality of generation models.
Outcome: The proposed framework generates prompts similar to those produced by human prompt engineers and provides user control over stylistic features via constraint set specification.
BeautifulPrompt: Towards Automatic Prompt Engineering for Text-to-Image Synthesis (2023.emnlp-industry)

Copied to clipboard

Challenge: Recent text-to-image models require multiple passes of prompt engineering by humans to produce satisfactory results for real-world applications.
Approach: They propose a deep generative model to generate high-quality prompts from raw descriptions using visual feedback.
Outcome: The proposed model produces high-quality prompts from simple raw descriptions . it can be integrated to a cloud-native AI platform to provide better image generation service in the cloud.
Follow the Flow: On Information Flow Across Textual Tokens in Text-to-Image Models (2026.acl-long)

Copied to clipboard

Challenge: Prior work focused on improving alignment by refining the diffusion process, ignoring the role of the text encoder, which guides the diffusion.
Approach: They investigate how semantic information is distributed across token representations in text-to-image prompts by patching techniques to uncover encoding patterns.
Outcome: The proposed model can improve alignment and generation quality by modifying the diffusion stage and the cross-attention mechanism.
RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation (2025.findings-naacl)

Copied to clipboard

Challenge: Text-to-image generation models exhibit a strong bias toward English-speaking cultures, ignoring or misrepresenting the unique characteristics of other language groups, countries, and nationalities.
Approach: They propose a RusCode benchmark to evaluate the quality of text-to-image generation containing elements of the Russian cultural code.
Outcome: The proposed model is based on 1250 text prompts in Russian and their translations into English.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations