Follow the Flow: On Information Flow Across Textual Tokens in Text-to-Image Models (2026.acl-long)
Copied to clipboard
| Challenge: | Prior work focused on improving alignment by refining the diffusion process, ignoring the role of the text encoder, which guides the diffusion. |
| Approach: | They investigate how semantic information is distributed across token representations in text-to-image prompts by patching techniques to uncover encoding patterns. |
| Outcome: | The proposed model can improve alignment and generation quality by modifying the diffusion stage and the cross-attention mechanism. |
Similar Papers
Padding Tone: A Mechanistic Analysis of Padding Tokens in T2I Models (2025.naacl-long)
Copied to clipboard
| Challenge: | Text-to-image (T2I) diffusion models rely on encoded prompts to guide the image generation process. |
| Approach: | They conduct the first in-depth analysis of the role padding tokens play in T2I diffusion models by using two causal techniques to analyze how information is encoded in the representation of tokens across different components of the pipeline. |
| Outcome: | The proposed techniques reveal that padding tokens may affect the model’s output during text encoding, during the diffusion process, or be effectively ignored. |
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines (2024.acl-long)
Copied to clipboard
| Challenge: | Text-to-image diffusion models use a latent text prompt to guide image generation . however, the process by which the encoder produces the text representation is unknown . |
| Approach: | They propose a method for analyzing the text encoder of T2I models by generating images from its intermediate representations. |
| Outcome: | The proposed method provides valuable insights into the text encoder component in T2I pipelines. |
Mechanistic Interpretability of Text-to-Image Diffusion Models via Cross-Attention Interventions (2026.findings-acl)
Copied to clipboard
| Challenge: | Text-to-image diffusion models generate high quality images through iterative denoising, but their internal mechanisms for grounding prompt semantics into visual structure remain unclear. |
| Approach: | They propose a mechanistic interpretability framework that probes how individual prompt tokens are represented and utilized during the denoising process. |
| Outcome: | The proposed framework enables module-wise and head-wise attribution of semantic changes across denoising timesteps. |
Prompt Expansion for Adaptive Text-to-Image Generation (2024.acl-long)
Copied to clipboard
| Challenge: | Text-to-image generation models are powerful but difficult to use. Users craft specific prompts to get better images, though the prompts can be repetitive. |
| Approach: | They propose a framework that takes a text query as input and outputs a set of expanded text prompts that are optimized to generate a wider variety of appealing images. |
| Outcome: | The proposed framework generates high-quality images from text prompts with less effort and is more aesthetically pleasing than baseline models. |
Uncovering Limitations in Text-to-Image Generation: A Contrastive Approach with Structured Semantic Alignment (2023.findings-emnlp)
Copied to clipboard
| Challenge: | a new method for text-to-image generation models is proposed to address these limitations . SSA focuses on learning structured semantic embeddings across different modalities . |
| Approach: | They propose a method to evaluate text-to-image generation models using structured semantic embeddings . they propose to learn mutated prompts by substituting words with equivalent or nonequivalent alternatives . |
| Outcome: | The proposed method improves the measurement of semantic consistency of text-to-image generation models. |
Pre-trained Language Models Do Not Help Auto-regressive Text-to-Image Generation (2024.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in image tokenizers have enabled text-to-image generation using auto-regressive methods, but these methods lack pre-trained language models for text-based models. |
| Approach: | They adapt a pre-trained language model for auto-regressive text-to-image generation and show that pre-train language models offer limited help. |
| Outcome: | The proposed model is compared with a pre-trained language model and shows that it is no more effective than random initialized models. |
STAIR: Learning Sparse Text and Image Representation in Grounded Tokens (2023.emnlp-main)
Copied to clipboard
Chen Chen, Bowen Zhang, Liangliang Cao, Jiguang Shen, Tom Gunter, Albin Jose, Alexander Toshev, Yantao Zheng, Jonathon Shlens, Ruoming Pang, Yinfei Yang
| Challenge: | State-of-the-art contrastive learning models like CLIP and ALIGN are less interpretable and suffer from inferior accuracy than dense representations. |
| Approach: | They extend CLIP and ALIGN models to build a sparse semantic representation that is interpretable and easy to integrate with existing retrieval systems. |
| Outcome: | The proposed model outperforms CLIP and ALIGN models on image and text retrieval tasks with a 4.9% and +4.3% improvement on COCO-5k textimage and imagetext retrieval respectively. |
Token Alignment via Character Matching for Subword Completion (2024.findings-acl)
Copied to clipboard
Ben Athiwaratkun, Shiqi Wang, Mingyue Shang, Yuchen Tian, Zijian Wang, Sujan Kumar Gonugondla, Sanjay Krishna Gouda, Robert Kwiatkowski, Ramesh Nallapati, Parminder Bhatia, Bing Xiang
| Challenge: | Generative models struggle with prompts corresponding to partial tokens due to tokenization, where partial token is out-of-distribution during inference. |
| Approach: | They propose a method to alleviate tokenization artifact on text completion by backtracking to the last complete tokens and aligning subsequent generations to match with the prompt. |
| Outcome: | The proposed method shows that it improves on partial token scenarios with only a minor time increase. |
DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models (2023.acl-long)
Copied to clipboard
| Challenge: | Recent advances in diffusion models have enabled high-quality image generation . generating images with desired details requires proper prompts . |
| Approach: | They analyze syntactic and semantic characteristics of diffusion models and their prompts . they pinpoint specific hyperparameter values and prompt styles that can lead to model errors . |
| Outcome: | The first large-scale text-to-image prompt dataset totals 6.5TB . it contains 14 million images generated by Stable Diffusion, 1.8 million unique prompts, and hyperparameters specified by real users. |
Underspecification in Scene Description-to-Depiction Tasks (2022.aacl-main)
Copied to clipboard
| Challenge: | Recent text-to-image generation systems have demonstrated impressive capabilities . recent work focuses on generating images depicting scenes from scene descriptions . |
| Approach: | They propose a conceptual framework to address implicitness, ambiguity and underspecification issues in multimodal image+text systems. |
| Outcome: | The proposed framework addresses key challenges concerning textual and visual ambiguity and risks that may be amplified by ambiguous and underspecified elements. |