| Challenge: | Existing models do not analyze human preferences at a finer granularity, which leads to quality issues. |
| Approach: | They propose a set of preference indicators across two major dimensions, text-image consistency and aesthetic quality, and a generative framework to steer the model toward a generation path that more closely aligns with human aesthetic sensibilities. |
| Outcome: | The proposed model improves target recognition accuracy and overall visual aesthetic presentation by focusing on human preferences. |
Similar Papers
Sentimental Image Generation for Aspect-based Sentiment Analysis (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent work on textual Aspect-Based Sentiment Analysis (ABSA) has demonstrated promising performance, but limited semantics derived from raw data. |
| Approach: | They propose a method that provides visual semantics to reinforce textual ABSA by adding additional augmentations to the input data. |
| Outcome: | The proposed method can provide visual semantics to reinforce the textual extraction. |
Impressions: Visual Semiotics and Aesthetic Impact Understanding (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing image captioning and conditional generation models struggle to simulate plausible human responses to images. |
| Approach: | They propose a dataset to investigate the semiotics of images and how visual features and design choices can elicit specific emotions, thoughts and beliefs. |
| Outcome: | The proposed dataset improves existing models for image captioning and conditional generation. |
Textual Aesthetics in Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies on image aesthetics have focused on content correctness and helpfulness of responses. |
| Approach: | They propose a textual aesthetics-powered fine-tuning method that leverages textual visual aesthetics without compromising content correctness. |
| Outcome: | The proposed method improves aesthetic scores and performs well on general evaluation datasets. |
Aspect-based Sentiment Analysis via Synthetic Image Generation (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in Aspect-Based Sentiment Analysis (ABSA) have shown promising results, yet the semantics derived solely from textual data remain limited. |
| Approach: | They propose a supervised image generation framework to generate synthetic images with alignment to text and sentiment information. |
| Outcome: | The proposed approach significantly outperforms state-of-the-art methods on multiple benchmark datasets. |
How well can Text-to-Image Generative Models understand Ethical Natural Language Interventions? (2022.emnlp-main)
Copied to clipboard
| Challenge: | Text-to-image generative models can generate high-quality photo-realistic images conditional on natural language text descriptions in a zero-shot fashion. |
| Approach: | They propose an Ethical NaTural Language Interventions in Text-to-Image GENeration benchmark dataset to evaluate the change in image generation conditional on ethical interventions across three social axes – gender, skin color, and culture. |
| Outcome: | The proposed model generations cover diverse social groups while preserving image quality. |
DreamSync: Aligning Text-to-Image Generation with Image Understanding Feedback (2025.naacl-long)
Copied to clipboard
Jiao Sun, Deqing Fu, Yushi Hu, Su Wang, Royi Rassin, Da-Cheng Juan, Dana Alon, Charles Herrmann, Sjoerd Van Steenkiste, Ranjay Krishna, Cyrus Rashtchian
| Challenge: | Text-to-Image models (T2I) still struggle to produce images that are both aesthetically pleasing and faithful to the user’s input text. |
| Approach: | They propose a training algorithm that trains T2I models to be faithful to the input text. |
| Outcome: | The proposed model improves both the semantic alignment and aesthetic appeal of two diffusion-based T2I models, evidenced by multiple benchmarks (+1.7% on TIFA, +2.9% on DSG1K, +3.4% on VILA aesthetic). |
Uncovering Limitations in Text-to-Image Generation: A Contrastive Approach with Structured Semantic Alignment (2023.findings-emnlp)
Copied to clipboard
| Challenge: | a new method for text-to-image generation models is proposed to address these limitations . SSA focuses on learning structured semantic embeddings across different modalities . |
| Approach: | They propose a method to evaluate text-to-image generation models using structured semantic embeddings . they propose to learn mutated prompts by substituting words with equivalent or nonequivalent alternatives . |
| Outcome: | The proposed method improves the measurement of semantic consistency of text-to-image generation models. |
The Face of Persuasion: Analyzing Bias and Generating Culture-Aware Ads (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Text-to-image models are appealing for customizing visual ads and targeting specific populations. |
| Approach: | We examine the disparate level of persuasiveness of ads that are identical except for gender/race of the people portrayed. |
| Outcome: | The proposed technique is based on a demographic bias analysis of ads for different topics and a disparate level of persuasiveness of ads that are identical except for gender/race of the people portrayed. |
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment (2025.emnlp-main)
Copied to clipboard
Zhipeng Bian, Jieming Zhu, Qijiong Liu, Wang Lin, Guohao Cai, Zhaocheng Du, Jiacheng Sun, Zhou Zhao, Zhenhua Dong
| Challenge: | Large language models and diffusion models have opened new possibilities for AI-generated content . personalized cover image generation remains underexplored despite its critical role in boosting user engagement on digital platforms. |
| Approach: | They propose a framework that integrates MLLM-based prompting with personalized preference alignment to generate high-quality, contextually relevant covers. |
| Outcome: | The proposed framework improves image quality, semantic fidelity, and personalization, leading to stronger user appeal and offline recommendation accuracy in downstream tasks. |
Prompt Expansion for Adaptive Text-to-Image Generation (2024.acl-long)
Copied to clipboard
| Challenge: | Text-to-image generation models are powerful but difficult to use. Users craft specific prompts to get better images, though the prompts can be repetitive. |
| Approach: | They propose a framework that takes a text query as input and outputs a set of expanded text prompts that are optimized to generate a wider variety of appealing images. |
| Outcome: | The proposed framework generates high-quality images from text prompts with less effort and is more aesthetically pleasing than baseline models. |