Challenge: Existing image captioning and conditional generation models struggle to simulate plausible human responses to images.
Approach: They propose a dataset to investigate the semiotics of images and how visual features and design choices can elicit specific emotions, thoughts and beliefs.
Outcome: The proposed dataset improves existing models for image captioning and conditional generation.

Similar Papers

AesX: Enhance Your Images with Stunning Aesthetic Beauty (2026.acl-industry)

Copied to clipboard

Challenge: Existing models do not analyze human preferences at a finer granularity, which leads to quality issues.
Approach: They propose a set of preference indicators across two major dimensions, text-image consistency and aesthetic quality, and a generative framework to steer the model toward a generation path that more closely aligns with human aesthetic sensibilities.
Outcome: The proposed model improves target recognition accuracy and overall visual aesthetic presentation by focusing on human preferences.
WikiArt Emotions: An Annotated Dataset of Emotions Evoked by Art (L18-1)

Copied to clipboard

Challenge: a dataset of 4,000 pieces of art has annotations for emotions evoked in the observer . the dataset can help answer questions about what makes art evocative, how does art convey different emotions, what attributes of a painting make it well liked, and how much does the title impact the affectual response to art.
Approach: They create a dataset of 4,000 western art pieces that has annotations for emotions . they use crowdsourcing to annotate the art for one or more of twenty emotion categories . fear, happiness, love, sadness were the dominant emotions that obtained consistent annotations .
Outcome: The dataset shows that the most popular emotions are fear, happiness, love and sadness . the dataset can be used to develop systems that detect emotions evoked by art .
EmoTag1200: Understanding the Association between Emojis and Emotions (2020.emnlp-main)

Copied to clipboard

Challenge: Emojis are increasingly used to convey affect, but their use is not trivial.
Approach: They propose to use human-solicited association ratings to explore the connection between emojis and emotions to conduct experiments.
Outcome: The proposed method can be inferred from word-level information when high-quality information is available.
PO-EMO: Conceptualization, Annotation, and Modeling of Aesthetic Emotions in German and English Poetry (2020.lrec-1)

Copied to clipboard

Challenge: a new study shows that literature enables engagement in a broader range of complex and subtle emotions.
Approach: They propose to use multiple emotion labels to capture mixed emotions in poetry . they evaluate an annotation experiment with experts and crowdsourcing .
Outcome: The proposed method shows that identifying aesthetic emotions is challenging in the German subset.
PVP: An Image Dataset for Personalized Visual Persuasion with Persuasion Strategies, Viewer Characteristics, and Persuasiveness Ratings (2025.acl-long)

Copied to clipboard

Challenge: Visual persuasion uses visual elements to influence cognition and behaviors . lack of comprehensive data sets connect persuasiveness of images with personal information .
Approach: They propose to use a dataset to connect persuasiveness with personal information . they find psychological characteristics enhance the generation and evaluation of persuasive images .
Outcome: The proposed dataset provides persuasiveness scores of images evaluated by human annotators along with demographic and psychological characteristics.
FeelingBlue: A Corpus for Understanding the Emotional Connotation of Color in Context (2023.tacl-1)

Copied to clipboard

Challenge: Experimental results shed light on the emotional connotation of color in context . color is a powerful tool for conveying emotion across cultures .
Approach: They propose a multimodal dataset for exploring the emotional connotation of color as mediated by line, stroke, texture, shape, and language.
Outcome: The proposed model sheds light on the emotional connotation of color in context and the potential for future studies.
Vanessa: Visual Connotation and Aesthetic Attributes Understanding Network for Multimodal Aspect-based Sentiment Analysis (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to analyze images focus on superficial features or descriptions, omitting subtle contextual information.
Approach: They propose a Visual Connotation and Aesthetic Attributes Understanding Network (Vanessa) for Multimodal Aspect-based Sentiment Analysis.
Outcome: The proposed network captures both implicit and explicit sentimental cues and can be used to enrich textual sentiment analysis.
Cross-Lingual and Cross-Cultural Variation in Image Descriptions (2025.naacl-long)

Copied to clipboard

Challenge: Behavioural and cognitive studies report cultural effects on perception, but these are limited in scope and hard to replicate.
Approach: They develop a method to accurately identify entities mentioned in captions and present in images, then measure how they vary across languages.
Outcome: The proposed method corroborates previous studies showing that languages that are geographically or genetically closer mention entities more frequently than others.
SNAG: Spoken Narratives and Gaze Dataset (P18-2)

Copied to clipboard

Challenge: Existing datasets that combine gaze and spoken descriptions of visual inputs are needed to provide insight into how humans process information and make decisions.
Approach: They propose a multimodal gaze and spoken descriptions dataset that can be used to label important image regions with appropriate linguistic labels.
Outcome: The proposed dataset can be used to label image regions with appropriate linguistic labels.
The Face of Persuasion: Analyzing Bias and Generating Culture-Aware Ads (2025.findings-emnlp)

Copied to clipboard

Challenge: Text-to-image models are appealing for customizing visual ads and targeting specific populations.
Approach: We examine the disparate level of persuasiveness of ads that are identical except for gender/race of the people portrayed.
Outcome: The proposed technique is based on a demographic bias analysis of ads for different topics and a disparate level of persuasiveness of ads that are identical except for gender/race of the people portrayed.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations