Challenge: Prior work on multimodal content classification has not addressed these challenges.
Approach: They propose to use two auxiliary tasks to fine-tune multimodal models to address hidden cross-modal semantics and weak image-text relationships when modeling text and images.
Outcome: The proposed model improves by up to 2.6 F1 score across five diverse social media datasets.

Similar Papers

Borrowing Human Senses: Comment-Aware Self-Training for Social Media Multimodal Classification (2022.emnlp-main)

Copied to clipboard

Challenge: Social media users are using images and text to voice opinions and share ideas.
Approach: They propose to use user comments to extract hinting features from user comments and explore them via self-training.
Outcome: The proposed framework improves on four social media benchmarks for image-text relation classification, sarcasm detection, sentiment classification, and hate speech detection.
Integrating Text and Image: Determining Multimodal Document Intent in Instagram Posts (D19-1)

Copied to clipboard

Challenge: Existing studies on text-image content have focused on image as primary content, and text as secondary content.
Approach: They propose a multimodal dataset of 1299 Instagram posts labeled for three orthogonal taxonomies . they show that employing both text and image improves intent detection by 9.6 .
Outcome: The proposed model shows that using both text and image improves intent detection by 9.6 compared to using only the image modality.
Bridging Modality Gap for Effective Multimodal Sentiment Analysis in Fashion-related Social Media (2025.coling-main)

Copied to clipboard

Challenge: Existing sentiment analysis tasks focus on text comprehension, but visual content is important for emotional expression.
Approach: They propose a multimodal framework that integrates information from various modalities for sentiment classification of fashion posts.
Outcome: The proposed framework outperforms existing unimodal and multimodal baselines on a comprehensive dataset and significantly outperformed existing unilmodal and multiple modal frameworks.
Visual Attention Model for Name Tagging in Multimodal Social Media (P18-1)

Copied to clipboard

Challenge: Name tagging is a key task for language understanding, but is often limited by the short textual components.
Approach: They propose a novel model architecture based on visual attention that outperforms other methods . they use multimodal datasets to analyze the name tagging task on social media .
Outcome: The proposed model outperforms existing methods and significantly outperformed existing methods.
Different Data, Different Modalities! Reinforced Data Splitting for Effective Multimodal Information Extraction from Social Media Posts (2022.coling-1)

Copied to clipboard

Challenge: Recent multimodal information extraction approaches overestimate the significance of images.
Approach: They propose a general data splitting strategy to divide social media posts into two sets to achieve better performance under information extraction models of the corresponding modalities.
Outcome: The proposed method outperforms existing models on two different multimodal information extraction tasks.
Reasoning with Multimodal Sarcastic Tweets via Modeling Cross-Modality Contrast and Semantic Association (2020.acl-main)

Copied to clipboard

Challenge: Existing methods for sarcasm detection rely on text data, but are insufficient to detect multimodal sarcasm.
Approach: They propose a method for modeling cross-modality contrast in the associated context by constructing the Decomposition and Relation Network.
Outcome: The proposed model can detect sarcasm in multimodal tweets using a dataset .
Multimodal Emoji Prediction (N18-2)

Copied to clipboard

Challenge: Emojis are small images that are commonly included in social media text messages.
Approach: They propose a multimodal approach that is able to predict emojis in Instagram posts by using both text and image.
Outcome: The proposed model incorporates both text and image to improve accuracy .
Exploiting Pseudo Image Captions for Multimodal Summarization (2023.findings-acl)

Copied to clipboard

Challenge: Existing approaches to multimodal summarization with multimodal output (MSMO) lack reference images for training, and exposure of image captions during training is inconsistent with MSMO’s task settings.
Approach: They propose a coarse-to-fine image-text alignment mechanism to identify the most relevant sentence of each image in a document, resembling the role of image captions in capturing visual knowledge.
Outcome: The proposed method sets up state-of-the-art on all intermodality and intramodality metrics and improves on image recommendation precision.
Multimodal Named Entity Recognition for Short Social Media Posts (N18-1)

Copied to clipboard

Challenge: Social media posts often contain inconsistent or incomplete syntax and lexical notations with limited textual contexts.
Approach: They propose a task called Multimodal Named Entity Recognition (MNER) for noisy user-generated data . they use a dataset called SnapCaptions to build upon the state-of-the-art NER models .
Outcome: The proposed model outperforms existing models on noisy user-generated data . it uses a deep image network and generic modality attention module .
Natural Disaster Tweets Classification Using Multimodal Data (2023.emnlp-main)

Copied to clipboard

Challenge: Social media platforms are used for expressing opinions or conveying information.
Approach: They propose a hierarchical system that can integrate multimodal data and perform sequential hierarchic classification.
Outcome: The proposed system can find the damage and its severity along with classify the data into humanitarian categories.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations