Challenge: Social media users are using images and text to voice opinions and share ideas.
Approach: They propose to use user comments to extract hinting features from user comments and explore them via self-training.
Outcome: The proposed framework improves on four social media benchmarks for image-text relation classification, sarcasm detection, sentiment classification, and hate speech detection.

Similar Papers

Improving Multimodal Classification of Social Media Posts by Leveraging Image-Text Auxiliary Tasks (2024.findings-eacl)

Copied to clipboard

Challenge: Prior work on multimodal content classification has not addressed these challenges.
Approach: They propose to use two auxiliary tasks to fine-tune multimodal models to address hidden cross-modal semantics and weak image-text relationships when modeling text and images.
Outcome: The proposed model improves by up to 2.6 F1 score across five diverse social media datasets.
Visual Attention Model for Name Tagging in Multimodal Social Media (P18-1)

Copied to clipboard

Challenge: Name tagging is a key task for language understanding, but is often limited by the short textual components.
Approach: They propose a novel model architecture based on visual attention that outperforms other methods . they use multimodal datasets to analyze the name tagging task on social media .
Outcome: The proposed model outperforms existing methods and significantly outperformed existing methods.
Exploring Unified Training Framework for Multimodal User Profiling (2025.coling-main)

Copied to clipboard

Challenge: Recent studies on user profiling focus on extracting multiple aspects of user attributes from textual reviews, but these studies do not fully exploit the potential of the rich multimodal data at hand.
Approach: They propose a task that utilizes both review texts and their accompanying images to generate comprehensive user profiles.
Outcome: The proposed training framework incorporates historical review texts and images for user profile generation.
SarcNet: A Multilingual Multimodal Sarcasm Detection Dataset (2024.lrec-main)

Copied to clipboard

Challenge: Sarcasm is an implicit form of sarcasm, involving an intended meaning that contradicts the literal expression . human use conflict between factual information and a statement as cues to detect sarcasm . sarkasmatic analysis is challenging due to its implicit nature .
Approach: They propose a multimodal sarcasm detection dataset that uses multiple modalities to detect sarcasm.
Outcome: The proposed model improves on previous models based on a single label . human sarcasm cannot be detected using a unified label across multiple modalities .
Leveraging Generative Large Language Models with Visual Instruction and Demonstration Retrieval for Multimodal Sarcasm Detection (2024.naacl-long)

Copied to clipboard

Challenge: Existing methods for multimodal sarcasm detection do not fully utilize cross-modal features, limiting their performance on in-domain datasets.
Approach: They propose a multimodal sarcasm detection model with a designed instruction template and a demonstration retrieval module.
Outcome: The proposed model outperforms existing methods on in-domain datasets and achieves state-of-the-art performance.
Multi-View Incongruity Learning for Multimodal Sarcasm Detection (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for multimodal sarcasm detection rely on spurious correlations, demonstrating poor generalizability beyond training environments.
Approach: They propose a method that integrates multimodal incongruities via contrastive learning for multimodal sarcasm detection by using three views to drive multi-view learning.
Outcome: The proposed method outperforms existing methods on benchmark datasets and shows that it is more generalizable than existing methods.
Different Data, Different Modalities! Reinforced Data Splitting for Effective Multimodal Information Extraction from Social Media Posts (2022.coling-1)

Copied to clipboard

Challenge: Recent multimodal information extraction approaches overestimate the significance of images.
Approach: They propose a general data splitting strategy to divide social media posts into two sets to achieve better performance under information extraction models of the corresponding modalities.
Outcome: The proposed method outperforms existing models on two different multimodal information extraction tasks.
Self-Supervised Multimodal Opinion Summarization (2021.acl-long)

Copied to clipboard

Challenge: Existing methods for opinion summarization use text data, but non-text data are less abundant.
Approach: They propose a self-supervised opinion summarization framework that uses non-text data to generate a summary from multiple reviews.
Outcome: The proposed framework is superior to existing methods on Yelp and Amazon datasets.
CofiPara: A Coarse-to-fine Paradigm for Multimodal Sarcasm Target Identification with Large Multimodal Models (2024.acl-long)

Copied to clipboard

Challenge: Current methods for multimodal sarcasm target identification focus on superficial indicators in an end-to-end manner, overlooking the nuanced understanding of multimodal content.
Approach: They propose a multimodal sarcasm target identification framework with a coarse-to-fine paradigm by augmenting sarcasm explainability with reasoning and pre-training knowledge.
Outcome: The proposed framework outperforms state-of-the-art methods and exhibits explainability in deciphering sarcasm as well.
MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media Texts (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for detecting social-media texts are limited to the English language and longer texts are not easily recognisable by humans.
Approach: They propose to use a multilingual and multi-platform dataset to compare machine-generated text detection methods in the social-media domain to compare them to human-written texts.
Outcome: The proposed dataset contains 472,097 texts, of which about 58k are human-written and approximately the same amount is generated by each of 7 multilingual LLMs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations