Santiago Castro, Devamanyu Hazarika, Verónica Pérez-Rosas, Roger Zimmermann, Rada Mihalcea, Soujanya Poria
| Challenge: | sarcasm is often expressed through multiple verbal and non-verbal cues, such as a change of tone, overemphasis, drawn-out syllables, or a straight looking face. |
| Approach: | They propose to use multimodal cues to improve sarcasm detection using audiovisual utterances annotated with sarcasm labels to improve the accuracy. |
| Outcome: | The proposed dataset reduces the error rate of sarcasm detection by 12.9% . it is based on audiovisual utterances annotated with sarcasm labels . |
Similar Papers
A Multimodal Corpus for Emotion Recognition in Sarcasm (2022.lrec-1)
Copied to clipboard
| Challenge: | sarcasm and emotion are often used in conversational systems to generate the right response. |
| Approach: | They use a sarcastic expression dataset pre-annotated with 9 emotions to detect emotion . they identify and correct 343 incorrect emotion labels and label each sarkastic utterance with one of four sarcasm types. |
| Outcome: | The proposed model outperforms state-of-the-art sarcasm detection methods by using a multimodal sarcastic detection dataset. |
SarcNet: A Multilingual Multimodal Sarcasm Detection Dataset (2024.lrec-main)
Copied to clipboard
| Challenge: | Sarcasm is an implicit form of sarcasm, involving an intended meaning that contradicts the literal expression . human use conflict between factual information and a statement as cues to detect sarcasm . sarkasmatic analysis is challenging due to its implicit nature . |
| Approach: | They propose a multimodal sarcasm detection dataset that uses multiple modalities to detect sarcasm. |
| Outcome: | The proposed model improves on previous models based on a single label . human sarcasm cannot be detected using a unified label across multiple modalities . |
Predict and Use: Harnessing Predicted Gaze to Improve Multimodal Sarcasm Detection (2023.emnlp-main)
Copied to clipboard
| Challenge: | sarcasm detection depends on content spoken, tonality, facial expressions, context, and personal traits like language proficiency and cognitive capabilities. |
| Approach: | They propose to use synthetic gaze data to improve sarcasm detection in conversational context . they collect gaze features for 20% of data instances and use them to predict gaze features . |
| Outcome: | The proposed model improves performance on a conversational dataset using gaze features . it achieves a gain of 6.6% points on the complete dataset with only predicted gaze features. |
Scale Is All You Need: Analyzing Modality Interaction and Speaker Intent Without Fine-Tuning (2026.eacl-srw)
Copied to clipboard
| Challenge: | Recent work on sarcasm and humor detection uses large multimodal Transformers, but they are computationally expensive and opaque. |
| Approach: | They propose a lightweight framework for multimodal sarcasm detection that combines frozen text, audio, and visual embeddings from pretrained encoders through compact fusion heads. |
| Outcome: | The proposed framework improves on the best unimodal baseline by combining text, audio, and visual embeddings from pretrained encoders with compact fusion heads. |
Leveraging Generative Large Language Models with Visual Instruction and Demonstration Retrieval for Multimodal Sarcasm Detection (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing methods for multimodal sarcasm detection do not fully utilize cross-modal features, limiting their performance on in-domain datasets. |
| Approach: | They propose a multimodal sarcasm detection model with a designed instruction template and a demonstration retrieval module. |
| Outcome: | The proposed model outperforms existing methods on in-domain datasets and achieves state-of-the-art performance. |
Action and Reaction Go Hand in Hand! a Multi-modal Dialogue Act Aided Sarcasm Identification (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing studies have shown that sarcasm is reflected by the intended meaning of the speaker's utterance. |
| Approach: | They propose to extend the MUStARD dataset to enclose dialogue acts for each dialogue . they propose a dialogue act-aided multi-modal transformer network for sarcasm identification model . |
| Outcome: | The proposed model improves performance in dialogue act-aided sarcasm identification compared to sardasmatic identification alone. |
Multi-View Incongruity Learning for Multimodal Sarcasm Detection (2025.coling-main)
Copied to clipboard
Diandian Guo, Cong Cao, Fangfang Yuan, Yanbing Liu, Guangjie Zeng, Xiaoyan Yu, Hao Peng, Philip S. Yu
| Challenge: | Existing methods for multimodal sarcasm detection rely on spurious correlations, demonstrating poor generalizability beyond training environments. |
| Approach: | They propose a method that integrates multimodal incongruities via contrastive learning for multimodal sarcasm detection by using three views to drive multi-view learning. |
| Outcome: | The proposed method outperforms existing methods on benchmark datasets and shows that it is more generalizable than existing methods. |
When did you become so smart, oh wise one?! Sarcasm Explanation in Multi-modal Multi-party Dialogues (2022.acl-long)
Copied to clipboard
| Challenge: | Indirect speech achieves a constellation of discourse goals in human communication, but it is challenging for AI agents to comprehend such idiosyncrasies. |
| Approach: | They propose a task to generate natural language explanations of satirical conversations using a multimodal and code-mixed dataset to capture multimodality. |
| Outcome: | The proposed task generates natural language explanations of satirical conversations in a multimodal and code-mixed setting and surpasses baselines on almost all metrics. |
Multi-Modal Sarcasm Detection via Cross-Modal Graph Convolutional Network (2022.acl-long)
Copied to clipboard
| Challenge: | Existing studies on multimodal sarcasm detection using textual and visual information have been limited to text-only approaches. |
| Approach: | They propose to construct a cross-modal graph for each multi-modal instance to explicitly draw the ironic relations between textual and visual modalities. |
| Outcome: | The proposed model achieves state-of-the-art in multi-modal sarcasm detection. |
Generalizable Sarcasm Detection is Just Around the Corner, of Course! (2024.naacl-long)
Copied to clipboard
| Challenge: | sarcasm can be used to hurt, criticize, or deride but also to be mocking, humorous, or to bond. |
| Approach: | They tested the robustness of sarcasm detection models by fine-tuning their behavior on four sarkasmatic datasets . they found that models performed better when fine- tuned with third-party labels than with author labels. |
| Outcome: | The proposed models performed better when fine-tuned with third-party labels than with author labels on the same dataset and across different datasets. |