Challenge: sarcasm is often expressed through multiple verbal and non-verbal cues, such as a change of tone, overemphasis, drawn-out syllables, or a straight looking face.
Approach: They propose to use multimodal cues to improve sarcasm detection using audiovisual utterances annotated with sarcasm labels to improve the accuracy.
Outcome: The proposed dataset reduces the error rate of sarcasm detection by 12.9% . it is based on audiovisual utterances annotated with sarcasm labels .

Similar Papers

A Multimodal Corpus for Emotion Recognition in Sarcasm (2022.lrec-1)

Copied to clipboard

Challenge: sarcasm and emotion are often used in conversational systems to generate the right response.
Approach: They use a sarcastic expression dataset pre-annotated with 9 emotions to detect emotion . they identify and correct 343 incorrect emotion labels and label each sarkastic utterance with one of four sarcasm types.
Outcome: The proposed model outperforms state-of-the-art sarcasm detection methods by using a multimodal sarcastic detection dataset.
SarcNet: A Multilingual Multimodal Sarcasm Detection Dataset (2024.lrec-main)

Copied to clipboard

Challenge: Sarcasm is an implicit form of sarcasm, involving an intended meaning that contradicts the literal expression . human use conflict between factual information and a statement as cues to detect sarcasm . sarkasmatic analysis is challenging due to its implicit nature .
Approach: They propose a multimodal sarcasm detection dataset that uses multiple modalities to detect sarcasm.
Outcome: The proposed model improves on previous models based on a single label . human sarcasm cannot be detected using a unified label across multiple modalities .
Predict and Use: Harnessing Predicted Gaze to Improve Multimodal Sarcasm Detection (2023.emnlp-main)

Copied to clipboard

Challenge: sarcasm detection depends on content spoken, tonality, facial expressions, context, and personal traits like language proficiency and cognitive capabilities.
Approach: They propose to use synthetic gaze data to improve sarcasm detection in conversational context . they collect gaze features for 20% of data instances and use them to predict gaze features .
Outcome: The proposed model improves performance on a conversational dataset using gaze features . it achieves a gain of 6.6% points on the complete dataset with only predicted gaze features.
Scale Is All You Need: Analyzing Modality Interaction and Speaker Intent Without Fine-Tuning (2026.eacl-srw)

Copied to clipboard

Challenge: Recent work on sarcasm and humor detection uses large multimodal Transformers, but they are computationally expensive and opaque.
Approach: They propose a lightweight framework for multimodal sarcasm detection that combines frozen text, audio, and visual embeddings from pretrained encoders through compact fusion heads.
Outcome: The proposed framework improves on the best unimodal baseline by combining text, audio, and visual embeddings from pretrained encoders with compact fusion heads.
Leveraging Generative Large Language Models with Visual Instruction and Demonstration Retrieval for Multimodal Sarcasm Detection (2024.naacl-long)

Copied to clipboard

Challenge: Existing methods for multimodal sarcasm detection do not fully utilize cross-modal features, limiting their performance on in-domain datasets.
Approach: They propose a multimodal sarcasm detection model with a designed instruction template and a demonstration retrieval module.
Outcome: The proposed model outperforms existing methods on in-domain datasets and achieves state-of-the-art performance.
Action and Reaction Go Hand in Hand! a Multi-modal Dialogue Act Aided Sarcasm Identification (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies have shown that sarcasm is reflected by the intended meaning of the speaker's utterance.
Approach: They propose to extend the MUStARD dataset to enclose dialogue acts for each dialogue . they propose a dialogue act-aided multi-modal transformer network for sarcasm identification model .
Outcome: The proposed model improves performance in dialogue act-aided sarcasm identification compared to sardasmatic identification alone.
Multi-View Incongruity Learning for Multimodal Sarcasm Detection (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for multimodal sarcasm detection rely on spurious correlations, demonstrating poor generalizability beyond training environments.
Approach: They propose a method that integrates multimodal incongruities via contrastive learning for multimodal sarcasm detection by using three views to drive multi-view learning.
Outcome: The proposed method outperforms existing methods on benchmark datasets and shows that it is more generalizable than existing methods.
When did you become so smart, oh wise one?! Sarcasm Explanation in Multi-modal Multi-party Dialogues (2022.acl-long)

Copied to clipboard

Challenge: Indirect speech achieves a constellation of discourse goals in human communication, but it is challenging for AI agents to comprehend such idiosyncrasies.
Approach: They propose a task to generate natural language explanations of satirical conversations using a multimodal and code-mixed dataset to capture multimodality.
Outcome: The proposed task generates natural language explanations of satirical conversations in a multimodal and code-mixed setting and surpasses baselines on almost all metrics.
Multi-Modal Sarcasm Detection via Cross-Modal Graph Convolutional Network (2022.acl-long)

Copied to clipboard

Challenge: Existing studies on multimodal sarcasm detection using textual and visual information have been limited to text-only approaches.
Approach: They propose to construct a cross-modal graph for each multi-modal instance to explicitly draw the ironic relations between textual and visual modalities.
Outcome: The proposed model achieves state-of-the-art in multi-modal sarcasm detection.
Generalizable Sarcasm Detection is Just Around the Corner, of Course! (2024.naacl-long)

Copied to clipboard

Challenge: sarcasm can be used to hurt, criticize, or deride but also to be mocking, humorous, or to bond.
Approach: They tested the robustness of sarcasm detection models by fine-tuning their behavior on four sarkasmatic datasets . they found that models performed better when fine- tuned with third-party labels than with author labels.
Outcome: The proposed models performed better when fine-tuned with third-party labels than with author labels on the same dataset and across different datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations