Papers with leave-one-speaker-out
Scale Is All You Need: Analyzing Modality Interaction and Speaker Intent Without Fine-Tuning (2026.eacl-srw)
Copied to clipboard
| Challenge: | Recent work on sarcasm and humor detection uses large multimodal Transformers, but they are computationally expensive and opaque. |
| Approach: | They propose a lightweight framework for multimodal sarcasm detection that combines frozen text, audio, and visual embeddings from pretrained encoders through compact fusion heads. |
| Outcome: | The proposed framework improves on the best unimodal baseline by combining text, audio, and visual embeddings from pretrained encoders with compact fusion heads. |