HapticCap: A Multimodal Dataset and Task for Understanding User Experience of Vibration Haptic Signals (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a dataset of vibration haptic signals is developed to match descriptions to vibrations . a lack of large datasets annotated with textual descriptions is a challenge . |
| Approach: | They propose a multimodal dataset and task to match user descriptions to vibration haptic signals. |
| Outcome: | The proposed dataset matches user descriptions to vibration haptic signals . the results show that language models and audio models perform better than existing models . |
Similar Papers
HapticLLaMA: A Multimodal Sensory Language Model for Haptic Captioning (2026.findings-eacl)
Copied to clipboard
| Challenge: | haptic captioning is the task of generating natural language descriptions from haptics, such as vibrations, for use in virtual reality and rehabilitation applications. |
| Approach: | They propose a multimodal sensory language model that interprets vibration signals into descriptions in a given sensory, emotional, or associative category. |
| Outcome: | The proposed model interprets vibration signals into descriptions in a given sensory, emotional, or associative category. |
Multimodality for NLP-Centered Applications: Resources, Advances and Frontiers (2022.lrec-1)
Copied to clipboard
| Challenge: | resurgence of multimodal datasets has attracted significant research interest, but there is no comprehensive survey for this task. |
| Approach: | They present a survey of a multimodal dataset with different modalities according to the applications. |
| Outcome: | The proposed datasets are available online and discuss the new frontier and motivate future researches. |
MemeCap: A Dataset for Captioning and Interpreting Memes (2023.emnlp-main)
Copied to clipboard
| Challenge: | a new dataset aims to understand meme captioning tasks using visual metaphors . vision and language models are proving to be effective in image captioning and visual question answering tasks . |
| Approach: | They present a dataset that contains 6.3K memes and 6.3k meme captions . they show that vision and language models still struggle with visual metaphors despite their advanced capabilities . |
| Outcome: | The proposed dataset contains 6.3K memes along with the title of the post containing the meme, meme captions, literal image caption, and visual metaphors. |
EmpathicStories++: A Multimodal Dataset for Empathy Towards Personal Experiences (2024.findings-acl)
Copied to clipboard
Jocelyn Shen, Yubin Kim, Mohit Hulse, Wazeer Zulfikar, Sharifa Alghowinem, Cynthia Breazeal, Hae Park
| Challenge: | Existing datasets for empathy modeling are limited in the ways they are not captured in the wild. |
| Approach: | They propose a multimodal dataset for empathy during personal experience sharing that contains 53 hours of video, audio, and text data of 41 participants. |
| Outcome: | The EmpathicStories++ dataset contains 53 hours of video, audio, and text data of 41 participants sharing vulnerable experiences and reading empathically resonant stories with an AI agent. |
Emotags: Computer-Assisted Verbal Labelling of Expressive Audiovisual Utterances for Expressive Multimodal TTS (2024.lrec-main)
Copied to clipboard
| Challenge: | We show that ascribing verbal descriptions to expressive audiovisual utterances is efficient and efficient. |
| Approach: | They propose a web app for ascribing verbal descriptions to expressive audiovisual utterances. |
| Outcome: | The proposed system can be deployed at a large scale to efficiently collect relevant verbal descriptions. |
Development of an Annotated Multimodal Dataset for the Investigation of Classification and Summarisation of Presentations using High-Level Paralinguistic Features (L18-1)
Copied to clipboard
| Challenge: | Existing summarisation methods take no account of multimodal high-level paralinguistic features which form part of audio-visual presentations. |
| Approach: | They propose to use audiovisual recordings to extract paralinguistic features from audio recordings . they use manual annotations to help users find relevant material . |
| Outcome: | The proposed method can identify the most important or emphasised material within a presentation. |
CHEER-Ekman: Fine-grained Embodied Emotion Classification (2025.acl-short)
Copied to clipboard
| Challenge: | Emotions manifest through physical experiences and bodily reactions, yet identifying such embodied emotions in text remains understudied. |
| Approach: | They propose to extend existing binary embodied emotion dataset with Ekman’s six basic emotion categories. |
| Outcome: | The proposed dataset outperforms existing methods with large language models. |
ALCAP: Alignment-Augmented Music Captioner (2023.emnlp-main)
Copied to clipboard
| Challenge: | Traditional approaches to music captioning ignore the intricate interplay between the two . however, a comprehensive understanding of music necessitates the integration of both these elements. |
| Approach: | They propose a method to learn multimodal alignment between audio and lyrics through contrastive learning. |
| Outcome: | The proposed method achieves new state-of-the-art on two music captioning datasets. |
A Survey of Data Augmentation Approaches for NLP (2021.findings-acl)
Copied to clipboard
Steven Y. Feng, Varun Gangal, Jason Wei, Sarath Chandar, Soroush Vosoughi, Teruko Mitamura, Eduard Hovy
| Challenge: | Data augmentation is a field of research that has been underexplored due to the discrete nature of language data. |
| Approach: | They present a comprehensive survey of data augmentation for NLP by summarizing the literature in a structured manner. |
| Outcome: | The proposed methods are used for popular NLP applications and tasks and highlight current challenges and directions for future research. |
Chinese Synesthesia Detection: New Dataset and Models (2022.findings-acl)
Copied to clipboard
| Challenge: | Synesthesia refers to the description of perceptions in one sensory modality through concepts from other modalities. |
| Approach: | They propose a task called synesthesia detection to extract the sensory word of a sentence and predict the original and synesthetic sensory modalities of the corresponding sensory word. |
| Outcome: | The proposed model achieves state-of-the-art on the Chinese synesthesia dataset. |