Challenge: a dataset of vibration haptic signals is developed to match descriptions to vibrations . a lack of large datasets annotated with textual descriptions is a challenge .
Approach: They propose a multimodal dataset and task to match user descriptions to vibration haptic signals.
Outcome: The proposed dataset matches user descriptions to vibration haptic signals . the results show that language models and audio models perform better than existing models .

Similar Papers

HapticLLaMA: A Multimodal Sensory Language Model for Haptic Captioning (2026.findings-eacl)

Copied to clipboard

Challenge: haptic captioning is the task of generating natural language descriptions from haptics, such as vibrations, for use in virtual reality and rehabilitation applications.
Approach: They propose a multimodal sensory language model that interprets vibration signals into descriptions in a given sensory, emotional, or associative category.
Outcome: The proposed model interprets vibration signals into descriptions in a given sensory, emotional, or associative category.
Multimodality for NLP-Centered Applications: Resources, Advances and Frontiers (2022.lrec-1)

Copied to clipboard

Challenge: resurgence of multimodal datasets has attracted significant research interest, but there is no comprehensive survey for this task.
Approach: They present a survey of a multimodal dataset with different modalities according to the applications.
Outcome: The proposed datasets are available online and discuss the new frontier and motivate future researches.
MemeCap: A Dataset for Captioning and Interpreting Memes (2023.emnlp-main)

Copied to clipboard

Challenge: a new dataset aims to understand meme captioning tasks using visual metaphors . vision and language models are proving to be effective in image captioning and visual question answering tasks .
Approach: They present a dataset that contains 6.3K memes and 6.3k meme captions . they show that vision and language models still struggle with visual metaphors despite their advanced capabilities .
Outcome: The proposed dataset contains 6.3K memes along with the title of the post containing the meme, meme captions, literal image caption, and visual metaphors.
EmpathicStories++: A Multimodal Dataset for Empathy Towards Personal Experiences (2024.findings-acl)

Copied to clipboard

Challenge: Existing datasets for empathy modeling are limited in the ways they are not captured in the wild.
Approach: They propose a multimodal dataset for empathy during personal experience sharing that contains 53 hours of video, audio, and text data of 41 participants.
Outcome: The EmpathicStories++ dataset contains 53 hours of video, audio, and text data of 41 participants sharing vulnerable experiences and reading empathically resonant stories with an AI agent.
Emotags: Computer-Assisted Verbal Labelling of Expressive Audiovisual Utterances for Expressive Multimodal TTS (2024.lrec-main)

Copied to clipboard

Challenge: We show that ascribing verbal descriptions to expressive audiovisual utterances is efficient and efficient.
Approach: They propose a web app for ascribing verbal descriptions to expressive audiovisual utterances.
Outcome: The proposed system can be deployed at a large scale to efficiently collect relevant verbal descriptions.
Development of an Annotated Multimodal Dataset for the Investigation of Classification and Summarisation of Presentations using High-Level Paralinguistic Features (L18-1)

Copied to clipboard

Challenge: Existing summarisation methods take no account of multimodal high-level paralinguistic features which form part of audio-visual presentations.
Approach: They propose to use audiovisual recordings to extract paralinguistic features from audio recordings . they use manual annotations to help users find relevant material .
Outcome: The proposed method can identify the most important or emphasised material within a presentation.
CHEER-Ekman: Fine-grained Embodied Emotion Classification (2025.acl-short)

Copied to clipboard

Challenge: Emotions manifest through physical experiences and bodily reactions, yet identifying such embodied emotions in text remains understudied.
Approach: They propose to extend existing binary embodied emotion dataset with Ekman’s six basic emotion categories.
Outcome: The proposed dataset outperforms existing methods with large language models.
ALCAP: Alignment-Augmented Music Captioner (2023.emnlp-main)

Copied to clipboard

Challenge: Traditional approaches to music captioning ignore the intricate interplay between the two . however, a comprehensive understanding of music necessitates the integration of both these elements.
Approach: They propose a method to learn multimodal alignment between audio and lyrics through contrastive learning.
Outcome: The proposed method achieves new state-of-the-art on two music captioning datasets.
A Survey of Data Augmentation Approaches for NLP (2021.findings-acl)

Copied to clipboard

Challenge: Data augmentation is a field of research that has been underexplored due to the discrete nature of language data.
Approach: They present a comprehensive survey of data augmentation for NLP by summarizing the literature in a structured manner.
Outcome: The proposed methods are used for popular NLP applications and tasks and highlight current challenges and directions for future research.
Chinese Synesthesia Detection: New Dataset and Models (2022.findings-acl)

Copied to clipboard

Challenge: Synesthesia refers to the description of perceptions in one sensory modality through concepts from other modalities.
Approach: They propose a task called synesthesia detection to extract the sensory word of a sentence and predict the original and synesthetic sensory modalities of the corresponding sensory word.
Outcome: The proposed model achieves state-of-the-art on the Chinese synesthesia dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations