Challenge: haptic captioning is the task of generating natural language descriptions from haptics, such as vibrations, for use in virtual reality and rehabilitation applications.
Approach: They propose a multimodal sensory language model that interprets vibration signals into descriptions in a given sensory, emotional, or associative category.
Outcome: The proposed model interprets vibration signals into descriptions in a given sensory, emotional, or associative category.

Similar Papers

HapticCap: A Multimodal Dataset and Task for Understanding User Experience of Vibration Haptic Signals (2025.findings-emnlp)

Copied to clipboard

Challenge: a dataset of vibration haptic signals is developed to match descriptions to vibrations . a lack of large datasets annotated with textual descriptions is a challenge .
Approach: They propose a multimodal dataset and task to match user descriptions to vibration haptic signals.
Outcome: The proposed dataset matches user descriptions to vibration haptic signals . the results show that language models and audio models perform better than existing models .
Exploring the Capability of Multimodal LLMs with Yonkoma Manga: The YManga Dataset and Its Challenging Tasks (2024.findings-emnlp)

Copied to clipboard

Challenge: YManga dataset is the first specifically designed for yonkoma manga understanding .
Approach: They propose to use a dataset of 1,015 yonkoma strips with 10,150 human annotations to define three tasks for panel sequence detection, intent generation and description generation for masked panels.
Outcome: The proposed dataset contains 1,015 high-quality yonkoma strips with 10,150 human annotations.
Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired Users (2025.acl-long)

Copied to clipboard

Challenge: Despite high adoption rate of Large Language Models, there are limitations related to contextual understanding, cultural sensitivity, and complex scene understanding.
Approach: They conduct a user survey to identify adoption patterns and key challenges users face with such technologies.
Outcome: The proposed models have high adoption rates but still face limitations in visual aids.
Unveiling Multimodal Processing: Exploring Activation Patterns in Multimodal LLMs for Interpretability and Efficiency (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in multimodal large language models have remained opaque.
Approach: They propose a method to convert dense MLLMs into fine-grained Mixture-of-Experts architectures.
Outcome: The proposed method outperforms random expert pruning and sparse activation and model pruning.
Retrieving Multimodal Information for Augmented Generation: A Survey (2023.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly using multimodality to augment their generation ability, but there is no unified perception of at which stage and how to incorporate different modalities.
Approach: They propose to use multimodality to augment Large Language Models (LLMs) this will provide scholars with a deeper understanding of the methods' applications and encourage them to adapt existing techniques to the fast-growing field of LLMs.
Outcome: The proposed methods improve factuality, reasoning, interpretability, and robustness of the generated content.
Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions (2025.naacl-long)

Copied to clipboard

Challenge: Recent work focuses on training vision-language models with long, detailed image captions, but small-scale VLMs struggle to balance the richness of these captions with the risk of hallucinations.
Approach: They propose an evaluation framework that breaks down generated captions into individual propositions, assessing each in isolation.
Outcome: The proposed framework outperforms baselines in both automatic metrics and human evaluations on small-scale vision-language models with long, detailed captions.
SignAlignLM: Integrating Multimodal Sign Language Processing into Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Deaf and Hard-of-Hearing (DHH) users increasingly utilize Large Language Models (LLMs), yet face significant challenges due to these models’ limited understanding of sign language grammar, multimodal sign inputs, and Deafic cultural contexts.
Approach: They propose to use sign language support in LLMs to integrate sign linguistic rules and conventions into prompting and fine-tuning strategies to address the needs of DHH users.
Outcome: The proposed model can be generalized interfaces for both spoken and signed languages if trained with a multitasking paradigm.
A Computational Acquisition Model for Multimodal Word Categorization (2022.naacl-main)

Copied to clipboard

Challenge: Recent advances in self-supervised modeling of text and images open new opportunities for computational models of child language acquisition.
Approach: They propose a multimodal language acquisition model trained from image-caption pairs on naturalistic data using cross-modal self-supervision.
Outcome: The proposed model learns word categories and object recognition abilities, the authors show . their model is trained from image-caption pairs on naturalistic data using cross-modal self-supervision .
Widget Captioning: Generating Natural Language Description for Mobile User Interface Elements (2020.emnlp-main)

Copied to clipboard

Challenge: Existing tools for examining and fixing missing captions are lacking in mobile UIs.
Approach: They propose a task for automatically generating language descriptions for UI elements from multimodal input including both the image and structural representations of user interfaces.
Outcome: The proposed task can generate captions from image and structural representations of UI elements.
VLIS: Unimodal Language Models Guide Multimodal Language Generation (2023.emnlp-main)

Copied to clipboard

Challenge: Existing vision-language models face challenges in tasks that require complex linguistic understanding.
Approach: They propose a framework that combines visual conditioning and linguistic understanding of unimodal text-only language models without further training to improve vision-language models.
Outcome: The proposed framework improves vision-language models on diverse tasks including commonsense understanding and complex text generation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations