HapticLLaMA: A Multimodal Sensory Language Model for Haptic Captioning (2026.findings-eacl)
Copied to clipboard
| Challenge: | haptic captioning is the task of generating natural language descriptions from haptics, such as vibrations, for use in virtual reality and rehabilitation applications. |
| Approach: | They propose a multimodal sensory language model that interprets vibration signals into descriptions in a given sensory, emotional, or associative category. |
| Outcome: | The proposed model interprets vibration signals into descriptions in a given sensory, emotional, or associative category. |
Similar Papers
HapticCap: A Multimodal Dataset and Task for Understanding User Experience of Vibration Haptic Signals (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a dataset of vibration haptic signals is developed to match descriptions to vibrations . a lack of large datasets annotated with textual descriptions is a challenge . |
| Approach: | They propose a multimodal dataset and task to match user descriptions to vibration haptic signals. |
| Outcome: | The proposed dataset matches user descriptions to vibration haptic signals . the results show that language models and audio models perform better than existing models . |
Exploring the Capability of Multimodal LLMs with Yonkoma Manga: The YManga Dataset and Its Challenging Tasks (2024.findings-emnlp)
Copied to clipboard
| Challenge: | YManga dataset is the first specifically designed for yonkoma manga understanding . |
| Approach: | They propose to use a dataset of 1,015 yonkoma strips with 10,150 human annotations to define three tasks for panel sequence detection, intent generation and description generation for masked panels. |
| Outcome: | The proposed dataset contains 1,015 high-quality yonkoma strips with 10,150 human annotations. |
Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired Users (2025.acl-long)
Copied to clipboard
Antonia Karamolegkou, Malvina Nikandrou, Georgios Pantazopoulos, Danae Sanchez Villegas, Phillip Rust, Ruchira Dhar, Daniel Hershcovich, Anders Søgaard
| Challenge: | Despite high adoption rate of Large Language Models, there are limitations related to contextual understanding, cultural sensitivity, and complex scene understanding. |
| Approach: | They conduct a user survey to identify adoption patterns and key challenges users face with such technologies. |
| Outcome: | The proposed models have high adoption rates but still face limitations in visual aids. |
Unveiling Multimodal Processing: Exploring Activation Patterns in Multimodal LLMs for Interpretability and Efficiency (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in multimodal large language models have remained opaque. |
| Approach: | They propose a method to convert dense MLLMs into fine-grained Mixture-of-Experts architectures. |
| Outcome: | The proposed method outperforms random expert pruning and sparse activation and model pruning. |
Retrieving Multimodal Information for Augmented Generation: A Survey (2023.findings-emnlp)
Copied to clipboard
Ruochen Zhao, Hailin Chen, Weishi Wang, Fangkai Jiao, Xuan Long Do, Chengwei Qin, Bosheng Ding, Xiaobao Guo, Minzhi Li, Xingxuan Li, Shafiq Joty
| Challenge: | Large Language Models (LLMs) are increasingly using multimodality to augment their generation ability, but there is no unified perception of at which stage and how to incorporate different modalities. |
| Approach: | They propose to use multimodality to augment Large Language Models (LLMs) this will provide scholars with a deeper understanding of the methods' applications and encourage them to adapt existing techniques to the fast-growing field of LLMs. |
| Outcome: | The proposed methods improve factuality, reasoning, interpretability, and robustness of the generated content. |
Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions (2025.naacl-long)
Copied to clipboard
| Challenge: | Recent work focuses on training vision-language models with long, detailed image captions, but small-scale VLMs struggle to balance the richness of these captions with the risk of hallucinations. |
| Approach: | They propose an evaluation framework that breaks down generated captions into individual propositions, assessing each in isolation. |
| Outcome: | The proposed framework outperforms baselines in both automatic metrics and human evaluations on small-scale vision-language models with long, detailed captions. |
SignAlignLM: Integrating Multimodal Sign Language Processing into Large Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Deaf and Hard-of-Hearing (DHH) users increasingly utilize Large Language Models (LLMs), yet face significant challenges due to these models’ limited understanding of sign language grammar, multimodal sign inputs, and Deafic cultural contexts. |
| Approach: | They propose to use sign language support in LLMs to integrate sign linguistic rules and conventions into prompting and fine-tuning strategies to address the needs of DHH users. |
| Outcome: | The proposed model can be generalized interfaces for both spoken and signed languages if trained with a multitasking paradigm. |
A Computational Acquisition Model for Multimodal Word Categorization (2022.naacl-main)
Copied to clipboard
| Challenge: | Recent advances in self-supervised modeling of text and images open new opportunities for computational models of child language acquisition. |
| Approach: | They propose a multimodal language acquisition model trained from image-caption pairs on naturalistic data using cross-modal self-supervision. |
| Outcome: | The proposed model learns word categories and object recognition abilities, the authors show . their model is trained from image-caption pairs on naturalistic data using cross-modal self-supervision . |
Widget Captioning: Generating Natural Language Description for Mobile User Interface Elements (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing tools for examining and fixing missing captions are lacking in mobile UIs. |
| Approach: | They propose a task for automatically generating language descriptions for UI elements from multimodal input including both the image and structural representations of user interfaces. |
| Outcome: | The proposed task can generate captions from image and structural representations of UI elements. |
VLIS: Unimodal Language Models Guide Multimodal Language Generation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing vision-language models face challenges in tasks that require complex linguistic understanding. |
| Approach: | They propose a framework that combines visual conditioning and linguistic understanding of unimodal text-only language models without further training to improve vision-language models. |
| Outcome: | The proposed framework improves vision-language models on diverse tasks including commonsense understanding and complex text generation. |