Papers by Biniyam Tolera
By My Eyes: Grounding Multimodal Large Language Models with Sensor Data via Visual Prompting (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing text-based prompts for large language models (LLMs) show performance degradation when handling long sensor data sequences. |
| Approach: | They propose a visual prompt that directs MLLMs to utilize visualized sensor data alongside descriptions of the target sensory task. |
| Outcome: | The proposed approach achieves 10% higher accuracy and reduces token costs by 15.8 times on nine sensory tasks involving four sensing modalities . |