Papers by Yongseop Shin
Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval (2026.acl-long)
Copied to clipboard
| Challenge: | Experiments with AudioCaps, Clotho, and MECAT show that OEA achieves comparable text-to-text retrieval performance to state-of-the-art M2D-CLAP. |
| Approach: | They propose a retrieval-oriented encoder leveraging multimodal LLMs with native audio understanding that allows users to express their queries in five different ways. |
| Outcome: | Experiments on AudioCaps, Clotho, and MECAT show that OEA achieves comparable text-to-audio retrieval performance to state-of-the-art M2D-CLAP while demonstrating clear advantages in two critical areas. |