| Challenge: | Sounding source localization is a challenging task due to the difficulty of cross-modal alignment. |
| Approach: | They propose an unsupervised method which enables pixel-level sounding source localization in unsupervised paradigm. |
| Outcome: | The proposed method achieves pixel-level sounding source localization without annotations. |
Similar Papers
Learning to See through Sound: From VggCaps to Multi2Cap for Richer Automated Audio Captioning (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing AAC datasets suffer from short and simplistic captions, limiting expressiveness and semantic depth. |
| Approach: | They propose a multi-modal dataset that pairs audio with corresponding video and leverages large language models to generate rich, descriptive captions. |
| Outcome: | The proposed framework outperforms existing benchmarks in caption length, lexical diversity, and human-rated quality. |
Unsupervised Multimodal Clustering for Semantics Discovery in Multimodal Utterances (2024.acl-long)
Copied to clipboard
| Challenge: | Existing methods for semantics discovery focus on text, video, and audio, failing to leverage the rich multimodal information in the real world. |
| Approach: | They propose a method to construct augmentation views for multimodal data and use them to perform pre-training to establish well-initialized representations for subsequent clustering. |
| Outcome: | The proposed method improves on benchmark multimodal intent and dialogue act datasets by 2-6% over state-of-the-art methods. |
The taste of IPA: Towards open-vocabulary keyword spotting and forced alignment in any language (2024.naacl-long)
Copied to clipboard
| Challenge: | a recent study shows that multilingual speech processing systems can generalize to unseen languages without adaptation. |
| Approach: | They propose a phoneme-based phoneme embedding model that can be generalized to unseen languages by using a neural forced aligner. |
| Outcome: | The proposed model can generalize to unseen languages without adaptation. |
FineLAP: Taming Heterogeneous Supervision for Fine-grained Language-Audio Pretraining (2026.acl-long)
Copied to clipboard
| Challenge: | Existing audio-language models excel at clip-level understanding but struggle with frame-level tasks. |
| Approach: | They propose a novel training paradigm that advances both clip- and frame-level alignment in CLAP with heterogeneous data. |
| Outcome: | The proposed training paradigm improves both clip- and frame-level alignment in CLAP with heterogeneous data. |
Beyond Transcription: Unified Audio Schema for Perception-Aware AudioLLMs (2026.findings-acl)
Copied to clipboard
Linhao Zhang, Yuhan Song, Aiwei Liu, Chuhan Wu, Sijun Zhang, Wei Jia, Yuan Liu, Houfeng Wang, Zhou Xiao
| Challenge: | Recent Audio Large Language Models (AudioLLMs) excel at reasoning tasks, but struggle at elementary auditory perception. |
| Approach: | They propose a framework that organizes audio information into three explicit components in a unified JSON format. |
| Outcome: | The proposed framework boosts fine-grained perception by 10.9% on MMSU over state-of-the-art models while preserving robust reasoning capabilities. |
Leveraging Unimodal Self-Supervised Learning for Multimodal Audio-Visual Speech Recognition (2022.acl-long)
Copied to clipboard
| Challenge: | Existing methods for audio-visual speech recognition use extra data to increase performance . a recent study shows that the use of unimodal self-supervised learning improves performance on multimodal tasks. |
| Approach: | They propose to use unimodal self-supervised learning to train AVSR models on unlabelled unilateral data. |
| Outcome: | The proposed model improves on lip reading sentences 2 by 30% even without an external language model. |
UniPSDA: Unsupervised Pseudo Semantic Data Augmentation for Zero-Shot Cross-Lingual Natural Language Understanding (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing studies rely on shallow unsupervised data generated by token surface matching regardless of global context-aware semantics of the surrounding text tokens. |
| Approach: | They propose an Unsupervised Pseudo Semantic Data Augmentation mechanism to enrich training data without human intervention. |
| Outcome: | The proposed model improves on general zero-shot cross-lingual understanding tasks on different languages without human intervention. |
Weakly-Supervised Spoken Video Grounding via Semantic Interaction Learning (2023.acl-long)
Copied to clipboard
| Challenge: | Recent work on spoken video grounding challenges extracting semantic information from speech . previous studies focused on textual queries, but recent work focuses on spoken queries . |
| Approach: | They propose a framework for weakly-supervised spoken video grounding to represent cross-modal semantics without expensive temporal annotations. |
| Outcome: | The proposed framework is more efficient than existing methods. |
Unsupervised Data Augmentation for Aspect Based Sentiment Analysis (2022.coling-1)
Copied to clipboard
| Challenge: | Recent approaches to Aspect-based Sentiment Analysis (ABSA) perform the subtasks of aspect term extraction (ATE) and aspect sentiment classification (ASC) simultaneously. |
| Approach: | They introduce an adaptation of Unsupervised Data Augmentation in semi-supervised learning that performs both aspects of Aspect-based Sentiment Analysis (ABSA) and aspect sentiment classification (ASC) they show that simple augmentations applied to modest-sized datasets along with consistency training lead to competitive performance with current ABSA state-of-the-art in restaurant and laptop domains . |
| Outcome: | The proposed approach performs well on a span-level classification task with minimal training data. |
Connecting the Dots between Audio and Text without Parallel Data through Visual Knowledge Transfer (2022.naacl-main)
Copied to clipboard
| Challenge: | Existing methods for learning audio-text connections rely on parallel audio- text data . a new approach allows for the representation of environmental soundscapes without using parallel data - a challenge for many applications . |
| Approach: | They propose a model that induces Audio-Text alignment without using parallel audio-text data. |
| Outcome: | The proposed model outperforms the current state-of-the-art for audio classification tasks with no audio-text data by 2.2% on the ESC50 and US8K tasks. |