Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization on Multi-party Conversation (2025.acl-long)
Copied to clipboard
Luyao Cheng, Hui Wang, Chong Deng, Siqi Zheng, Yafeng Chen, Rongjie Huang, Qinglin Zhang, Qian Chen, Xihao Li, Wen Wang
| Challenge: | Mainstream speaker diarization systems rely only on acoustic information, making it challenging in complex aural environments. |
| Approach: | They propose a multimodal approach that integrates audio, visual, and semantic cues to enhance speaker diarization. |
| Outcome: | The proposed approach outperforms state-of-the-art methods on multi-party conversations . it integrates audio-visual-semantic cues into the clustering process for acoustic speaker embeddings . |
Similar Papers
Exploring Speaker-Related Information in Spoken Language Understanding for Better Speaker Diarization (2023.findings-acl)
Copied to clipboard
| Challenge: | Current speaker diarization systems consider only acoustic information, resulting in performance degradation when encountering adverse acustic environment. |
| Approach: | They propose methods to extract speaker-related information from conversational semantics in multi-party meetings. |
| Outcome: | The proposed method improves on AISHELL-4 and AliMeeting datasets on speakers diarization and speaker-turn detection. |
Speaker Clustering in Textual Dialogue with Pairwise Utterance Relation and Cross-corpus Dialogue Act Supervision (2022.coling-1)
Copied to clipboard
| Challenge: | Existing models for textual dialogues do not include speaker annotations. |
| Approach: | They propose a speaker clustering model for textual dialogues that groups utterances without annotations so that the actual speakers are identical inside each cluster. |
| Outcome: | The proposed model outperforms the sequence classification baseline and benefits from the auxiliary dialogue act classification task. |
A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models (2026.eacl-long)
Copied to clipboard
Iwona Christop, Mateusz Czyżnikiewicz, Paweł Skórzewski, Łukasz Bondaruk, Jakub Kubiak, Marcin Lewandowski, Marek Kubis
| Challenge: | Existing benchmarks for testing audio modality of multimodal large language models focus on testing audio tasks in isolation. |
| Approach: | They propose a new benchmark to assess multimodal large language models' ability to combine audio tasks. |
| Outcome: | The proposed benchmarks show that multimodal models can solve problems that require reasoning over audio signals with satisfactory results. |
Listening Like Humans: Semantics-Guided Noise-Robust Multimodal Speech Recognition (2026.acl-long)
Copied to clipboard
| Challenge: | Severe acoustic degradation results in unreliable ASR outputs . et al., 2024b): critical concerns regarding reliability and fairness of ASR . |
| Approach: | They propose a multimodal framework that reframes ASR as semantics-guided speech reconstruction. |
| Outcome: | The proposed framework achieves an average reduction in WER while also attaining 98.71% BERTScore and 96.7% USE over advanced baselines. |
Speaker Overlap-aware Neural Diarization for Multi-party Meeting Analysis (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to speaker diarization treat speaker dependency and overlaps as multi-label classification problems. |
| Approach: | They propose to reformulate overlapped speaker diarization task as a single-label prediction problem via power set encoding (PSE) to overcome the disadvantages, they propose a speaker overlap-aware neural diarisation model which incorporates a context-independent scorer and a contextual-dependent score. |
| Outcome: | The proposed model outperforms the state-of-the-art methods on speaker voice activity detection and improves relative diarization error reduction by 6.30%. |
Speaker-Aware Discourse Parsing on Multi-Party Dialogues (2022.coling-1)
Copied to clipboard
| Challenge: | Discourse parsing on multi-party dialogues is an important but difficult task in dialogue systems and conversational analysis. |
| Approach: | They propose a speaker-aware model for parsing on multi-party dialogues using interaction features between different speakers. |
| Outcome: | The proposed model achieves the best-reported performance on two standard benchmark datasets. |
TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding (2026.acl-long)
Copied to clipboard
| Challenge: | a critical ambiguity persists regarding what constitutes "joint ASR and diarization" a unified framework for multi-speaker ASR is proposed, but it is not yet clear what constitute "diarization." |
| Approach: | They propose a unified LLM-based framework that uses Temporal Anchor Grounding for joint multi-speaker ASR and diarization. |
| Outcome: | The proposed framework improves on AMI and AliMeeting benchmarks on speaker-content alignment . the proposed framework achieves consistent improvements in Diarization Error Rate over strong baselines . |
AV-Dialog: Spoken Dialogue Models with Audio-Visual Input (2026.acl-long)
Copied to clipboard
| Challenge: | AV-Dialog uses audio and visual cues to track the target speaker, predict turn-taking, and generate coherent responses. |
| Approach: | They propose a multimodal dialog framework that uses both audio and visual cues to track the target speaker. |
| Outcome: | AV-Dialog outperforms audio-only models under interference, reducing transcription errors, improving turn-taking prediction and human-rated dialogue quality. |
Leveraging Unimodal Self-Supervised Learning for Multimodal Audio-Visual Speech Recognition (2022.acl-long)
Copied to clipboard
| Challenge: | Existing methods for audio-visual speech recognition use extra data to increase performance . a recent study shows that the use of unimodal self-supervised learning improves performance on multimodal tasks. |
| Approach: | They propose to use unimodal self-supervised learning to train AVSR models on unlabelled unilateral data. |
| Outcome: | The proposed model improves on lip reading sentences 2 by 30% even without an external language model. |
Bazinga! A Dataset for Multi-Party Dialogues Structuring (2022.lrec-1)
Copied to clipboard
Paul Lerner, Juliette Bergoënd, Camille Guinaudeau, Hervé Bredin, Benjamin Maurice, Sharleyne Lefevre, Martin Bouteiller, Aman Berhe, Léo Galmant, Ruiqing Yin, Claude Barras
| Challenge: | a dataset of 16 TV and movie series is filled with challenging multi-party dialogues. |
| Approach: | They propose a dataset built around 16 TV and movie series with challenging multi-party dialogues. |
| Outcome: | The proposed dataset is a step towards better multi-party dialogue structuring and understanding. |