Challenge: Mainstream speaker diarization systems rely only on acoustic information, making it challenging in complex aural environments.
Approach: They propose a multimodal approach that integrates audio, visual, and semantic cues to enhance speaker diarization.
Outcome: The proposed approach outperforms state-of-the-art methods on multi-party conversations . it integrates audio-visual-semantic cues into the clustering process for acoustic speaker embeddings .

Similar Papers

Exploring Speaker-Related Information in Spoken Language Understanding for Better Speaker Diarization (2023.findings-acl)

Copied to clipboard

Challenge: Current speaker diarization systems consider only acoustic information, resulting in performance degradation when encountering adverse acustic environment.
Approach: They propose methods to extract speaker-related information from conversational semantics in multi-party meetings.
Outcome: The proposed method improves on AISHELL-4 and AliMeeting datasets on speakers diarization and speaker-turn detection.
Speaker Clustering in Textual Dialogue with Pairwise Utterance Relation and Cross-corpus Dialogue Act Supervision (2022.coling-1)

Copied to clipboard

Challenge: Existing models for textual dialogues do not include speaker annotations.
Approach: They propose a speaker clustering model for textual dialogues that groups utterances without annotations so that the actual speakers are identical inside each cluster.
Outcome: The proposed model outperforms the sequence classification baseline and benefits from the auxiliary dialogue act classification task.
A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Existing benchmarks for testing audio modality of multimodal large language models focus on testing audio tasks in isolation.
Approach: They propose a new benchmark to assess multimodal large language models' ability to combine audio tasks.
Outcome: The proposed benchmarks show that multimodal models can solve problems that require reasoning over audio signals with satisfactory results.
Listening Like Humans: Semantics-Guided Noise-Robust Multimodal Speech Recognition (2026.acl-long)

Copied to clipboard

Challenge: Severe acoustic degradation results in unreliable ASR outputs . et al., 2024b): critical concerns regarding reliability and fairness of ASR .
Approach: They propose a multimodal framework that reframes ASR as semantics-guided speech reconstruction.
Outcome: The proposed framework achieves an average reduction in WER while also attaining 98.71% BERTScore and 96.7% USE over advanced baselines.
Speaker Overlap-aware Neural Diarization for Multi-party Meeting Analysis (2022.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to speaker diarization treat speaker dependency and overlaps as multi-label classification problems.
Approach: They propose to reformulate overlapped speaker diarization task as a single-label prediction problem via power set encoding (PSE) to overcome the disadvantages, they propose a speaker overlap-aware neural diarisation model which incorporates a context-independent scorer and a contextual-dependent score.
Outcome: The proposed model outperforms the state-of-the-art methods on speaker voice activity detection and improves relative diarization error reduction by 6.30%.
Speaker-Aware Discourse Parsing on Multi-Party Dialogues (2022.coling-1)

Copied to clipboard

Challenge: Discourse parsing on multi-party dialogues is an important but difficult task in dialogue systems and conversational analysis.
Approach: They propose a speaker-aware model for parsing on multi-party dialogues using interaction features between different speakers.
Outcome: The proposed model achieves the best-reported performance on two standard benchmark datasets.
TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding (2026.acl-long)

Copied to clipboard

Challenge: a critical ambiguity persists regarding what constitutes "joint ASR and diarization" a unified framework for multi-speaker ASR is proposed, but it is not yet clear what constitute "diarization."
Approach: They propose a unified LLM-based framework that uses Temporal Anchor Grounding for joint multi-speaker ASR and diarization.
Outcome: The proposed framework improves on AMI and AliMeeting benchmarks on speaker-content alignment . the proposed framework achieves consistent improvements in Diarization Error Rate over strong baselines .
AV-Dialog: Spoken Dialogue Models with Audio-Visual Input (2026.acl-long)

Copied to clipboard

Challenge: AV-Dialog uses audio and visual cues to track the target speaker, predict turn-taking, and generate coherent responses.
Approach: They propose a multimodal dialog framework that uses both audio and visual cues to track the target speaker.
Outcome: AV-Dialog outperforms audio-only models under interference, reducing transcription errors, improving turn-taking prediction and human-rated dialogue quality.
Leveraging Unimodal Self-Supervised Learning for Multimodal Audio-Visual Speech Recognition (2022.acl-long)

Copied to clipboard

Challenge: Existing methods for audio-visual speech recognition use extra data to increase performance . a recent study shows that the use of unimodal self-supervised learning improves performance on multimodal tasks.
Approach: They propose to use unimodal self-supervised learning to train AVSR models on unlabelled unilateral data.
Outcome: The proposed model improves on lip reading sentences 2 by 30% even without an external language model.
Bazinga! A Dataset for Multi-Party Dialogues Structuring (2022.lrec-1)

Copied to clipboard

Challenge: a dataset of 16 TV and movie series is filled with challenging multi-party dialogues.
Approach: They propose a dataset built around 16 TV and movie series with challenging multi-party dialogues.
Outcome: The proposed dataset is a step towards better multi-party dialogue structuring and understanding.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations