Challenge: Existing methods for audio-visual speech recognition use extra data to increase performance . a recent study shows that the use of unimodal self-supervised learning improves performance on multimodal tasks.
Approach: They propose to use unimodal self-supervised learning to train AVSR models on unlabelled unilateral data.
Outcome: The proposed model improves on lip reading sentences 2 by 30% even without an external language model.

Similar Papers

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks (2026.findings-acl)

Copied to clipboard

Challenge: Survey aims to identify challenges of multimodal unlearning for vision, language, audio and video . retraining after deletion requests or policy updates is often impractical, survey finds .
Approach: They propose to enable selective removal across modalities while retaining overall utility.
Outcome: This study compares models with existing models to identify weaknesses and improves performance.
CLASP: Cross-modal Alignment Using Pre-trained Unimodal Models (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in speech-text pretraining rely on parallel speech- text data . however, data accessibility is a challenge due to the limited data available.
Approach: They propose a framework for jointly performing speech and text processing without parallel corpora during pre-training but only downstream.
Outcome: The proposed framework extracts distinct representations for speech and text, aligning them effectively in a newly defined space using a multi-level contrastive learning mechanism.
Self-supervised Cross-modal Pretraining for Speech Emotion Recognition and Sentiment Analysis (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to multimodal speech emotion recognition and sentiment analysis have not improved results due to their relatively simple fusion mechanisms and lack of proper cross-modal pretraining.
Approach: They propose a deep-fused audio-text bi-modal transformer with carefully designed cross-modal fusion mechanism and stage-wise cross-mod pretraining scheme to facilitate cross-modulation.
Outcome: The proposed method exceeds benchmarks on public IEMOCAP emotion and CMU-MOSEI sentiment datasets by a large margin.
Multimodal and Multiresolution Speech Recognition with Transformers (2020.acl-main)

Copied to clipboard

Challenge: Existing audio visual automatic speech recognition systems rely on audio input to produce transcriptions.
Approach: They propose an audio visual automatic speech recognition system using a transformer-based architecture and incorporate a multitask training criterion for multiresolution ASR.
Outcome: The proposed system can generate character and subword transcriptions with visual information.
UniS-MMC: Multimodal Classification via Unimodality-supervised Multimodal Contrastive Learning (2023.findings-acl)

Copied to clipboard

Challenge: Existing multimodal fusion methods ignore inter-modality relationship, treat each modality equally, suffer sensor noise, and thus reduce multimodal learning performance.
Approach: They propose a multimodal contrastive method to explore more reliable multimodal representations under the weak supervision of unimodal predicting.
Outcome: The proposed method outperforms current state-of-the-art multimodal learning methods on image-text classification benchmarks UPMC-Food-101 and N24News.
MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens (2025.findings-acl)

Copied to clipboard

Challenge: Recent Large Language Model (LLM) based AVSR systems incur high computational costs due to high temporal resolution of audio-visual speech.
Approach: They propose an efficient multimodal speech LLM framework that minimizes token length while preserving essential linguistic content.
Outcome: The proposed approach reduces token usage by 86% while using only 3.5 tokens per second.
MIR-GAN: Refining Frame-Level Modality-Invariant Representations with Adversarial Network for Audio-Visual Speech Recognition (2023.acl-long)

Copied to clipboard

Challenge: Audio-visual speech recognition (AVSR) leverages multimodal signals to understand human speech.
Approach: They propose an adversarial network to refine frame-level modality-invariant representations to bridge the distribution gap between modalities.
Outcome: The proposed approach outperforms the state-of-the-art on public benchmarks LRS3 and LRS2 on the modalities of AVSR.
A Large-Scale Chinese Multimodal NER Dataset with Speech Clues (2021.acl-long)

Copied to clipboard

Challenge: Using a large-scale dataset, we explore Chinese named entity recognition (NER) with both textual and acoustic contents.
Approach: They propose a Chinese multimodal named entity recognition dataset . their corpus contains 42,987 annotated sentences and 71 hours of speech data .
Outcome: The proposed model yields state-of-the-art (SoTA) results on Chinese multimodal named entity recognition (NER) based on 42,987 annotated sentences and 71 hours of speech data.
Bridging the Temporal Gap in Multimodal LLMs: Deeply Stacking Temporal Tokens for Audio-Visual Speech Recognition (2026.findings-acl)

Copied to clipboard

Challenge: Existing audio-visual speech recognition systems suffer from a temporal gap . visual speech patterns captured from lip movements provide complementary information that remains inherently robust to acoustic noise.
Approach: They propose a framework that deeply stacks temporal tokens across both encoding and decoding stages to bridge this temporal gap.
Outcome: The proposed framework outperforms existing supervised, self-supervised, and LLM-based methods by 6.1% on LRS2 and 7.8% on LLS3.
Multi-Stage Multi-Modal Pre-Training for Automatic Speech Recognition (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for pre-training for automatic speech recognition (ASR) focus on single-stage pre-train followed by fine-tuning on downstream task.
Approach: They propose a multi-modal pre-training method that combines unsupervised pre-training with translation-based supervised mid-training.
Outcome: The proposed method improves WERs by 38.45% over baselines on both Librispeech and SUPERB.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations