Challenge: Recent years have witnessed significant advancements in large language models (LLMs) but still struggle with integrating vision and audio.
Approach: They propose a self-knowledge distillation method to improve vision-audio capabilities of OLLMs by learning from the vision-text components.
Outcome: The proposed method improves vision-audio capabilities of OLLMs by learning from vision-text components, which improves interaction between audio and images and results in improved performance on multimodal tasks.

Similar Papers

Leveraging Unimodal Self-Supervised Learning for Multimodal Audio-Visual Speech Recognition (2022.acl-long)

Copied to clipboard

Challenge: Existing methods for audio-visual speech recognition use extra data to increase performance . a recent study shows that the use of unimodal self-supervised learning improves performance on multimodal tasks.
Approach: They propose to use unimodal self-supervised learning to train AVSR models on unlabelled unilateral data.
Outcome: The proposed model improves on lip reading sentences 2 by 30% even without an external language model.
Probing Audio-Visual Reasoning in Multimodal Language Models through the Lens of Audio (2026.acl-long)

Copied to clipboard

Challenge: Recent multimodal large language models lack robust audio-visual integration ability and performance on DeafTest is highly correlated with AV-Odyssey accuracy.
Approach: They propose a benchmarking tool that integrates audio-visual reasoning with audio-video cues to infer solutions.
Outcome: The proposed model performs well on DeafTest, but lacks audio perception in simple audio tasks.
MASSV: Multimodal Adaptation and Self-Data Distillation for Speculative Decoding of Vision-Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Speculative decoding of vision-language models provides a novel way to accelerate language model inference by enabling a lightweight draft model to propose multiple tokens that a larger target model verifies simultaneously.
Approach: They propose a technique that allows a lightweight draft model to propose multiple tokens that a larger target model verifies simultaneously.
Outcome: The proposed technique increases accepted length by 30% and delivers speedups of up to 1.46x compared to conventional text-only drafting baselines on visually-grounded tasks.
MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct (2025.findings-acl)

Copied to clipboard

Challenge: a new framework for image-text instruction data evolution improves MLLM performance . lack of high-quality instruction data remains a major bottleneck in ML modeling .
Approach: They propose a multimodal instruction data evolution framework that iteratively enhances data quality through fine-grained perception, cognitive reasoning, and interaction evolution.
Outcome: The proposed approach improves MLLM performance in nine vision-language tasks while using significantly less data.
Self-Improvement in Multimodal Large Language Models: A Survey (2025.findings-emnlp)

Copied to clipboard

Challenge: Using data and data, self-improvement for Large Language Models has improved model capabilities without significantly increasing costs.
Approach: This survey provides a comprehensive overview of self-improvement for Large Language Models . it includes commonly used evaluations and downstream applications .
Outcome: The authors provide a comprehensive overview of self-improvement in Multimodal LLMs.
With Ears to See and Eyes to Hear: Sound Symbolism Experiments with Multimodal Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models and Vision Language Model (VLMs) have demonstrated aptitude as potential substitutes for human participants in psycholinguistic experiments.
Approach: They examine whether large language models and vision language models implicitly understand sound-based phenomena via orthography and imagery alone.
Outcome: The proposed models demonstrate sound symbolism and ability to "hear" using language and vision modules.
KD-VLP: Improving End-to-End Vision-and-Language Pretraining with Object Knowledge Distillation (2022.findings-naacl)

Copied to clipboard

Challenge: Existing vision-and-language pretraining approaches rely on external object detectors to encode images in a multi-modal transformer framework.
Approach: They propose an object-aware end-to-end VLP framework which feeds image grid features from CNNs into the Transformer and learns the multi-modal representations jointly.
Outcome: The proposed framework achieves competitive or superior performances on vision-language tasks.
ROSCO-Omni: Multimodal LLM-Based Communication Understanding for Non- and Minimally-Speaking Autistic Individuals (2026.findings-acl)

Copied to clipboard

Challenge: 30% of autistic individuals remain non- or minimally-speaking throughout their lives . however, caregivers rely on simultaneous integration of visual cues, auditory signals, and contextual understanding to infer intent.
Approach: They propose a framework that fine tunes a teacher-student MLLM for domain-specialized inference.
Outcome: The proposed framework achieves comparable performance to closed-source models .
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale (2025.acl-long)

Copied to clipboard

Challenge: Current instruction-tuning datasets focus on simplistic visual question answering tasks, and provide phrase-level answers without any intermediate rationales.
Approach: They propose to use open-source multimodal large language models to train MLLMs on a dataset with 12M instruction-response pairs to elicit CoT reasoning.
Outcome: The proposed model achieves state-of-the-art performance on benchmarks such as MathVerse, MMMU-Pro, and MuirBench, and gains improvements of up to 4% on non-reasoning-based benchmarks.
Enabling Multimodal Generation on CLIP via Vision-Language Knowledge Distillation (2022.findings-acl)

Copied to clipboard

Challenge: Recent large-scale vision-language pre-training models are powerful in multimodal classification and retrieval tasks.
Approach: They propose to augment a vision-language pre-training model with a textual pre-trained language model . the model achieves 44.5% zero-shot accuracy on multimodal generation tasks .
Outcome: The proposed model achieves 44.5% zero-shot accuracy on open-ended visual question answering and image captioning tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations