Challenge: Existing models that represent word meanings from word co-occurrences ignore associations between modalities and lack ability to transfer information between .
Approach: They propose a novel associative multichannel autoencoder that integrates textual, visual and auditory inputs to learn multimodal word representations.
Outcome: The proposed model outperforms strong unimodal models and state-of-the-art models on six benchmark concepts similarity tests.

Similar Papers

Retrieving Multimodal Information for Augmented Generation: A Survey (2023.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly using multimodality to augment their generation ability, but there is no unified perception of at which stage and how to incorporate different modalities.
Approach: They propose to use multimodality to augment Large Language Models (LLMs) this will provide scholars with a deeper understanding of the methods' applications and encourage them to adapt existing techniques to the fast-growing field of LLMs.
Outcome: The proposed methods improve factuality, reasoning, interpretability, and robustness of the generated content.
Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks (2026.findings-acl)

Copied to clipboard

Challenge: Survey aims to identify challenges of multimodal unlearning for vision, language, audio and video . retraining after deletion requests or policy updates is often impractical, survey finds .
Approach: They propose to enable selective removal across modalities while retaining overall utility.
Outcome: This study compares models with existing models to identify weaknesses and improves performance.
Word Representation Learning in Multimodal Pre-Trained Transformers: An Intrinsic Evaluation (2021.tacl-1)

Copied to clipboard

Challenge: Existing models for linguistic representations of words are based on information extracted from large text corpora, and the sensory-motor experiences humans have with the world play an important role in determining word meaning.
Approach: They propose to use contextualized word representations to learn semantic representations of words that align with human semantic intuitions.
Outcome: The proposed models are shown to be more efficient on concrete word pairs than on abstract ones.
How do Multimodal Foundation Models Encode Text and Speech? An Analysis of Cross-Lingual and Cross-Modal Representations (2025.naacl-short)

Copied to clipboard

Challenge: Recent advances in foundation models have sparked growing interest in expanding their text processing capabilities to speech.
Approach: They analyze the model activations from semantically equivalent sentences across languages in the text and speech modalities and examine how text and spoken are represented in recent multimodal foundation models.
Outcome: The proposed models exhibit cross-lingual differences, but are not explicitly trained for modality-agnostic representations.
Leveraging Unimodal Self-Supervised Learning for Multimodal Audio-Visual Speech Recognition (2022.acl-long)

Copied to clipboard

Challenge: Existing methods for audio-visual speech recognition use extra data to increase performance . a recent study shows that the use of unimodal self-supervised learning improves performance on multimodal tasks.
Approach: They propose to use unimodal self-supervised learning to train AVSR models on unlabelled unilateral data.
Outcome: The proposed model improves on lip reading sentences 2 by 30% even without an external language model.
Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis (2023.emnlp-main)

Copied to clipboard

Challenge: Multimodal Sentiment Analysis (MSA) is effective when using rich information from multiple sources, but the potential sentiment-irrelevant information across modalities may hinder the performance from being further improved.
Approach: They propose an Adaptive Language-guided Multimodal Transformer (ALMT) that learns an irrelevance/conflict-suppressing representation from visual and audio features under guidance of language features at different scales.
Outcome: The proposed model achieves state-of-the-art on several popular datasets and an abundance of ablation shows the effectiveness of the proposed model.
Multimodal Contrastive Learning via Uni-Modal Coding and Cross-Modal Prediction for Multimodal Sentiment Analysis (2022.findings-emnlp)

Copied to clipboard

Challenge: Recent work on multimodal representation learning has focused on uni-modality pre-training or cross-modalities integration.
Approach: They propose a framework for multimodal representation learning that uses uni-modal contrastive coding and an efficient unimodal feature augmentation strategy to capture intermodal dynamics.
Outcome: The proposed framework surpasses state-of-the-art methods on two public datasets.
MCSE: Multimodal Contrastive Learning of Sentence Embeddings (2022.naacl-main)

Copied to clipboard

Challenge: Existing approaches to learning semantically meaningful sentence embeddings are limited by the complexity of pre-trained models.
Approach: They propose a sentence embedding learning approach that exploits both visual and textual information via a multimodal contrastive objective.
Outcome: The proposed approach improves the state-of-the-art average Spearman’s correlation by 1.7% on a variety of semantic textual similarity tasks.
Unsupervised Natural Language Inference via Decoupled Multimodal Contrastive Learning (2020.emnlp-main)

Copied to clipboard

Challenge: a recent study shows that humans are not supervised by the natural language inference .
Approach: They propose to solve the natural language inference problem via task-agnostic multimodal pretraining.
Outcome: The proposed network outperforms fully-supervised BiLSTM and BiLS+ELMO on plain text inference datasets.
Tutorial on Multimodal Machine Learning (2022.naacl-tutorials)

Copied to clipboard

Challenge: Multimodal machine learning is a challenging but crucial area with numerous applications in multimedia, affective computing, robotics, finance, HCI, and healthcare.
Approach: This tutorial will describe an updated taxonomy on multimodal machine learning synthesizing its core technical challenges and major directions for future research.
Outcome: The proposed taxonomy synthesizes the core technical challenges and major directions for future research.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations