Papers by Haifeng Hu

9 papers
CORD: Bridging the Audio–Text Reasoning Gap via Weighted On-policy Cross-modal Distillation (2026.findings-acl)

Copied to clipboard

Challenge: Large Audio Language Models (LALMs) exhibit a degradation in knowledge and reasoning capabilities . empirical results show that CORD significantly bridges the audio–text performance gap .
Approach: They propose a framework that performs online cross-modal self-distillation to bridge the acoustic-semantic gap between LALMs and text-based models.
Outcome: The proposed framework bridges the acoustic-semantic gap between LALMs and text-based models . it employs on-policy reverse KL divergence with importance-aware weighting .
Learning to Contrast the Counterfactual Samples for Robust Visual Question Answering (2020.emnlp-main)

Copied to clipboard

Challenge: Existing methods of generating counterfactual samples are not fully utilized in the task of Visual Question Answering (VQA).
Approach: They propose a self-supervised contrastive learning mechanism to learn the relationship between original samples, factual samples and counterfactual samples.
Outcome: The proposed method surpasses state-of-the-art models on the VQA-CP dataset, a diagnostic benchmark for assessing the VQ model’s robustness.
Which is Making the Contribution: Modulating Unimodal and Cross-modal Dynamics for Multimodal Sentiment Analysis (2021.findings-emnlp)

Copied to clipboard

Challenge: Recent studies focus on learning cross-modal dynamics, but neglect to explore optimal solution for unimodal networks.
Approach: They propose a new MSA framework to identify contribution of modalities and reduce impact of noisy information.
Outcome: The proposed model outperforms state-of-the-art methods on publicly available datasets.
Curriculum Learning Meets Weakly Supervised Multimodal Correlation Learning (2022.emnlp-main)

Copied to clipboard

Challenge: Existing studies have used the correlation information stored in samples for self-supervised learning, but they feed the training pairs in a random order without consideration of difficulty.
Approach: They propose to inject curriculum learning into weakly supervised multimodal correlation learning by scoring and feeding pairs according to difficulty.
Outcome: The proposed model achieves state-of-the-art on multimodal sentiment analysis without human annotation.
MoE Adapter for Large Audio Language Models: Sparsity, Disentanglement, and Gradient-Conflict-Free (2026.findings-acl)

Copied to clipboard

Challenge: Existing research on Large Language Models (LLMs) limited to textual input modality . acoustic information is intrinsically heterogeneous, entangling attributes such as speech, music, and environmental context.
Approach: They propose a sparse Mixture-of-Experts architecture to decouple acoustic information by routing audio tokens to specialized experts.
Outcome: The proposed architecture outperforms existing models on audio semantic and paralinguistic tasks while retaining shared experts for global context.
Divide, Conquer and Combine: Hierarchical Feature Fusion Network with Local and Global Perspectives for Multimodal Affective Computing (P19-1)

Copied to clipboard

Challenge: Existing approaches to multimodal fusion are based on fusing features at holistic level instead of focusing on local and global interactions.
Approach: They propose a general strategy called ‘divide, conquer and combine’ for multimodal fusion that combines local and global interactions in a hierarchy.
Outcome: The proposed strategy achieves state-of-the-art performance on multimodal affective computing with higher efficiency.
Evaluating the Expressive Appropriateness of Speech in Rich Contexts (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for evaluating expressive speech focus on word accuracy, naturalness, signal quality, or emotional intensity at the utterance level.
Approach: They propose a framework for Evaluating Expressive Appropriateness in speech that assesses whether a speech sample aligns with the underlying communicative intent implied by its discourse-level narrative context.
Outcome: The proposed framework outperforms existing speech evaluation and analysis systems on a human-annotated test set.
Multimodal Contrastive Learning via Uni-Modal Coding and Cross-Modal Prediction for Multimodal Sentiment Analysis (2022.findings-emnlp)

Copied to clipboard

Challenge: Recent work on multimodal representation learning has focused on uni-modality pre-training or cross-modalities integration.
Approach: They propose a framework for multimodal representation learning that uses uni-modal contrastive coding and an efficient unimodal feature augmentation strategy to capture intermodal dynamics.
Outcome: The proposed framework surpasses state-of-the-art methods on two public datasets.
Supervised Attention Mechanism for Low-quality Multimodal Data (2025.emnlp-main)

Copied to clipboard

Challenge: Current studies address missing and noisy modalities separately in multimodal data . missing modality is often caused by unavailable data collection equipment or sensor failures .
Approach: They propose a framework for multimodal affective computing that addresses missing and noisy modalities to enhance model robustness in low-quality data scenarios.
Outcome: The proposed model outperforms state-of-the-art baselines on multiple datasets under the settings of complete modalities, missing modalités, and noisy modality.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations