Challenge: EmoOmni is a data paradigm for omni-modal large language models that can be used for emotion reasoning.
Approach: They propose a data paradigm that interleaves guided tokens into reasoning traces to enforce structured evidence extraction.
Outcome: The proposed paradigm over-relys on a dominant modality while neglecting complementary cues.

Similar Papers

Omni-RewardBench: Toward a Comprehensive Evaluation of Generative Reward Models Across Modalities (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluation benchmarks for ORMs are largely text-centric or limited to bimodal tasks . a new study examines the effectiveness of Omni-RewardBench for ORms across modalities .
Approach: They propose a hybrid automatic-annotation and human-verification pipeline to construct high-quality evaluation data.
Outcome: The proposed model is the first benchmark for comprehensive evaluation of ORMs across modalities.
Cross-Modal Coreference Alignment: Enabling Reliable Information Transfer in Omni-LLMs (2026.acl-long)

Copied to clipboard

Challenge: Experiments on 13 Omni-LLMs reveal systematic weaknesses in cross-modal coreference . cross-module coreference is a crucial missing piece for advancing robust omni-modal reasoning.
Approach: They propose a cross-modal coreference problem to evaluate and enhance Omni-LLMs' reasoning capabilities.
Outcome: Experiments on 13 Omni-LLMs show they lack coreference-aware thinking patterns . the CROSSOMNI dataset yields significant performance gains and generalizes well to collaborative reasoning tasks.
e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings (2026.findings-acl)

Copied to clipboard

Challenge: Recent omni-modal embeddings rely heavily on implicit alignment from pretrained visionlanguage models.
Approach: They propose a lightweight explicit alignment recipe that adapts off-the-shelf VLMs into robust omni-modal embedding models.
Outcome: The proposed model improves on MMEB-V2 and AudioCaps with a lightweight explicit alignment recipe.
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities (2025.findings-acl)

Copied to clipboard

Challenge: MLLMs are able to integrate multiple modalities into a single model to tackle complex tasks in real-world scenarios.
Approach: They propose a comprehensive survey of Omni-MLLMs to address the challenges and opportunities of multimodal modeling.
Outcome: The proposed model can integrate multiple modalities into a single model and provide novel perspectives.
OMHBench: Benchmarking Balanced and Grounded Omni-Modal Multi-Hop Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluation frameworks for multimodal large language models suffer from limitations . modality shortcuts and biased reasoning paths are common in such models .
Approach: a new benchmark evaluates omni-modal multi-hop reasoning using 6,144 questions . authors propose OMHBench to address these limitations by comparing modalities .
Outcome: OMHBench evaluates omni-modal multi-hop reasoning on 6,144 questions with balanced reasoning paths . evaluation of 13 state-of-the-art models shows performance gap exists between MLLMs and open-source models .
VAPO: End-to-end Slide-Enhanced Speech Recognition with Omni-modal Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Current Automatic Speech Recognition models, such as Whisper, have demonstrated impressive performance in general domains, but their accuracy often deteriorates significantly in specialized scenarios.
Approach: They propose a visually-anchored policy optimization approach to decouple visual perception from auditory processing to optimize the model's inference process.
Outcome: The proposed model eliminates visual interference and achieves state-of-the-art performance on SlideASR-Bench and public datasets.
GRPO-Guided Modality Selection Enhanced LoRA-Tuned LLMs for Multimodal Emotion Recognition (2025.findings-emnlp)

Copied to clipboard

Challenge: Multimodal emotion recognition in conversation (MERC) aims to identify speakers’ emotional states by utilizing text, audio, and visual modalities.
Approach: They propose an adaptive modality selection framework for multimodal emotion recognition in conversation that integrates all available modalities into one .
Outcome: The proposed framework outperforms existing methods on multimodal dialogue datasets and is available at https://github.com/youflyaway/Modality-Selection-Enhanced-LoRA-Tuned-LLMs.
Recognizing Everything from All Modalities at Once: Grounded Multimodal Universal Information Extraction (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies on IE tasks have focused on recognizing and analyzing cross-modal information . a multimodal large language model (MLLM) is developed to analyze IE across modalities .
Approach: They propose a multimodal large language model (MLLM) capable of grounding information from all modalities.
Outcome: The proposed framework provides a framework to analyze IE tasks over various modalities and their fine-grained groundings.
Layer-wise Fusion with Modality Independence Modeling for Multi-modal Emotion Recognition (2023.acl-long)

Copied to clipboard

Challenge: Existing studies focus on developing models that exploit the unification of multiple modalities.
Approach: They propose to maintain modality independence by using a multi-modal transformer model that fuses all modalities.
Outcome: The proposed model outperforms state-of-the-art models in multi-modal emotion recognition.
MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for expert parallelism inference suffer from a significant efficiency bottleneck . existing methods fail to address information heterogeneity and modality dynamics .
Approach: They propose a training-free inference framework that scales experts without training . they propose an Entropy-Weighted Load mechanism to quantify the semantic value of visual tokens .
Outcome: Experiments show that MACS outperforms existing methods on multimodal benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations