Challenge: Existing models struggle to maintain temporal, spatial, and narrative coherence across image sequences . existing models lack depth and engagement of human-authored stories .
Approach: They propose a topic-driven narrative optimizer that integrates image descriptions, topic generation, and GPT-4-based refinements.
Outcome: The proposed framework outperforms existing models in visual relevance, coherence, and fluency.

Similar Papers

Are Large Language Models Capable of Generating Human-Level Narratives? (2024.emnlp-main)

Copied to clipboard

Challenge: a recent HCI study has pointed to gaps in machine storytelling ability at the global level . authors show that LLMs have less suspense and less tension than human stories .
Approach: They propose a computational framework to analyze narratives through three discourse-level aspects.
Outcome: The proposed framework analyzes narratives through three discourse-level aspects . it shows that LLMs fall short of human abilities in discourse understanding .
MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct (2025.findings-acl)

Copied to clipboard

Challenge: a new framework for image-text instruction data evolution improves MLLM performance . lack of high-quality instruction data remains a major bottleneck in ML modeling .
Approach: They propose a multimodal instruction data evolution framework that iteratively enhances data quality through fine-grained perception, cognitive reasoning, and interaction evolution.
Outcome: The proposed approach improves MLLM performance in nine vision-language tasks while using significantly less data.
The Revolution of Multimodal Large Language Models: A Survey (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models have led to the development of multimodal large language model.
Approach: They present a review of recent visual-based Large Language Models and analyze their architectures and alignment strategies.
Outcome: The proposed models can integrate visual and textual modalities while providing a dialogue-based interface and instruction-following capabilities.
Learning to Model Multimodal Semantic Alignment for Story Visualization (2022.findings-emnlp)

Copied to clipboard

Challenge: Story visualization aims to generate sequence of images to narrate each sentence in a multi-sentence story . current methods face semantic misalignment because of their fixed architecture and diversity of input modalities .
Approach: They propose to use a GAN-based generative model to match semantic levels between text and image representations to solve the semantic misalignment problem.
Outcome: Experiments show that the proposed approach improves image quality and story consistency compared with state-of-the-art methods.
Multimodal Large Language Models for Text-rich Image Understanding: A Comprehensive Review (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in vision-language models have unified perception and understanding tasks within Visual Question Answering paradigms.
Approach: They propose to outline timeline, architecture, and pipeline of nearly all TIU MLLMs and review their performance on mainstream benchmarks.
Outcome: The proposed models perform well on mainstream benchmarks and are compared with other models.
Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback (2024.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models have influenced the development of video large multimodal models (VLMMs).
Approach: They propose a method that integrates video descriptions as context into a multimodal AI system to enrich the understanding of video content.
Outcome: Empirical evaluations show that the proposed approach outperforms existing approaches for video large multimodal models (VLMMs)
LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture (2025.findings-emnlp)

Copied to clipboard

Challenge: Long-context Large Language Models (MLLMs) are critical for video understanding and image analysis.
Approach: They propose a hybrid architecture that integrates Mamba and Transformer blocks . they introduce data construction methods that capture both temporal and spatial dependencies .
Outcome: The proposed model achieves competitive results across various benchmarks while maintaining high throughput and low memory consumption.
DataNarrative: Automated Data-Driven Storytelling with Visualizations and Texts (2024.emnlp-main)

Copied to clipboard

Challenge: Data-driven storytelling uses visual aids and visualizations to convey insights.
Approach: They propose a task for data story generation using large language models and a benchmark containing 1,449 stories from diverse sources.
Outcome: The proposed framework outperforms non-agentic counterparts in both model-based and human evaluations, but also reveals unique challenges in data story generation.
Towards Unified Multimodal Large Language Models: A survey (2026.findings-acl)

Copied to clipboard

Challenge: unified multimodal large language models (MLLMs) are emerging but lack a systematic framework to connect them and situate current trends within a broader landscape.
Approach: They present a systematic review of unified Multimodal Large Language Models . they outline the foundational concepts and prerequisites for understanding them .
Outcome: The present review provides a systematic and systematic overview of unified MLLMs . it discusses persistent challenges and identify promising directions for future research .
VFA: Empowering Multilingual MLLMs via Vision-Free Adaptation (2026.acl-long)

Copied to clipboard

Challenge: Multimodal large language models have advanced rapidly, yet most remain English-centric . scaling multilingual multimodal instruction tuning is limited by the scarcity and high cost of non-English image–text supervision.
Approach: They propose a framework that decouples multilingual language enhancement from visual alignment by composing complementary task vectors over a shared LLM backbone.
Outcome: The proposed framework achieves competitive performance with a fully multimodally trained model using less than 2% of the text data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations