Challenge: Large Language Models (LLMs) can better capture cultural and social factors such as viewing intensity and geographic spread of video content.
Approach: They propose to use Large Language Models to capture cultural and social factors that influence video popularity and generate interpretable, attribute-based explanations.
Outcome: The proposed model captures both engagement intensity and geographic spread on 13,639 popular videos, while the neural network's predictions reach 82% without fine-tuning.

Similar Papers

Are Large Language Models (LLMs) Good Social Predictors? (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies suggest that Large Language Models can generate human-like responses, but it is unclear how well they work and where the plausible predictions derive from.
Approach: They propose to use LLMs to generate human-like responses by mutability and accessibility of social inputs to perform a social prediction task.
Outcome: The proposed model performs well in three realistic settings and a novel social prediction task.
Uncovering Bias in Large Vision-Language Models at Scale with Counterfactuals (2025.naacl-long)

Copied to clipboard

Challenge: Large Vision-Language Models (LVLMs) have been proposed to augment LLMs with visual inputs.
Approach: They propose large vision-Language Models to augment LLMs with visual inputs.
Outcome: The proposed models condition generated text on both an input image and a visual prompt, enabling a variety of use cases such as visual question answering and multimodal chat.
Large Language Models for Generative Recommendation: A Survey and Visionary Discussions (2024.lrec-main)

Copied to clipboard

Challenge: Large language models (LLMs) have revolutionized the field of natural language processing but are not fully able to leverage the generative power of LLM.
Approach: They examine the progress, methods, and future directions of large language models . they examine what generative recommendation is, why RS should advance to generative recommendations .
Outcome: The proposed approach can be simplified to generate recommendations from the entire pool of items.
A Video Is Worth 4096 Tokens: Verbalize Videos To Understand Them In Zero Shot (2023.emnlp-main)

Copied to clipboard

Challenge: Existing annotated training datasets hinder development of supervised learning models for multimedia content . lack of annotating benchmarks hinders development of models with satisfactory performance . a recent study shows that large language models have zero-shot performance in multimedia understanding .
Approach: They propose to verbalize long videos to generate their descriptions in natural language . they then perform video-understanding tasks on the generated story as opposed to the original video .
Outcome: The proposed method achieves better results than baselines for video understanding.
Unveiling Multi-level and Multi-modal Semantic Representations in the Human Brain using Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Recent studies have assessed different levels of semantic content, such as speech, objects, and stories, separately.
Approach: They used functional magnetic resonance imaging to record brain activity while watching 8.3 hours of dramas and movies.
Outcome: The findings show that LLMs predict human brain activity more accurately than traditional language models, particularly for complex background stories.
Multimodal Large Language Models for Text-rich Image Understanding: A Comprehensive Review (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in vision-language models have unified perception and understanding tasks within Visual Question Answering paradigms.
Approach: They propose to outline timeline, architecture, and pipeline of nearly all TIU MLLMs and review their performance on mainstream benchmarks.
Outcome: The proposed models perform well on mainstream benchmarks and are compared with other models.
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection (2024.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to visual-language understanding lack unified tokenization for images and videos . lack of unified visual representations makes it difficult to learn multi-modal interactions from poor projection layers.
Approach: They propose to unify visual representation into the language feature space to advance the foundational LLM towards a unified LVLM.
Outcome: The proposed model outperforms Video-ChatGPT on image benchmarks and on 9 image benchmark benchmarks.
From Static Inference to Dynamic Interaction: A Survey of Streaming Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing definitions of streaming LLMs are fragmented and lack a systematic taxonomy . large language models are pre-trained on static and full-context corpora .
Approach: They propose a systematic taxonomy of current streaming Large Language Models and propose underlying methodologies for streaming LLMs.
Outcome: The proposed model is based on data flow and dynamic interaction to clarify existing ambiguities.
Eliciting In-Context Learning in Vision-Language Models for Videos Through Curated Data Distributional Properties (2024.emnlp-main)

Copied to clipboard

Challenge: Emergent In-context Learning on Videos induces in-contact learning over video and text . eILeV-trained models outperform other off-the-shelf VLMs in few-shot video narration for novel, rare actions.
Approach: They implement Emergent In-context Learning on Videos (EILeV) that induces in-contact learning over video and text by capturing key properties of pre-training data.
Outcome: The proposed training paradigm outperforms off-the-shelf VLMs in few-shot video narration for novel, rare actions.
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models (2024.acl-long)

Copied to clipboard

Challenge: a surge of deep learning applications for video understanding have led to major advancements in video-related tasks.
Approach: They propose a multimodal video-based conversation model that merges a video-adapted visual encoder with an LLM and a dataset that is easily scalable and robust to label noise.
Outcome: The proposed model can understand and generate detailed conversations about videos.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations