Papers by Muhammad Maaz

2 papers
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models (2024.acl-long)

Copied to clipboard

Challenge: a surge of deep learning applications for video understanding have led to major advancements in video-related tasks.
Approach: They propose a multimodal video-based conversation model that merges a video-adapted visual encoder with an LLM and a dataset that is easily scalable and robust to label noise.
Outcome: The proposed model can understand and generate detailed conversations about videos.
A Culturally-diverse Multilingual Multimodal Video Benchmark & Model (2025.emnlp-main)

Copied to clipboard

Challenge: Large multimodal models have gained attention for their effectiveness to understand and generate descriptions of visual content.
Approach: They propose a multilingual Video LMM benchmark to evaluate video LMMs across 14 languages . they also introduce a machine translated multilingual video training set .
Outcome: The proposed video LMM benchmark is designed to evaluate video Lmms across 14 languages including Arabic, Bengali, Chinese, English, French, German, Hindi, Japanese, Russian, Sinhala, Spanish, Swedish, Tamil, and Urdu.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations