Papers by Guohui Zhang

2 papers
MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluation benchmarks for Multimodal Large Language Models (MLLMs) focus on single-turn question answering, overlooking the complexity of multi-turn dialogues in real-world scenarios.
Approach: They propose a video understanding benchmark for MLLMs in multi-turn dialogues that assesses six core competencies that focus on perceptivity and interactivity.
Outcome: The MT-Video-Bench evaluates 1,000 multi-turn dialogues from diverse domains and reveals significant performance discrepancies and limitations in handling multi-turned video dialogues.
Intention Reasoning Network for Multi-Domain End-to-end Task-Oriented Dialogue (2021.emnlp-main)

Copied to clipboard

Challenge: Recent years has witnessed the remarkable success in end-to-end task-oriented dialog system, especially when incorporating external knowledge information.
Approach: They propose a mechanism to model deterministic entity knowledge by using an intention reasoning network to obtain intention-aware representations of conceptual tokens.
Outcome: The proposed mechanism captures concept shifts and generates accurate responses on two representative multi-domain dialog datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations