Papers by Haoning Zhang

6 papers
Generative Frame Sampler for Long Video Understanding (2025.findings-acl)

Copied to clipboard

Challenge: Existing video large language models (LMMs) employ an impedance of thousands of frames to understand long videos.
Approach: They propose a plug-and-play module integrated with VideoLLMs to facilitate efficient lengthy video perception.
Outcome: The proposed module boosts the performance of open-source VideoLLMs and proprietary assistants on long-form video benchmarks.
MTR-DuplexBench: Towards a Comprehensive Evaluation of Multi-Round Conversations for Full-Duplex Speech Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks focus on evaluating single-round interactions, neglecting other critical aspects.
Approach: They propose a benchmark to evaluate full-duplex speech language models in multi-round settings . they segment continuous full-dual dialogues into discrete turns for evaluation .
Outcome: The proposed benchmark compared full-duplex speech language models with full-dual speech models . the results show that the models perform better in multi-round settings than standard models compared to benchmarks .
MoNET: Tackle State Momentum via Noise-Enhanced Training for Dialogue State Tracking (2023.findings-acl)

Copied to clipboard

Challenge: Experimental results show that MoNET outperforms previous DST methods in alleviating state momentum issues and improving the anti-noise ability.
Approach: They propose to use previous state of each turn in training data as input to learn to predict current state.
Outcome: The proposed model outperforms existing methods on multiWOZ datasets and shows that it can update and correct slot values and improve anti-noise ability.
MMRA: A Benchmark for Evaluating Multi-Granularity and Multi-Image Relational Association Capabilities in Large Visual Language Models (2026.findings-eacl)

Copied to clipboard

Challenge: Current multimodal benchmarks focus on facts within individual images, but neglect associative relations among multiple images.
Approach: They propose a multi-image relational association task and a MMRA benchmark to evaluate LVLMs.
Outcome: The proposed benchmarks show that entity-level multi-image perception tasks pose greater challenges than image-level tasks.
CSS: Combining Self-training and Self-supervised Learning for Few-shot Dialogue State Tracking (2022.aacl-short)

Copied to clipboard

Challenge: Existing few-shot dialogue state tracking (DST) methods transfer knowledge from labeled data into DST, but collecting large amount of labeles is laborious.
Approach: They propose a few-shot dialogue state tracking framework that integrates self-training and self-supervised learning methods into the framework.
Outcome: The proposed framework achieves competitive performance in several few-shot scenarios.
LIME: Less Is More for MLLM Evaluation (2025.findings-acl)

Copied to clipboard

Challenge: Existing MLLM benchmarks and unified evaluation frameworks cannot accurately and efficiently reflect the ability of MLMLs.
Approach: They propose a semi-automated benchmark curated using a pipeline that filters out uninformative samples and eliminates answer leakage by focusing on tasks that require image-based understanding.
Outcome: The proposed benchmark reduces the number of samples by 76% and evaluation time by 77% while it can more effectively distinguish different models’ abilities.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations