Papers by Zezhong Ding
See or Say Graphs: Agent-Driven Scalable Graph Understanding with Vision-Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing studies have explored textual graph descriptions and visual modalities for VLMs to understand graphs. |
| Approach: | They propose a unified framework that enhances both scalability and modality coordination in graph understanding by integrating textual and visual modalities. |
| Outcome: | GraphVista scales to large graphs, 200 larger than those used in existing benchmarks, and consistently outperforms existing textual, visual, and fusion-based methods. |
GraphInsight: Unlocking Insights in Large Language Models for Graph Structure Understanding (2025.acl-long)
Copied to clipboard
| Challenge: | Large language models struggle with comprehending graphical structure information through prompts of graph description sequences, especially as the graph size increases. |
| Approach: | They propose a framework to improve LLMs’ comprehension of both macro- and micro-level graphical information by placing critical graphical data in positions where LLM's exhibit stronger memory performance. |
| Outcome: | The proposed framework outperforms all other graph description methods in understanding graph structures of varying sizes. |
MVP: Enhancing Video Large Language Models via Self-supervised Masked Video Prediction (2026.acl-long)
Copied to clipboard
| Challenge: | Recent research has attempted to transfer reinforcement learning paradigms to Video Large Language Models (MLLMs) but these methods lack explicit supervision for intrinsic temporal coherence and inter-frame correlations. |
| Approach: | They propose a novel post-training objective: Masked Video Prediction (MVP) that requires the model to reconstruct a masked continuous segment from a set of challenging distractors and employs Group Relative Policy Optimization (GRPO) with a fine-grained reward function to enhance the model's understanding of video context and temporal properties. |
| Outcome: | The proposed model improves video reasoning capabilities by reinforcing temporal reasoning and causal understanding. |