Papers by Quanfeng Lu
TVWorld: Foundations for Remote-Control TV Agents (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing work on large vision–language models focuses on point-and-click interaction, while remote-control interaction is underexplored. |
| Approach: | They propose a topology-aware training framework that injects topology awareness into LVLMs. |
| Outcome: | The proposed model achieves 68.3% success rate on TVWorld-N, surpassing closed-source benchmarks and state-of-the-art (SOTA) benchmarks show that existing agents lack topology awareness for focus-based, long-horizon TV navigation. |
ChartAssistant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning (2024.findings-acl)
Copied to clipboard
| Challenge: | Charts are an effective tool for understanding data patterns, but their combination of graphical elements and textual components poses challenges for general-purpose multimodal models. |
| Approach: | They propose a chart-based vision-language model for universal chart comprehension and reasoning that leverages a dataset of chart-related tasks. |
| Outcome: | The proposed model outperforms the state-of-the-art charts with zero-shot setting on various chart tasks. |