Papers by Xiaoyang Yi
CoreGaze: Core Subgraph-Driven Visual Gaze Diffusion for Training-Free Referring Multimodal Large Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods rely on extensive fine-tuning to mitigate attention distraction, leading to redundant outputs or hallucinations. |
| Approach: | They propose a training-free framework that simulates human visual gaze diffusion for fine-grained comprehension by combining a sparse semantic graph with a core subgraph with amplified initial influence. |
| Outcome: | The proposed framework simulates human visual gaze diffusion for fine-grained comprehension. |
Integrating Structural Semantic Knowledge for Enhanced Information Extraction Pre-training (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing pre-training methods focus on exploiting textual knowledge, which limits scalability and versatility of resulting models. |
| Approach: | They propose a pre-training framework that integrates structural semantic knowledge via contrastive learning. |
| Outcome: | The proposed framework outperforms state-of-the-art pre-training methods across multiple tasks. |
RISER: Orchestrating Latent Reasoning Skills for Adaptive Activation Steering (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for domain-specific reasoning with large language models require updating parameter updates. |
| Approach: | They propose a plug-and-play intervention framework that adaptively steers LLM reasoning in activation space. |
| Outcome: | The proposed framework achieves zero-shot accuracy improvements of 3.4–6.5% over the base model while outperforming chain-of-thought-style reasoning with 2–3 higher token efficiency and robust accuracy gains. |
HVGuard: Utilizing Multimodal Large Language Models for Hateful Video Detection (2025.emnlp-main)
Copied to clipboard
Yiheng Jing, Mingming Zhang, Yong Zhuang, Jiacheng Guo, Juan Wang, Xiaoyang Xu, Wenzhe Yi, Keyan Guo, Hongxin Hu
| Challenge: | Existing methods for hateful video detection rely on unimodal analysis or feature fusion . Existing tools struggle to capture cross-modal interactions and reason through implicit hate in sarcasm and metaphor . |
| Approach: | They propose a reasoning-based hateful video detection framework with multimodal large language models . they integrate Chain-of-Thought reasoning to enhance multimodal interaction modeling . |
| Outcome: | The proposed framework outperforms existing tools on two public datasets covering English and Chinese. |
Beyond the Panorama: Training-Free Hierarchical Perception-Reasoning for Fine-Grained Vision in MLLMs (2026.acl-long)
Copied to clipboard
| Challenge: | Existing multimodal large language models (MLLMs) face challenges in fine-grained visual tasks. |
| Approach: | They propose a training-free hierarchical perception-reasoning framework that enhances fine-grained visual understanding by simulating human perception mechanisms. |
| Outcome: | The proposed framework enhances fine-grained visual understanding by simulating human perception mechanisms. |