Papers by Xiaoyang Yi

5 papers
CoreGaze: Core Subgraph-Driven Visual Gaze Diffusion for Training-Free Referring Multimodal Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing methods rely on extensive fine-tuning to mitigate attention distraction, leading to redundant outputs or hallucinations.
Approach: They propose a training-free framework that simulates human visual gaze diffusion for fine-grained comprehension by combining a sparse semantic graph with a core subgraph with amplified initial influence.
Outcome: The proposed framework simulates human visual gaze diffusion for fine-grained comprehension.
Integrating Structural Semantic Knowledge for Enhanced Information Extraction Pre-training (2024.emnlp-main)

Copied to clipboard

Challenge: Existing pre-training methods focus on exploiting textual knowledge, which limits scalability and versatility of resulting models.
Approach: They propose a pre-training framework that integrates structural semantic knowledge via contrastive learning.
Outcome: The proposed framework outperforms state-of-the-art pre-training methods across multiple tasks.
RISER: Orchestrating Latent Reasoning Skills for Adaptive Activation Steering (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for domain-specific reasoning with large language models require updating parameter updates.
Approach: They propose a plug-and-play intervention framework that adaptively steers LLM reasoning in activation space.
Outcome: The proposed framework achieves zero-shot accuracy improvements of 3.4–6.5% over the base model while outperforming chain-of-thought-style reasoning with 2–3 higher token efficiency and robust accuracy gains.
HVGuard: Utilizing Multimodal Large Language Models for Hateful Video Detection (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for hateful video detection rely on unimodal analysis or feature fusion . Existing tools struggle to capture cross-modal interactions and reason through implicit hate in sarcasm and metaphor .
Approach: They propose a reasoning-based hateful video detection framework with multimodal large language models . they integrate Chain-of-Thought reasoning to enhance multimodal interaction modeling .
Outcome: The proposed framework outperforms existing tools on two public datasets covering English and Chinese.
Beyond the Panorama: Training-Free Hierarchical Perception-Reasoning for Fine-Grained Vision in MLLMs (2026.acl-long)

Copied to clipboard

Challenge: Existing multimodal large language models (MLLMs) face challenges in fine-grained visual tasks.
Approach: They propose a training-free hierarchical perception-reasoning framework that enhances fine-grained visual understanding by simulating human perception mechanisms.
Outcome: The proposed framework enhances fine-grained visual understanding by simulating human perception mechanisms.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations