Papers by Fanheng Kong

6 papers
TIGER: A Unified Generative Model Framework for Multimodal Dialogue Response Generation (2024.lrec-main)

Copied to clipboard

Challenge: Existing research on multimodal dialogues focuses on textual response generation and visual response selection based on the dialogue context.
Approach: They propose a generative model framework for multimodal dialogue response generation that ground the conversation on an image.
Outcome: The proposed system provides users with an enhanced conversational experience.
STICKERCONV: Generating Multimodal Empathetic Responses from Scratch (2024.acl-long)

Copied to clipboard

Challenge: Prior studies on stickers focused on sentiment analysis and recommendation systems, overlooking their vast potential in empathetic response generation.
Approach: They propose a multimodal empathetic dialogue dataset, STICKERCONV, which simulates human behavior with stickers, and propose evaluative metrics based on LLM.
Outcome: The proposed framework generates contextually relevant and emotionally resonant multimodal empathetic responses, contributing to the advancement of more nuanced and engaging e-dialog systems.
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search (2025.acl-long)

Copied to clipboard

Challenge: Existing benchmarks and evaluation protocols suffer from inadequate or homogeneous creation of key points, exorbitant cost of data creation, and limited evaluation scopes.
Approach: They propose an automatic framework which leverages Monte Carlo Tree Search to construct numerous and diverse descriptive sentences that thoroughly represent video content in an iterative way.
Outcome: The proposed framework improves MCTS-VCB and DREAM-1K on video captioning tasks by 25.0% and 16.3% respectively.
DPN-LE: Dual Personality Neuron Localization and Editing for Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Current methods for editing personality traits in large language models can change personalities but reduce performance.
Approach: They propose a novel paradigm for personality editing that locates and edits LLM neurons and enables competitive personality control at inference time.
Outcome: Experiments on LLaMA-3-8B-Instruct and Qwen2.5-7B-instruct show that the proposed approach can improve performance and improve performance.
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos (2025.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for video understanding often focus on specific aspects, overlooking the holistic nature of video content.
Approach: They propose a temporal-oriented benchmark for fine-grained understanding on dense dynamic videos with two complementary tasks: captioning and QA.
Outcome: The proposed model performs well on diverse video scenarios and dynamic videos, with interpretable and robust evaluation criteria.
RATION: Entropy-Driven Task-Adaptive Visual Attention Allocation Framework for Multimodal Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Prior studies have focused on strengthening multimodal reasoning by improving representation alignment or increasing computation, but these methods do not characterize the differences in visual demands across tasks.
Approach: They propose an entropy-driven task-adaptive visual attention allocation framework that uses visual attention entropic as a control signal to dynamically allocate attention according to task demands.
Outcome: The proposed framework achieves consistent performance gains across diverse reasoning tasks, datasets, and models, providing a clear direction toward more reliable multimodal reasoning.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations