Papers by Chunhui Zhang
M2DF: Multi-grained Multi-curriculum Denoising Framework for Multimodal Aspect-based Sentiment Analysis (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing work mainly utilizes image information to improve the performance of MABSA task. |
| Approach: | They propose a multimodal Aspect-based Sentiment Analysis task that uses image information to improve model performance. |
| Outcome: | The proposed framework outperforms state-of-the-art work on three sub-tasks of MABSA. |
Pretrained Image-Text Models are Secretly Video Captioners (2025.naacl-short)
Copied to clipboard
| Challenge: | Current video captioning methods often incorporate intricate designs tailored to video inputs. |
| Approach: | They adapt an image-based captioning model to address dynamic video sequences without modifications. |
| Outcome: | The proposed model outperforms specialised captioning systems on major benchmarks. |
EFUF: Efficient Fine-Grained Unlearning Framework for Mitigating Hallucinations in Multimodal Large Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to eliminate hallucinations require expensive human annotation . hallucination in multimodal large language models poses unique challenges for current research . |
| Approach: | They propose a fine-grained unlearning framework that performs gradient ascent to eliminate hallucinations without paired data. |
| Outcome: | The proposed method reduces hallucinations while preserving quality with modest computational overhead. |
Working Memory Identifies Reasoning Limits in Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Using large language models, we examine the limitations of their cognitive capabilities and their working memory. |
| Approach: | They examine the limitations of large language models from a scaling perspective . they also assess various prompting strategies, revealing their diverse impacts on LLM performance. |
| Outcome: | The proposed models perform poorly on n-back tasks and on prompting strategies. |
Temporal Working Memory: Query-Guided Segment Refinement for Enhanced Multimodal Understanding (2025.findings-naacl)
Copied to clipboard
Xingjian Diao, Chunhui Zhang, Weiyi Wu, Zhongyu Ouyang, Peijun Qing, Ming Cheng, Soroush Vosoughi, Jiang Gui
| Challenge: | Multimodal foundation models have demonstrated significant success in tasks such as visual captioning, question answering, and image-text retrieval. |
| Approach: | They propose a specialized cognitive module, temporal working memory, which selectively retains task-relevant information across temporal dimensions. |
| Outcome: | The module retains task-relevant information across temporal dimensions, ensuring that critical details are preserved throughout the processing of video and audio content. |
Learning Sparsity for Effective and Efficient Music Performance Question Answering (2025.acl-short)
Copied to clipboard
| Challenge: | Existing Music AVQA methods rely on dense and unoptimized representations, leading to inefficiencies in the isolation of key information, reduction of redundancy, and prioritization of critical samples. |
| Approach: | They propose a sparse learning framework specifically designed for Music AVQA to address these challenges. |
| Outcome: | The proposed framework reduces training time by 28.32% while maintaining accuracy while maintaining state-of-the-art performance on the Music AVQA datasets. |
Knowing More, Acting Better: Hierarchical Representation for Embodied Decision-Making (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Modern embodied AI uses multimodal large language models as policy models, predicting actions from final-layer hidden states. |
| Approach: | They propose a hierarchical action probing method that aggregates representations from all layers, mirroring the brain's multi-level organization. |
| Outcome: | Experiments show that hierarchical probing improves on last-layer embodied models and achieves a 46.6% success rate and a 62.5% gain in spatial reasoning tasks. |
Modality-Aware Neuron Pruning for Unlearning in Multimodal Large Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Large Language Models and Multimodal Large Language Modells can memorize sensitive information, raising ethical and privacy concerns. |
| Approach: | They propose a novel unlearning framework that selectively clips neurons based on their relative importance to the targeted forget data. |
| Outcome: | The proposed framework selectively clips neurons based on their relative importance to the targeted forget data, curated for different modalities. |
Music Audio-Visual Question Answering Requires Specialized Multimodal Designs (2026.findings-acl)
Copied to clipboard
Wenhao You, Xingjian Diao, Wenjun Huang, Chunhui Zhang, Keyi Kong, Weiyi Wu, Chiyu Ma, Zhongyu Ouyang, Tingxuan Wu, Ming Cheng, Soroush Vosoughi, Jiang Gui
| Challenge: | Music audio-visual question answering presents unique challenges with dense audio-visual content, intricate temporal dynamics, and the need for domain-specific knowledge. |
| Approach: | They analyze Music AVQA datasets and analyze their results to identify key design patterns . they propose concrete future directions for incorporating musical priors . |
| Outcome: | The proposed architectures are critical for success in Music AVQA, the authors argue . they suggest concrete future directions for incorporating musical priors . |
Addressing Overthinking in Large Vision-Language Models via Gated Perception-Reasoning Optimization (2026.findings-acl)
Copied to clipboard
Xingjian Diao, Zheyuan Liu, Chunhui Zhang, Weiyi Wu, Keyi Kong, Lin Shi, Kaize Ding, Soroush Vosoughi, Jiang Gui
| Challenge: | Prior work has attempted to mitigate this issue by using adaptive reasoning strategies, but these methods overlook a fundamental bottleneck: visual perception failures. |
| Approach: | They propose a meta-reasoning controller that dynamically routes computation among three decision paths at each generation step. |
| Outcome: | The proposed method outperforms slow-thinking methods while producing shorter responses. |
Superficial Self-Improved Reasoners Benefit from Model Merging (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) rely heavily on large-scale reasoning data, but as data becomes scarce, model self-improvement offers a promising alternative. |
| Approach: | They propose to merge the weights of original and self-improved LLMs to mitigate model collapse and improve generalized reasoning capability. |
| Outcome: | The proposed model merge mitigates model collapse and improves generalized reasoning capability. |
What Makes LLMs Effective Sequential Recommenders? A Study on Preference Intensity and Temporal Context (2026.acl-long)
Copied to clipboard
| Challenge: | Existing preference-alignment approaches rely on binary pairwise comparisons, overlooking preference intensity and temporal context. |
| Approach: | They propose a unified preference optimization framework that maps both explicit and implicit feedback into a common preference signal and constructs adaptive reward margins that jointly account for preference intensity and interaction recency. |
| Outcome: | The proposed framework outperforms state-of-the-art recommendations while maintaining behavioral patterns aligned with human decision-making. |
Is It Navajo? Accurate Language Detection for Endangered Athabaskan Languages (2025.naacl-short)
Copied to clipboard
| Challenge: | Endangered languages are significantly underrepresented in contemporary language technologies, exacerbating the challenges of their preservation and revitalization. |
| Approach: | They propose a random forest classifier trained on Navajo and 20 erroneously suggested languages by Google's Language Identification tool. |
| Outcome: | The proposed classifier achieves near-perfect accuracy across other Athabaskan languages suggesting its potential for broader application. |
Expedited Training of Visual Conditioned Language Generation via Redundancy Reduction (2024.acl-long)
Copied to clipboard
| Challenge: | EVLGen is a framework for visual-language pre-training with high computational demands. |
| Approach: | They propose a streamlined framework for the pre-training of visually conditioned language generation models with high computational demands. |
| Outcome: | The proposed framework accelerates training of vision-language models by a factor of 5 without compromising performance. |
Learning Musical Representations for Music Performance Question Answering (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for audio-visual learning fail to consider the distinctive characteristics of instruments and music. |
| Approach: | They propose to integrate multimodal interactions within the context of music data and annotate and release rhythmic and music sources in the current music datasets to enable the model to learn music characteristics. |
| Outcome: | The proposed model can learn music characteristics from the current music datasets and align its predictions with the temporal dimension. |
SoundMind: RL-Incentivized Logic Reasoning for Audio-Language Models (2025.emnlp-main)
Copied to clipboard
Xingjian Diao, Chunhui Zhang, Keyi Kong, Weiyi Wu, Chiyu Ma, Zhongyu Ouyang, Peijun Qing, Soroush Vosoughi, Jiang Gui
| Challenge: | Recent large language models have demonstrated impressive reasoning abilities, but their extension to the audio modality remains underexplored. |
| Approach: | They propose a rule-based reinforcement learning algorithm to equip LALMs with robust reasoning capabilities. |
| Outcome: | The proposed algorithm improves on the SoundMind benchmark. |
Visibility as Survival: Generalizing NLP for Native Alaskan Language Identification (2025.findings-acl)
Copied to clipboard
| Challenge: | Indigenous languages are largely invisible in commercial language identification systems, a stark reality exemplified by Google Translate’s LangID tool, which excludes all 150 Indigenous languages of North America. |
| Approach: | They propose a framework that shows how large language models and specialized classifiers can effectively identify these languages with minimal data. |
| Outcome: | The proposed framework shows that large language models and specialized classifiers can effectively identify these languages with minimal data. |
Demystifying Uncertainty in LLMs: Active Calibration between Concepts and Human Evaluations (2026.acl-long)
Copied to clipboard
| Challenge: | Existing static strategies for mitigating hallucinations do not explicitly model the information gain from interacting with the external environment. |
| Approach: | They propose a calibration-driven interactive learning strategy that selects clarification queries by optimizing calibration error. |
| Outcome: | The proposed method provides theoretical guarantees and empirical gains for reliability. |
What Makes a Good Curriculum? Disentangling the Effects of Data Ordering on LLM Mathematical Reasoning (2026.acl-long)
Copied to clipboard
| Challenge: | Curriculum learning (CL) orders data corpus by difficulty, but prior work employs disparate difficulty metrics and training setups. |
| Approach: | They propose a framework that decomposes curriculum difficulty into five dimensions: Problem Difficulty, Model Surprisal, Confidence Margin, Predictive Uncertainty and Decision Variability. |
| Outcome: | The proposed framework decomposes curriculum difficulty into five dimensions . the results show that no curriculum strategy dominates universally . |
Growing Through Experience: Scaling Episodic Grounding in Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Language models (LMs) require effective episodic grounding to perform well at physical planning tasks due to their limited ability to learn from and apply past experiences. |
| Approach: | They propose a weak-to-strong episodic learning framework that integrates episodic memory into hierarchical representations and pre-trained knowledge to unlock larger LMs' potential for grounding. |
| Outcome: | The proposed framework outperforms top proprietary LMs by 3.45% across diverse planning and question-answering tasks. |
Behavior Knowledge Merge in Reinforced Agentic Models (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods for supervised fine-tuning (SFT) are suboptimal to preserve task-specific capabilities on RL-trained agentic models. |
| Approach: | They propose a distribution-aware merging framework specifically designed for RL-trained agentic models that disentangles shared and task-specific unique parameter updates while selectively preserving and rescaling unique ones. |
| Outcome: | Experiments across multiple agent domains and model architectures show that the proposed framework surpasses baselines and unlocks synergistic potential among agents. |