Papers by Jianbing Zhang
On Prefix-tuning for Lightweight Out-of-distribution Detection (2023.acl-long)
Copied to clipboard
| Challenge: | Out-of-distribution (OOD) detection is a fundamental task vexing real-world applications . fine-tuning based methods require storing fine- tuned models for each scenario . |
| Approach: | They propose an unsupervised prefix-tuning based OOD detection framework called PTO . they propose to take advantage of optional training data labels and targeted OOD data . |
| Outcome: | The proposed framework performs better than existing methods under a wide range of metrics, detection settings, and OOD types. |
M2DF: Multi-grained Multi-curriculum Denoising Framework for Multimodal Aspect-based Sentiment Analysis (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing work mainly utilizes image information to improve the performance of MABSA task. |
| Approach: | They propose a multimodal Aspect-based Sentiment Analysis task that uses image information to improve model performance. |
| Outcome: | The proposed framework outperforms state-of-the-art work on three sub-tasks of MABSA. |
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents (2024.acl-long)
Copied to clipboard
| Challenge: | Existing GUI agents interact with the environment through extracted structured data, which can be notably lengthy (e.g., HTML) and occasionally inaccessible (e-book). |
| Approach: | They propose to enhance SeeClick with GUI grounding pre-training and devise a method to automate curation of GUI ground data. |
| Outcome: | The proposed agent improves ScreenSpot, the first realistic GUI grounding benchmark that encompasses mobile, desktop, and web environments. |
Vision-Language Models Can Self-Improve Reasoning via Reflection (2025.naacl-long)
Copied to clipboard
| Challenge: | Chain-of-thought (CoT) has been shown to improve the reasoning capability of large language models (LLMs). |
| Approach: | They propose a framework which iteratively enhances the model’s Vision-language Reasoning by Reflecting on CoT Rationales. |
| Outcome: | The proposed framework improves multimodal reasoning on vision-language tasks by 23% to 60% over baselines. |
EFUF: Efficient Fine-Grained Unlearning Framework for Mitigating Hallucinations in Multimodal Large Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to eliminate hallucinations require expensive human annotation . hallucination in multimodal large language models poses unique challenges for current research . |
| Approach: | They propose a fine-grained unlearning framework that performs gradient ascent to eliminate hallucinations without paired data. |
| Outcome: | The proposed method reduces hallucinations while preserving quality with modest computational overhead. |
MixRED: A Mix-lingual Relation Extraction Dataset (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing research focuses on monolingual relation extraction, but there is a significant gap in understanding relation extraction in the mix-lingual scenario. |
| Approach: | They propose a task of considering relation extraction in the mix-lingual scenario . they construct a human-annotated dataset to support the task . |
| Outcome: | The proposed task evaluates state-of-the-art supervised models and large language models on the human-annotated dataset MixRED. |
CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era (2025.findings-acl)
Copied to clipboard
Kanzhi Cheng, Wenpo Song, Jiaxin Fan, Zheng Ma, Qiushi Sun, Fangzhi Xu, Chenyang Yan, Nuo Chen, Jianbing Zhang, Jiajun Chen
| Challenge: | Image captioning has been a challenge for vision-language researchers for decades . current VLMs focus on tasks like visual question answering (YA) but image captioning is not as advanced as expected. |
| Approach: | They evaluate VLMs' performance on image captioning using human annotations . they find that some metrics show high caption-level agreement with humans . |
| Outcome: | The proposed model outperforms open-source models on image captioning . it achieves 93.4% correlation with human rankings at $4 per test . |
Local Interpretation of Transformer Based on Linear Decomposition (2023.acl-long)
Copied to clipboard
| Challenge: | Existing work on local explanation generation attempts to understand model dynamics on word-level or phraselevel by assigning importance scores on input features. |
| Approach: | They propose to interpret neural networks by linear decomposition by a Transformer model on a single input and a linear decomposing of the output to generate local explanations. |
| Outcome: | The proposed method achieves competitive performance in sentiment classification and machine translation, and fidelity of explanation. |
Addressing Linguistic Bias through a Contrastive Analysis of Academic Writing in the NLP Domain (2023.emnlp-main)
Copied to clipboard
| Challenge: | a reviewer’s opinion of the nativeness of expression in an academic paper affects the likelihood of it being accepted for publication. |
| Approach: | They conduct a statistical analysis of paper abstracts from the natural language processing domain to identify how authors from different linguistic backgrounds differ in the lexical, morphological, syntactic and cohesive aspects of their writing. |
| Outcome: | The results suggest that there is potential for linguistic bias in the domain of natural language processing. |
Probing Cross-modal Semantics Alignment Capability from the Textual Perspective (2022.findings-emnlp)
Copied to clipboard
| Challenge: | In recent years, vision and language pre-training (VLP) models have advanced the state-of-the-art results in a variety of cross-modal downstream tasks. |
| Approach: | They propose a new probing method that is based on image captioning to first empirically study the cross-modal semantics alignment of VLP models. |
| Outcome: | The proposed method analyzes captions generated by five popular VLP models to reveal how well they align with visual words and how well these align with images. |
Learning Representation Mapping for Relation Detection in Knowledge Base Question Answering (P19-1)
Copied to clipboard
| Challenge: | Existing approaches to detect relation detection only get high accuracy for questions whose relations have been seen in training data. |
| Approach: | They propose a method to learn representation mapping for both seen and unseen relations based on previously learned relation embedding. |
| Outcome: | The proposed method improves the performance of unseen relations while keeping the performance comparable to the state-of-the-art. |
Improving Long-Context Translation via Self-Supervised Dual Learning (2026.acl-long)
Copied to clipboard
| Challenge: | Large language models with long context windows suffer from catastrophic information distortion, undermining the strict faithfulness required for translation. |
| Approach: | They propose a self-supervised post-training framework that improves long-document translation reliability via round-trip consistency. |
| Outcome: | The proposed framework improves long-document translation reliability via round-trip consistency. |