Challenge: Visual instruction tuning is the predominant technology in eliciting multimodal task-solving capabilities of large vision-language models.
Approach: They propose a visual instruction-free fine-tuning framework for large vision-language models . they require only text-only instructions and image caption data during training .
Outcome: The proposed framework is based on visual instruction tuning, but requires images as input . it can achieve state-of-the-art performance on several downstream benchmarks with less training data.

Similar Papers

VIT-Pro: Visual Instruction Tuning for Product Images (2025.naacl-industry)

Copied to clipboard

Challenge: general-purpose vision-language models struggle to understand and converse about real-world e-commerce product images.
Approach: a new approach is proposed to use large-scale image-text pairs to train a generative VLM for e-commerce product images.
Outcome: The proposed model outperforms general-purpose VLMs on multiple vision tasks in the e-commerce domain.
Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning (2024.findings-acl)

Copied to clipboard

Challenge: Recent vision-language models (VLMs) have shown impressive capabilities as general visual assistants, but there are two challenges to their performance: (1) lacking task diversity in pretraining and visual instruction tuning; (2) annotation error and bias in GPT-4 synthesized instruction tuning data.
Approach: They propose a two-stage instruction tuning framework that fine tunes VLMs firstly and further tuned on GPT-4 synthesized data.
Outcome: The proposed framework outperforms the traditional single-stage visual instruction tuning framework and achieves state-of-the-art performance across a wide range of multi-modal evaluation benchmarks.
VFA: Empowering Multilingual MLLMs via Vision-Free Adaptation (2026.acl-long)

Copied to clipboard

Challenge: Multimodal large language models have advanced rapidly, yet most remain English-centric . scaling multilingual multimodal instruction tuning is limited by the scarcity and high cost of non-English image–text supervision.
Approach: They propose a framework that decouples multilingual language enhancement from visual alignment by composing complementary task vectors over a shared LLM backbone.
Outcome: The proposed framework achieves competitive performance with a fully multimodally trained model using less than 2% of the text data.
Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in instruction-tuned Large Vision-Language Models (LVLMs) have imbued the models with the ability to generate high-level, image-grounded explanations with ease.
Approach: They propose to use a multiple granularity attribute-centric benchmark and training mixture to evaluate LVLMs’ fine-grained visual comprehension ability.
Outcome: The proposed model improves on LLaVa-1.5, InstructBLIP and GPT-4V and demonstrates that they struggle to generate descriptive visual attributes based on a concept that appears within an input image despite their prominent zero-shot image captioning ability.
An Empirical Study on Parameter-Efficient Fine-Tuning for MultiModal Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Multimodal Large Language Models fine-tuned with multimodal instruction-following data have demonstrated formidable capabilities in multimodal tasks.
Approach: They propose to employ four PEFT methods to fine-tune the LLM component of open-source MLLMs.
Outcome: The proposed method is the best performing on seven datasets, while fine-tuning the connector layers leads to improved performance in most MLLMs.
Vision-Language Model Fine-Tuning via Simple Parameter-Efficient Modification (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in fine-tuning Vision-Language Models have seen the success of prompt tuning and adapter tuning.
Approach: They propose a method to fine-tune CLIP without introducing any overhead of extra parameters.
Outcome: The proposed method improves CLIP by 7.27% average harmonic mean accuracy.
Your Vision-Language Model Itself Is a Strong Filter: Towards High-Quality Instruction Tuning with Data Selection (2024.findings-acl)

Copied to clipboard

Challenge: Existing data selection methods for instruction-following large language models rely on unreliable scores or use downstream tasks for selection.
Approach: They propose a method that utilizes the VLM itself as a filter to select high-quality instruction-tuning data.
Outcome: The proposed method can reach better results compared to full data settings with merely about 15% samples and can achieve superior performance against competitive baselines.
SIFT-50M: A Large-Scale Multilingual Dataset for Speech Instruction Fine-Tuning (2025.acl-long)

Copied to clipboard

Challenge: Recent years have witnessed significant advancements in integrating speech and audio capabilities into large language models.
Approach: They propose a 50M-example dataset for instruction fine-tuning and pre-training of speech-text large language models (LLMs) the dataset spans five languages and enables a diverse range of speech understanding and controllable speech generation instructions.
Outcome: The proposed dataset outperforms existing speech-text LLMs on instruction-following benchmarks while achieving competitive performance on foundational speech tasks.
PEFT-Factory: Unified Parameter-Efficient Fine-Tuning of Autoregressive Large Language Models (2026.eacl-demo)

Copied to clipboard

Challenge: Parameter-Efficient Fine-Tuning (PEFT) methods address the increasing size of Large Language Models (LLMs).
Approach: They propose a framework for efficient fine-tuning Large Language Models (LLMs) they aim to train only a small percentage of the full model's parameters .
Outcome: Xu et al., 2023; Ding e t al, 2024; Lialin e al. 2023) show that using PEFT methods can improve performance.
Let’s Focus on Neuron: Neuron-Level Supervised Fine-tuning for Large Language Model (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are composed of neurons that exhibit diverse behaviors and roles.
Approach: They propose a novel approach that refines the granularity of parameter training down to the individual neuron, enabling a more parameter-efficient fine-tuning model.
Outcome: The proposed approach exceeds the performance of full-parameter fine-tuning and PEFT and provides insights into the analysis of neurons.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations