Challenge: Prior work has shown that language models can be tuned to follow user instructions using only a small set of high-quality instructions.
Approach: They analyze popular selection strategies across different datasets and benchmarks to find out whether they generalize poorly.
Outcome: The proposed methods outperform random baselines and cost-performance trade-offs on the full dataset and a random subset.

Similar Papers

Demystifying Instruction Mixing for Fine-tuning Large Language Models (2024.acl-srw)

Copied to clipboard

Challenge: Instruction tuning is effective for aligning large language models with human instructions, but the procedure to optimizing the mixing of instruction datasets is still unclear.
Approach: They categorize instructions into three primary types: NLP downstream tasks, coding, and general chat.
Outcome: The proposed method improves performance of large language models (LLMs) but it is difficult to combine different instruction datasets to optimize overall performance.
Parameter-Efficient Fine-Tuning: Is There An Optimal Subset of Parameters to Tune? (2024.findings-eacl)

Copied to clipboard

Challenge: Recent research has illuminated the possibility of selective parameter-efficient fine-tuning, which retains the inference speed of the original model and comes at no additional computational cost.
Approach: They propose to selectively update only a small subset of parameters during the fine-tuning process, keeping the remaining parameters frozen during training.
Outcome: The proposed methods retain the inference speed of the original model and come at no additional computational cost.
Unveiling the Generalization Power of Fine-Tuned Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated exceptional multitasking abilities, but the comprehensive effects of fine-tuning on the LLMs’ generalization ability are not fully understood.
Approach: They conduct extensive experiments across five distinct language tasks on different datasets to investigate whether fine-tuning affects the generalization ability intrinsic to LLMs.
Outcome: The proposed model can generalize to different domains and tasks by integrating the in-context learning strategy during fine-tuning on generation tasks.
Rethinking Data Selection at Scale: Random Selection is Almost All You Need (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing data selection techniques are designed for small data pools, a study finds . filtering data by token length is an efficient method for improving results .
Approach: They use self-scoring methods that do not rely on external help to perform fine-tuning . they also find that filtering data by token length offers a stable and efficient method .
Outcome: The proposed methods outperform random selection on large datasets on large data pools.
Take the essence and discard the dross: A Rethinking on Data Selection for Fine-Tuning Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Existing studies focus on data selection but lack a clear, unified framework . variability in experimental settings complicates systematic comparisons .
Approach: They propose a three-stage scheme to standardize data selection for fine-tuning large language models . they propose unified comparison approach that incorporates ratio-based efficiency and ranking-based feasibility metrics to address inconsistencies across experiments.
Outcome: The proposed scheme outperforms existing methods in a dozen key studies and identifies key challenges.
Do Models Really Learn to Follow Instructions? An Empirical Study of Instruction Tuning (2023.acl-short)

Copied to clipboard

Challenge: Recent studies on instruction tuning (IT) have achieved great performance with zero-shot generalizability to unseen tasks.
Approach: They analyze how models utilize instructions during IT by comparing model training with altered vs. original instructions.
Outcome: The proposed model outperforms naive models in low resource setting.
Smaller Language Models are capable of selecting Instruction-Tuning Training Data for Larger Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Instruction tuning language models can be expensive and expensive to train . current methods require extensive training on large datasets, resulting in high training costs.
Approach: They propose a novel approach to selecting training data based on the learning percentage of the samples.
Outcome: The proposed model performs better on models ranging from 1B to 13B in size compared to training on the entire dataset.
Instruction Matters: A Simple yet Effective Task Selection for Optimized Instruction Tuning of Specific Tasks (2024.emnlp-main)

Copied to clipboard

Challenge: Experimental results show that instruction tuning improves zero-shot generalization across various tasks and improves performance of specific tasks.
Approach: They propose a task selection method that leverages instruction information alone to identify relevant tasks and optimize instruction tuning for specific tasks.
Outcome: The proposed method is significantly more efficient than traditional approaches, which require complex measurements of pairwise transferability between tasks or the creation of data samples for the target task.
Fine-Tuning Large Language Models with Sequential Instructions (2025.naacl-long)

Copied to clipboard

Challenge: Existing instruction-tuned models struggle to adhere to a query with multiple intentions, which impairs their performance when the completion of several tasks is demanded by a single command.
Approach: They develop an automatic process that turns existing data into diverse and complex task chains and a new benchmark to evaluate a model’s ability to follow all the instructions in a sequence.
Outcome: The proposed model can follow instructions better and deliver higher results in coding, maths, and open-ended generation.
From Selection to Refinement: Iterative Optimization for Instruction Data (2026.acl-long)

Copied to clipboard

Challenge: Existing methods to optimize instruction tuning datasets face two main challenges: unreasonable pruning of potentially valuable low-quality data and the persistence of noise or semantic drift during revision.
Approach: They propose an automated iterative framework for instruction data optimization that prunes low-quality data and refines low quality data using feedback-driven iteration.
Outcome: The proposed framework outperforms state-of-the-art methods on seven public benchmark datasets with high data efficiency.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations