Clustering and Ranking: Diversity-preserved Instruction Selection through Expert-aligned Quality Estimation (2024.emnlp-main)
Copied to clipboard
Yuan Ge, Yilun Liu, Chi Hu, Weibin Meng, Shimin Tao, Xiaofeng Zhao, Mahong Xia, Zhang Li, Boxing Chen, Hao Yang, Bei Li, Tong Xiao, JingBo Zhu
| Challenge: | Existing methods for instruction data selection have limitations such as relying on fragile external APIs, being affected by biases in GPT models, or reducing the diversity of the selected instruction dataset. |
| Approach: | They propose an industrial-friendly, expert-aligned and diversity-preserved instruction data selection method: Clustering and Ranking (CaR). |
| Outcome: | The proposed method outperforms Alpaca's existing methods by 32.1% in GPT-4 evaluations. |
Similar Papers
Data Diversity Matters for Robust Instruction Tuning (2024.findings-emnlp)
Copied to clipboard
Alexander Bukharin, Shiyang Li, Zhengyang Wang, Jingfeng Yang, Bing Yin, Xian Li, Chao Zhang, Tuo Zhao, Haoming Jiang
| Challenge: | Recent studies have shown that by curating high quality and diverse instruction tuning datasets, we can significantly improve instruction-following capabilities. |
| Approach: | They propose an algorithm to control diversity and quality of instruction tuning datasets and validate it. |
| Outcome: | The proposed algorithm significantly improves worst and average case performance on large scale instruction tuning datasets. |
From Selection to Refinement: Iterative Optimization for Instruction Data (2026.acl-long)
Copied to clipboard
Hang Hu, Ziyan Liu, Rujie Wen, Ruihui Hou, Xueyan Wu, Mu Zhang, Jianxing Yu, Tong Ruan, Jingping Liu
| Challenge: | Existing methods to optimize instruction tuning datasets face two main challenges: unreasonable pruning of potentially valuable low-quality data and the persistence of noise or semantic drift during revision. |
| Approach: | They propose an automated iterative framework for instruction data optimization that prunes low-quality data and refines low quality data using feedback-driven iteration. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on seven public benchmark datasets with high data efficiency. |
RECOST: External Knowledge Guided Data-efficient Instruction Tuning (2024.findings-acl)
Copied to clipboard
| Challenge: | Considering the high computing power overhead, data-efficient instruction tuning is proposed to reduce the training data size. |
| Approach: | They propose a framework to improve instruction tuning by integrating external knowledge into a single pipeline. |
| Outcome: | The proposed method achieves better results with only 1% of the full dataset. |
MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for instruction-tuning datasets prioritize instance quality and use heuristic rules to maintain diversity. |
| Approach: | They propose a method that quantifies diversity based on the distribution of information within a label graph. |
| Outcome: | The proposed method outperforms state-of-the-art methods on 5% Tulu3 datasets and base models. |
MergeIT: From Selection to Merging for Efficient Instruction Tuning (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for instruction tuning rely on LLMs to score instruction quality . existing methods rely only on Llms to rank instruction quality, but this approach is expensive and time-consuming . |
| Approach: | They propose a novel LLM-based Merging strategy for better Instruction Tuning that shifts the focus from selection to synthesis. |
| Outcome: | The proposed method reduces time and computational cost while preserving diversity and reducing redundancy. |
Measuring Data Diversity for Instruction Tuning: A Systematic Analysis and A Reliable Metric (2025.acl-long)
Copied to clipboard
Yuming Yang, Yang Nan, Junjie Ye, Shihan Dou, Xiao Wang, Shuo Li, Huijie Lv, Tao Gui, Qi Zhang, Xuanjing Huang
| Challenge: | Existing studies have explored various diversity-aware data selection methods to construct high-quality datasets and enhance model performance. |
| Approach: | They propose to use data diversity to measure instruction tuning of large language models. |
| Outcome: | The proposed diversity metric outperforms existing methods on simulated and real-world data and shows that it captures diversity variations and achieves a 0.97 correlation with instruction tuning. |
Stratified Selective Sampling for Instruction Tuning with Dedicated Scoring Strategy (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Recent work shows that post-training datasets can be substantially downsampled without noticeably deteriorating performance. |
| Approach: | They propose a method that efficiently bins data into groups and scores difficulty using specialized models. |
| Outcome: | The proposed method can be efficient and universally applied to post-training datasets. |
CommonIT: Commonality-Aware Instruction Tuning for Large Language Models via Data Partitions (2024.emnlp-main)
Copied to clipboard
| Challenge: | Current studies have focused on fine-tuning, but the use of instruction tuning is not as effective as fine-cuning. |
| Approach: | They propose a commonality-aware instruction tuning strategy to cluster instruction datasets into distinct groups with three proposed metrics Task, Embedding and Length. |
| Outcome: | The proposed strategy boosts an average improvement of 2.1% on the general domain and 5.2% on the special domain. |
LimaCost: Data Valuation for Instruction Tuning of Large Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Instruction tuning is an effective approach for aligning large language models with human intentions. |
| Approach: | They propose a data quality measure that exhibits a strong correlation with model performance. |
| Outcome: | The proposed measure exhibits a strong correlation with model performance. |
Optimizing Instruction Synthesis: Effective Exploration of Evolutionary Space with Tree Search (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Extensive research has highlighted the quality of instruction data is essential for the success of this alignment. |
| Approach: | They propose a framework for iteratively improving existing instruction data by using Monte Carlo tree search to find suitable prompts that align the language model to effectively learn multiple skills. |
| Outcome: | The proposed framework improves the evaluation scores of seed instruction data, raising the average evaluation scores from 2.19 to 3.81. |