| Challenge: | Existing methods for fine-tuning and reinforcement learning use only positive examples, limiting their efficiency in low-resource scenarios. |
| Approach: | They propose a method that leverages both successful and failed trajectories for fine-tuning, maximizing the utility of limited resources. |
| Outcome: | The proposed method surpasses existing methods, including SFT, DPO, and PPO, across various tasks. |
Similar Papers
Learning from Mistakes: Negative Reasoning Samples Enhance Out-of-Domain Generalization (2026.acl-long)
Copied to clipboard
Tian Xueyun, MingHua Ma, Bingbing Xu, Nuoyan Lyu, Wei Li, Heng Dong, Zheng Chu, Yuanzhuo Wang, Huawei Shen
| Challenge: | Recent studies show that supervised fine-tuning (SFT) is a common approach for reasoning in large language models. |
| Approach: | They propose to use supervised fine-tuning (SFT) on chain-of-thought trajectories demonstrations . they find that incorporating negative traxories yields substantial OOD generalization gains . |
| Outcome: | The proposed scheme yields 5.51% OOD gain over positive-only training. |
ATLAS: Agent Tuning via Learning Critical Steps (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing agent tuning approaches employ supervised finetuning on entire expert trajectories, but behavior-cloning of full traitories introduces expert bias and weakens generalization to states not covered by the expert data. |
| Approach: | They propose a method that finetunes LLMs on critical steps in expert trajectories and identifies and finetuns them on these steps with reduced costs. |
| Outcome: | The proposed method outperforms existing methods and open-source LLM agents on only 30% critical steps in extensive experiments. |
RIFT: Repurposing Negative Samples via Reward-Informed Fine-Tuning (2026.findings-acl)
Copied to clipboard
| Challenge: | Reward Informed Fine-Tuning (RIFT) is an effective and robust alternative to expensive expert data for LLM alignment. |
| Approach: | They propose a reward-informed fine-tuning framework that utilizes all self-generated samples to learn from both positive and negative trajectories. |
| Outcome: | The proposed framework outperforms both RFT and Supervised Fine-Tuning (SFT) on mathematical benchmarks. |
AgentBank: Towards Generalized LLM Agents via Fine-Tuning on 50000+ Interaction Trajectories (2024.findings-emnlp)
Copied to clipboard
Yifan Song, Weimin Xiong, Xiutian Zhao, Dawei Zhu, Wenhao Wu, Ke Wang, Cheng Li, Wei Peng, Sujian Li
| Challenge: | Existing studies focus on specialized agents designed for particular tasks. |
| Approach: | They propose to scale annotated interaction trajectories and fine-tune LLMs on AgentBank to get a series of agent models, Samoyed. |
| Outcome: | The proposed model can scale to get generalized agent capabilities. |
Enhancing the General Agent Capabilities of Low-Paramter LLMs through Tuning and Multi-Branch Reasoning (2024.findings-naacl)
Copied to clipboard
| Challenge: | Open-source pre-trained Large Language Models exhibit strong language understanding and generation capabilities, making them highly successful in a variety of tasks. |
| Approach: | They propose a method to construct agent-specific data using GPT-4 and supervised fine-tuning . they find that supervised tunning can significantly reduce hallucination outputs and formatting errors in agent tasks . |
| Outcome: | The proposed method improves on five agent tasks of AgentBench. |
AgentTuning: Enabling Generalized Agent Abilities for LLMs (2024.findings-acl)
Copied to clipboard
| Challenge: | Open large language models (LLMs) with great performance in various tasks are far inferior to commercial models such as ChatGPT and GPT-4 when acting as agents to tackle complex tasks in the real world. |
| Approach: | They propose a method to enhance the agent capabilities of LLMs while maintaining their general abilities. |
| Outcome: | The AgentLM-70B is comparable to GPT-3.5-turbo on unseen agent tasks, demonstrating generalized agent capabilities. |
CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution (2026.acl-long)
Copied to clipboard
| Challenge: | Extensive experiments on AppWorld and BFCL demonstrate consistent and significant improvements over strong base models, yielding absolute gains of 19.43%, 15.58%, and 18.14%, respectively. |
| Approach: | They propose a framework that extracts feedback signals such as forgetting and uncertainty from rollout trajectories and utilizes them to guide LLM-based task synthesis. |
| Outcome: | Extensive experiments on AppWorld and BFCL show that the proposed framework improves over strong base models. |
Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing studies focus on prompt engineering or framework scheduling of one/multiple LLMs. |
| Approach: | They propose to integrate LLMs as agents into their training corpus by decomposition and redesigning the training corpu . they propose to use LLM-FLAN to effectively fine-tune LANguage models for Agents by reducing hallucinations. |
| Outcome: | The proposed model outperforms prior best models by 3.5% across agent evaluation datasets. |
STeCa: Step-level Trajectory Calibration for LLM Agent Learning (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing work focuses on behavior cloning from expert demonstrations or preference learning through exploratory trajectory sampling, but these methods often struggle to address long-horizon tasks where suboptimal actions accumulate step by step, causing agents to deviate from correct task trajectories. |
| Approach: | They propose a framework for LLM-based agent learning that identifies suboptimal actions through a step-level reward comparison during exploration and constructs calibrated trajectories using LLM reflection. |
| Outcome: | The proposed framework outperforms existing methods in long-horizon tasks where suboptimal actions accumulate step by step, causing agents to deviate from correct task trajectories. |
Reinforcement Learning with Supervised Alignment (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Supervised fine-tuning (SFT) is a widely used method for adapting Large Language Models to specific tasks. |
| Approach: | They propose a method that uses supervised fine-tuning to train a reward model for reinforcement learning. |
| Outcome: | The proposed method outperforms existing methods on in-domain benchmarks but surpasses them 50 times on out-of-domain and cross-task evaluations. |