SteerLM: Attribute Conditioned SFT as an (User-Steerable) Alternative to RLHF (2023.findings-emnlp)
Copied to clipboard
| Challenge: | supervised fine-tuning and reinforcement learning from human feedback (RLHF) are not effective in generating useful and high-quality responses. |
| Approach: | They propose a supervised fine-tuning method that empowers end-users to control responses during inference. |
| Outcome: | Experiments show that supervised fine-tuning and reinforcement learning from human feedback (RLHF) can generate helpful and high-quality responses while maintaining customizability. |
Similar Papers
Reinforcement Learning with Supervised Alignment (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Supervised fine-tuning (SFT) is a widely used method for adapting Large Language Models to specific tasks. |
| Approach: | They propose a method that uses supervised fine-tuning to train a reward model for reinforcement learning. |
| Outcome: | The proposed method outperforms existing methods on in-domain benchmarks but surpasses them 50 times on out-of-domain and cross-task evaluations. |
Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Reinforcement learning from human feedback (RLHF) is the primary method for aligning large language models with human preferences. |
| Approach: | They propose to train an Absolute-Rating Multi-Objective Reward Model with multi-dimensional absolute-rating data. |
| Outcome: | The proposed model outperforms the LLM-as-a-judge method on RewardBench . it achieves state-of-the-art performance on the benchmark . |
DEFT: Distribution-guided Efficient Fine-Tuning for Human Alignment (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Experimental results show that the methods enhanced by DEFT outperform the original methods in both alignment capability and generalization ability, with significantly reduced training time. |
| Approach: | They propose a distribution-based alignment framework that integrates data filtering and distributional guidance to improve alignment efficiency and generalization ability. |
| Outcome: | The proposed framework outperforms existing methods in alignment capability and generalization ability with significantly reduced training time. |
Aligning Large Language Models via Fully Self-Synthetic Data (2026.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to reinforcement learning from human feedback (RLHF) require expensive human-annotated datasets and proprietary models like GPT-4 to annotate preference pairs. |
| Approach: | They propose a self-synthetic framework for LLM alignment where all training data, including prompts (i.e., user queries), responses, and preferences, are generated by the model itself. |
| Outcome: | The proposed framework enhances the model’s chat capabilities on standard benchmarks like AlpacaEval 2.0 while maintaining strong performance on downstream objective tasks. |
LIRE: listwise reward enhancement for preference alignment (2024.findings-acl)
Copied to clipboard
| Challenge: | prevailing approaches to preference alignment focus on pairwise comparisons, with limited exploration into multi-response scenarios. |
| Approach: | They propose a listwise reward enhancement approach that integrates offline rewards of multiple responses into a streamlined listwise framework. |
| Outcome: | The proposed approach outperforms existing methods on dialogue and summarization tasks with good transferability to out-of-distribution data. |
CLHA: A Simple Yet Effective Contrastive Learning Framework for Human Alignment (2024.lrec-main)
Copied to clipboard
Feiteng Fang, Liang Zhu, Xi Feng, Jinchang Hou, Qixuan Zhao, Chengming Li, Xiping Hu, Ruifeng Xu, Min Yang
| Challenge: | Large language models (LLMs) have attracted considerable attention from academic and industrial communities due to their outstanding performance in various natural language processing tasks. |
| Approach: | They propose a Contrastive Learning Framework for Human Alignment to evaluate the noise within the data and dynamically adjust the training process. |
| Outcome: | The proposed framework surpasses other algorithms in terms of reward model scores, automatic evaluations, and human assessments on the widely used dataset "Helpful and Harmless" |
Aligning Large Language Models with Human Preferences through Representation Engineering (2024.acl-long)
Copied to clipboard
Wenhao Liu, Xiaohua Wang, Muling Wu, Tianlong Li, Changze Lv, Zixuan Ling, Zhu JianHao, Cenyuan Zhang, Xiaoqing Zheng, Xuanjing Huang
| Challenge: | Existing methods for achieving this alignment involve employing reinforcement learning from human feedback (RLHF) Existing approaches involve using RLHF to fine-tune LLMs based on human labels . however, RLRF is susceptible to instability during fine- tuning and presents challenges in implementation. |
| Approach: | They propose to use reinforcement learning from human feedback to fine-tune large language models with human preferences to achieve precise control of model behavior. |
| Outcome: | Experiments show that RAHF can be used to capture and manipulate representations to align with a broad spectrum of human preferences or values rather than being confined to a single concept or function. |
trlX: A Framework for Large Scale Reinforcement Learning from Human Feedback (2023.emnlp-main)
Copied to clipboard
Alexander Havrilla, Maksym Zhuravinskyi, Duy Phung, Aman Tiwari, Jonathan Tow, Stella Biderman, Quentin Anthony, Louis Castricato
| Challenge: | Current RLHF paradigms rely on Proximal Policy Optimization (PPO), which quickly becomes a challenge to implement and scale up to large architectures. |
| Approach: | They propose an open-source framework for reinforcement learning from human feedback . it allows for offline fine-tuning of large language models . |
| Outcome: | The framework can be used to fine-tune models up to and exceeding 70 billion parameters. |
Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback (2024.acl-long)
Copied to clipboard
| Challenge: | Recent advances in large language models have influenced the development of video large multimodal models (VLMMs). |
| Approach: | They propose a method that integrates video descriptions as context into a multimodal AI system to enrich the understanding of video content. |
| Outcome: | Empirical evaluations show that the proposed approach outperforms existing approaches for video large multimodal models (VLMMs) |
Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards (2024.acl-long)
Copied to clipboard
| Challenge: | Reinforcement Learning from Human Feedback (RLHF) relies on scalar rewards to capture user preferences. |
| Approach: | They propose a framework that integrates multi-objective reward modeling to represent diverse preference profiles. |
| Outcome: | The proposed method improves performance across reward objectives and targets. |