Papers by Kenshi Abe
Filtered Direct Preference Optimization (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies on the impact of RLHF on text quality have focused on reward-model-free RL. |
| Approach: | They propose an extension of direct preference optimization to improve model performance by analyzing the quality of the preference dataset. |
| Outcome: | The proposed method improves the performance of models optimized with DPO over those optimized with reward-model-based RLHF. |
Regularized Best-of-N Sampling with Minimum Bayes Risk Objective for Language Model Alignment (2025.naacl-long)
Copied to clipboard
| Challenge: | Best-of-N (BoN) sampling is an effective strategy for aligning Large Language Models (LLMs) to human preferences at the time of decoding. |
| Approach: | They propose a variant of BoN that incorporates the Minimum Bayes Risk objective as a proximity regularizer for BoN sampling. |
| Outcome: | The proposed method outperforms both BoN sampling and MBR decoding on the AlpacaFarm and Anthropic datasets. |