Papers by Daiki Matsunaga

1 papers
GDPO: Learning to Directly Align Language Models with Diversity Using GFlowNets (2024.emnlp-main)

Copied to clipboard

Challenge: Reinforcement learning with human feedback (RLHF) and its offline variant Direct Preference Optimization (DPO) are two of the most important methods for language model (LM) alignment.
Approach: They propose to use a diversity-seeking RL algorithm called GFlowNet-DPO in an offline preference alignment setting to optimize a model's behavior.
Outcome: Empirical results show that the proposed algorithm generates far more diverse responses than the baseline methods and is still relatively aligned with human values in dialog generation and summarization tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations