Papers by Rahul Bhotika
DPL: Diverse Preference Learning Without A Reference Model (2025.naacl-long)
Copied to clipboard
Abhijnan Nath, Andrey Volozin, Saumajit Saha, Albert Aristotle Nanda, Galina Grunin, Rahul Bhotika, Nikhil Krishnaswamy
| Challenge: | Existing methods to direct preference alignment do not utilize diversity in preference annotations which limits their applicability. |
| Approach: | They propose a reference-model-free method that learns a baseline desirability in LLM responses while being robust to the diversity of preference annotations. |
| Outcome: | The proposed method learns a baseline desirability in LLM responses while being robust to the diversity of preference annotations. |