Papers by Rahul Bhotika

1 papers
DPL: Diverse Preference Learning Without A Reference Model (2025.naacl-long)

Copied to clipboard

Challenge: Existing methods to direct preference alignment do not utilize diversity in preference annotations which limits their applicability.
Approach: They propose a reference-model-free method that learns a baseline desirability in LLM responses while being robust to the diversity of preference annotations.
Outcome: The proposed method learns a baseline desirability in LLM responses while being robust to the diversity of preference annotations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations