Papers by Mohammad Azar

1 papers
Contrastive Policy Gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashion (2024.emnlp-main)

Copied to clipboard

Challenge: Reinforcement Learning (RL) is a method used to fine tune Large Language Models (LLMs) using a reward model trained from preference data to better align with human judgment.
Approach: They propose a Reinforcement Learning (RL) algorithm that can estimate the optimal policy even from off-policy data.
Outcome: The proposed algorithm can estimate the optimal policy even from off-policy data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations