Papers by Karim Galliamov

1 papers
Enhancing RLHF with Human Gaze Modeling (2025.emnlp-main)

Copied to clipboard

Challenge: Reinforcement Learning from Human Feedback (RLHF) is a powerful paradigm for aligning language models with human values and preferences.
Approach: They propose to use gaze-aware reward models and gaze-based distribution of sparse rewards to enhance RLHF.
Outcome: The proposed models achieve faster convergence while maintaining or slightly improving performance, reducing computational requirements during policy training.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations