Papers by Saket Reddy
BiasGRPO: Stabilizing Bias Mitigation in High-Variance Reward Landscapes via Group-Relative Policy Optimization (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent preference-based fine-tuning methods have limited exploration in offline training . previous methods have been limited by the lack of exploration inherent in offline learning . |
| Approach: | They propose a method that normalizes rewards across a group of completed tasks to mitigate social bias in Large Language Models. |
| Outcome: | The proposed approach outperforms DPO and PPO in multiple benchmarks . it can overcome limitations of previous preference-based methods . |