BiasGRPO: Stabilizing Bias Mitigation in High-Variance Reward Landscapes via Group-Relative Policy Optimization (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent preference-based fine-tuning methods have limited exploration in offline training . previous methods have been limited by the lack of exploration inherent in offline learning . |
| Approach: | They propose a method that normalizes rewards across a group of completed tasks to mitigate social bias in Large Language Models. |
| Outcome: | The proposed approach outperforms DPO and PPO in multiple benchmarks . it can overcome limitations of previous preference-based methods . |
Similar Papers
MDP-GRPO: Stabilized Group Relative Policy Optimization for Multi-Constraint Instruction Following (2026.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) can follow many natural-language instructions, yet they remain brittle when a request bundles multiple explicit constraints, such as asking the LLM to respond in a particular structure with an exact ending phrase. |
| Approach: | They propose a method which stabilizes learning through multi-temperature sampling to increase reward dispersion, dual-anchor advantages to restore gradients in homogeneous groups, prospect-theoretic shaping to bound updates and penalize violations based on Kahneman Tversky’s theory and asymmetric KL regularization. |
| Outcome: | The proposed method outperforms standard GRPO on FollowBench, IFEval, and a curated multi-constraint dataset, improving strict constraint satisfaction by up to 5.0% on Llama-3.2-3B. |
DISCO Balances the Scales: Adaptive Domain- and Difficulty-Aware Reinforcement Learning on Imbalanced Data (2025.findings-emnlp)
Copied to clipboard
Yuhang Zhou, Jing Zhu, Shengyi Qian, Zhuokai Zhao, Xiyao Wang, Xiaoyu Liu, Ming Li, Paiheng Xu, Wei Ai, Furong Huang
| Challenge: | Large Language Models (LLMs) are increasingly aligned with human preferences through Reinforcement Learning from Human Feedback (RLHF). |
| Approach: | a new study proposes a domain-informed self-consistency policy optimization extension to GRPO that addresses inter-group imbalance. |
| Outcome: | a new extension of GRPO addresses inter-group imbalance with two key innovations . the proposed method outperforms existing GR PO variants by 5% on Qwen3 models . |
Graph-GRPO: Stabilizing Multi-Agent Topology Learning via Group Relative Policy Optimization (2026.findings-acl)
Copied to clipboard
Yueyang Cang, Xiaoteng Zhang, Erlu Zhao, Zehua Ji, Yuhang Liu, Yuchen He, Zhiyuan Ning, Chen Yijun, Wenge Que, Li Shi
| Challenge: | Recent approaches to optimize communication topology rely on single-sample policy gradients with absolute rewards. |
| Approach: | They propose a topology optimization framework that integrates Group Relative Policy Optimization. |
| Outcome: | The proposed topology optimization framework outperforms state-of-the-art methods on reasoning and code generation benchmarks. |
DRA-GRPO: Your GRPO Needs to Know Diverse Reasoning Paths for Mathematical Reasoning (2026.findings-acl)
Copied to clipboard
Xiwen Chen, Wenhui Zhu, Peijie Qiu, Xuanzhao Dong, Hao Wang, Haiyu Wu, Huayu Li, Aris Sotiras, Yalin Wang, Abolfazl Razi
| Challenge: | Existing methods for group-relative policy optimization rely on scalar correctness rewards that are often non-injective with respect to semantic content. |
| Approach: | They propose a framework that calibrates the reward signal using the semantic density of sampled groups. |
| Outcome: | The proposed framework outperforms strong baselines on five math benchmarks with 7,000 samples and 55 cost. |
Mitigating Selection Bias in Large Language Models via Permutation-Aware GRPO (2026.acl-long)
Copied to clipboard
| Challenge: | Existing inference-time debiasing ignores that the same question should yield consistent answers across permutations. |
| Approach: | They propose a permutation-aware group-relative policy optimization which enforces permutations-consistent semantic reasoning. |
| Outcome: | The proposed model outperforms strong baselines across seven benchmarks while maintaining high overall performance. |
Reinforcement Learning for Large Language Models via Group Preference Reward Shaping (2025.emnlp-main)
Copied to clipboard
Huaisheng Zhu, Siyuan Xu, Hangfan Zhang, Teng Xiao, Zhimeng Guo, Shijie Zhou, Shuyue Hu, Vasant G. Honavar
| Challenge: | Existing methods for fine-tuning Large Language Models (LLMs) are expensive and sensitive to reward model quality. |
| Approach: | They propose a method that leverages preference-based comparisons rather than precise numerical rewards. |
| Outcome: | Experiments show that GPRS outperforms critic-model-free RL algorithms on RLHF and reasoning tasks. |
On the Hidden Objective Biases of Group-based Reinforcement Learning (2026.acl-short)
Copied to clipboard
| Challenge: | Recent studies have reported unexpected behaviors during training, including lengthrelated biases, formatting tokens, and reward hacking in multi-objective settings. |
| Approach: | They propose to analyze group-based reinforcement learning methods within a unified surrogate formulation. |
| Outcome: | The proposed methods exhibit structural mismatches between reward optimization and the underlying training objective. |
Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening (2025.emnlp-main)
Copied to clipboard
| Challenge: | Reinforcement learning is emerging as a primary driver for improving language model reasoning capabilities. |
| Approach: | They propose a method for explicitly up-weighting rare but correct solutions to overcome rank bias in group relative policy optimization (GRPO) . |
| Outcome: | The proposed method mitigates rank bias and improves pass@N across a large range of N in both synthetic and real theorem proving settings. |
RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models (2024.findings-naacl)
Copied to clipboard
| Challenge: | Reinforcement learning with human feedback (RLHF) is widely employed to align large language models with user intent. |
| Approach: | They propose to combine rejection sampling and direct preference optimization to improve alignment with user intent by identifying pairs of contrastive samples from human annotator and alternative LLMs. |
| Outcome: | The proposed method outperforms existing methods including RS, PPO, and DPO in a limited resource environment. |
Auto-Weighted Group Relative Preference Optimization for Multi-Objective Text Generation Tasks (2025.emnlp-industry)
Copied to clipboard
| Challenge: | Failing to balance the objectives in advance can lead to overfitting or insufficient learning of each reward function. |
| Approach: | They propose a method that adjusts reward weights according to learning progress . they evaluate AW-GRPO on advertising text generation problem . |
| Outcome: | The proposed method outperforms GRPO on advertising text generation tasks . it overfits BLEURT and jReadability at the expense of BLeurT performance . |