Challenge: Hybrid offline–online reinforcement learning (O2O RL) promises both sample efficiency and robust exploration, but suffers from instability due to distribution shift between offline and online data.
Approach: They propose a framework that decouples policy optimization from safety enforcement . they propose dynamic curricula that gradually extend temporal horizons and anneal offline–online data mixing .
Outcome: The proposed framework preserves the exploratory value of online interactions without collapsing to conservative policies.

Similar Papers

SOLAR-RL: Semi-Online Long-horizon Assignment Reinforcement Learning (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to training GUI agents on dynamic tasks are based on SFT or Behavior Cloning.
Approach: They propose a framework that integrates global trajectory insights directly into offline learning . they reconstruct diverse rollout candidates from static data and detect first failure point .
Outcome: The proposed framework improves long-horizon task completion rates and robustness compared to baselines.
MORE-3S:Multimodal-based Offline Reinforcement Learning with Shared Semantic Spaces (2024.lrec-main)

Copied to clipboard

Challenge: Existing approaches to offline reinforcement learning (RL) focus on learning value functions or policy gradients, but they view it as a sequence modeling task.
Approach: They propose a method that integrates multimodal and pre-trained language models to transform offline reinforcement learning into a supervised learning task by integrating state information derived from images and action-related data obtained from text.
Outcome: The proposed approach outperforms baselines on Atari and OpenAI Gym environments while promoting long-term strategic thinking.
TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to safety alignment of large language models rely on costly manual annotations or human review.
Approach: They propose a closed-loop reinforcement learning framework called TriPlay-RL that enables iterative collaboration among three roles with near-zero manual annotation.
Outcome: The proposed framework achieves 20%–50% improvement in adversarial effectiveness while preserving high output diversity while achieving 10%–30% gains in safety performance without degrading general reasoning capability.
Human-centric dialog training via offline reinforcement learning (2020.emnlp-main)

Copied to clipboard

Challenge: a novel offline RL method can train dialog models to produce better conversations without the risk of humans teaching it harmful chat behaviors.
Approach: They develop offline reinforcement learning algorithms that use human feedback to train dialog models . they use language similarity, laughter, sentiment, and more to identify positive feedback .
Outcome: The proposed method improves on existing methods with 80 users in an open-domain setting.
Scalable and Safe Remediation of Defective Actions in Self-Learning Conversational Systems (2023.acl-industry)

Copied to clipboard

Challenge: Off-Policy reinforcement learning has been used to improve conversational AIs, but in large-scale commercial environments it is challenging to balance between policy improvements and experience continuity.
Approach: They propose to curate and leverage regression incident reports to validate, safe-guard, and improve policies prior to the online deployment.
Outcome: The proposed method validates, safe-guards, and improves policies prior to the online deployment.
trlX: A Framework for Large Scale Reinforcement Learning from Human Feedback (2023.emnlp-main)

Copied to clipboard

Challenge: Current RLHF paradigms rely on Proximal Policy Optimization (PPO), which quickly becomes a challenge to implement and scale up to large architectures.
Approach: They propose an open-source framework for reinforcement learning from human feedback . it allows for offline fine-tuning of large language models .
Outcome: The framework can be used to fine-tune models up to and exceeding 70 billion parameters.
Mobile-R1: Towards Interactive Capability for VLM-Based Mobile Agent via Systematic Training (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to training agents for visual-language models trap them in local optima, hindering exploration and error correction with the environment.
Approach: They propose a hierarchical training recipe that bridges atomic action execution and strategic task completion.
Outcome: The proposed training recipe bridges atomic action execution and strategic task completion.
From log 𝜋 to 𝜋: Taming Divergence in Soft Clipping via Bilateral Decoupled Decay of Probability Gradient Weight (2026.acl-long)

Copied to clipboard

Challenge: Standard algorithms for Large Language Models (LLMs) enforce stability via "hard clipping" but relying on log-probability gradient yields divergent weights as probabilities vanish, destabilizing LLM training.
Approach: They propose a decoupled gradient policy optimization that uses a decay mechanism to decouple the probability of a boundary token.
Outcome: The proposed algorithm outperforms baselines on various mathematical benchmarks.
Targeted Exploration via Unified Entropy Control for Reinforcement Learning (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for group relative policy optimization suffer from entropy collapse . Existing exploration methods introduce additional bias or variance during exploration, making it difficult to maintain stability.
Approach: They propose a framework that provides targeted mechanisms for exploration and stabilization.
Outcome: The proposed framework expands search space on difficult prompts while preventing entropy growth uncontrollably.
Reflect-RL: Two-Player Online RL Fine-Tuning for LMs (2024.acl-long)

Copied to clipboard

Challenge: supervised fine-tuning (SFT) on a limited offline dataset does not yield good performance.
Approach: They propose a two-player system to fine-tune an LM using SFT and online RL . they use negative example generation to enhance error-correction ability of the reflection model .
Outcome: The proposed system outperforms SFT and online RL without reflection on a GPT-2 XL 1.56B model.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations