Papers with RLTR

    1 papers
    Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning (2025.emnlp-industry)

    Copied to clipboard

    Challenge: Currently, the dominant end-to-end reinforcement learning paradigm for agents in Large Language Models (LLMs) employs multi-objective optimization that jointly trains both planning and answer summarization capabilities.
    Approach: They propose a framework that decouples the training process to enable a focused, single-objective optimization of the planning module.
    Outcome: The proposed framework achieves an 8%–12% improvement in planning performance compared to end-to-end baselines.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations