EPO: Hierarchical LLM Agents with Environment Preference Optimization (2024.emnlp-main)
Copied to clipboard
| Challenge: | Long-horizon decision-making tasks require extensive planning over multiple steps, maintaining coherence and goal orientation, which is difficult for LLMs that are typically designed for more immediate and localized predictions. |
| Approach: | They propose a hierarchical framework that decomposes complex tasks into manageable subgoals, utilizing separate LLMs for subgoal prediction and low-level action generation. |
| Outcome: | The proposed framework achieves first place on the ALFRED public leaderboard and demonstrates its potential to improve long-horizon decision-making in diverse environments. |
Similar Papers
MPO: Boosting LLM Agents with Meta Plan Optimization (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for interactive planning tasks suffer from planning hallucinations and require retraining for each new agent. |
| Approach: | They propose a framework that leverages explicit guidance through meta plans to assist agent planning and enables continuous optimization based on feedback from the agent’s task execution. |
| Outcome: | The proposed framework outperforms existing baselines on two representative tasks and significantly improves task completion efficiency and generalization capabilities. |
EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement Learning (2025.acl-long)
Copied to clipboard
Xiaoqian Liu, Ke Wang, Yongbin Li, Yuchuan Wu, Wentao Ma, Aobo Kong, Fei Huang, Jianbin Jiao, Junge Zhang
| Challenge: | Existing methods for strategic reasoning face challenges in adaptability, scalability, and transferring strategies to new contexts. |
| Approach: | They propose an explicit policy optimization model that provides strategies in open-ended action space and can be plugged into arbitrary LLM agents to motivate goal-directed behavior. |
| Outcome: | The proposed model provides strategies in open-ended action space and can be plugged into arbitrary LLM agents to motivate goal-directed behavior. |
HiAgent: Hierarchical Working Memory Management for Solving Long-Horizon Agent Tasks with Large Language Model (2025.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to optimize agent performance by incorporating entire historical action-observation pairs into LLMs are redundant in long-horizon tasks. |
| Approach: | They propose a framework that leverages subgoals as memory chunks to manage working memory of LLM-based agents hierarchically. |
| Outcome: | The proposed framework achieves a twofold increase in success rate and reduces the average number of steps required by 3.8. |
SELFGOAL: Your Language Agents Already Know How to Achieve High-level Goals (2025.naacl-long)
Copied to clipboard
Ruihan Yang, Jiangjie Chen, Yikai Zhang, Siyu Yuan, Aili Chen, Kyle Richardson, Yanghua Xiao, Deqing Yang
| Challenge: | Existing approaches to improve the performance of language agents without training are not available. |
| Approach: | They propose an automatic approach to break down high-level goals into tree structure of more practical subgoals during interaction with environments while identifying the most useful subgoal. |
| Outcome: | The proposed approach significantly improves the performance of language agents across various tasks, including competitive, cooperative, and deferred feedback environments. |
Teaching LLM to be Persuasive: Reward-Enhanced Policy Optimization for Alignment from Heterogeneous Rewards (2026.acl-industry)
Copied to clipboard
| Challenge: | a large language model (LLM) is used as a business development agent for persuasive price negotiation in online travel agencies. |
| Approach: | They propose a reward-enhancing policy optimization method that integrates three complementary reward sources-a preference-trained reward model and an LLM-as-a-judge. |
| Outcome: | The proposed method improves average dialogue rating to 4.63 (+0.33 over GRPO) and raises share of conversations with at least one excellent response to 66.67% (+23.34 pp over grepo). |
SPIO: Ensemble and Selective Strategies via LLM-Based Multi-Agent Planning in Automated Data Science (2026.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have enabled dynamic reasoning in automated data analytics, but rigid, single-path workflows restrict strategic exploration and often lead to suboptimal outcomes. |
| Approach: | a new framework replaces rigid workflows with adaptive, multi-path planning . the framework offers two operating modes: SPIO-S and SPIO -E . |
| Outcome: | a new framework outperforms state-of-the-art pipelines on Kaggle and OpenML benchmarks. |
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models (2025.findings-acl)
Copied to clipboard
Qi Liu, Jingqing Ruan, Hao Li, Haodong Zhao, Desheng Wang, Jiansong Chen, Wan Guanglu, Xunliang Cai, Zhi Zheng, Tong Xu
| Challenge: | Existing multi-objective preference alignment methods for large language models face limitations such as auxiliary reward/reference models and computational complexity. |
| Approach: | They propose a framework that achieves dynamic balance across preference dimensions by using dimension-aware generation metrics as implicit rewards. |
| Outcome: | Empirical results show that AMoPO outperforms state-of-the-art methods by 28.5% . |
CoEx – Co-evolving World-model and Exploration (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing LLM agents fail to assimilate new observations into dynamic updates of the world model, leading to divergent and erroneous plans. |
| Approach: | They propose a hierarchical agent architecture that allows LLM planning to co-evolve with a dynamically updated model of the world. |
| Outcome: | The proposed agent outperforms existing agent paradigms in planning and exploration. |
C-World: A Computer Use Agent Environment Creator (2026.acl-long)
Copied to clipboard
Ziqiao Xi, Shuang Liang, Qi Liu, Jiaqing Zhang, Letian Peng, Fang Nan, Meshal Nayim, Tianhui Zhang, Rishika Mundada, Lianhui Qin, Biwei Huang, Kun Zhou
| Challenge: | C-World enables users to build agent environments on demand. |
| Approach: | They propose a system that enables users to build agent environments on demand. |
| Outcome: | The proposed system outperforms baselines on 119k samples and achieves Spearman = 0.883 ranking correlation with real execution. |
User Feedback Alignment for LLM-powered Exploration in Large-scale Recommendation Systems (2025.acl-industry)
Copied to clipboard
Jianling Wang, Yifan Liu, Yinghao Sun, Xuejian Ma, Yueqi Wang, He Ma, Zhengyang Su, Minmin Chen, Mingyan Gao, Onkar Dalal, Ed H. Chi, Lichan Hong, Ningren Han, Haokai Lu
| Challenge: | Large Language Models (LLMs) can be used to broaden user experiences beyond established preferences and reinforce feedback loops. |
| Approach: | They propose a hierarchical approach that combines hierarchic planning with LLM inference-time scaling to improve recommendation relevancy without compromising novelty. |
| Outcome: | The proposed approach shows significant gains in both user satisfaction and exploration diversity. |