Papers by Yali Du
Generalization in Text-based Games via Hierarchical Reinforcement Learning (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Reinforcement Learning (RL) based agents are promising for text-based games, but their generalization remains a challenge. |
| Approach: | They propose a hierarchical framework for reinforcement learning based on knowledge graphs . they propose to decompose the game into subtasks and execute a sub-policy in the low level to conduct goal-conditioned reinforcement learning. |
| Outcome: | The proposed framework enjoys favorable generalizability on a set of difficulty levels and is able to handle complex training tasks. |
CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation (2025.emnlp-main)
Copied to clipboard
| Challenge: | Prior implicit CoT methods have underperformed in terms of efficiency and robustness by relying on natural language tokens for reasoning. |
| Approach: | They propose a training framework that compresses natural language CoT into continuous space by aligning hidden states of a designated token. |
| Outcome: | The proposed framework outperforms the existing state-of-the-art in 3.1x compression rate and 28.2% accuracy on GSM8k scale. |
Spiral of Silence in Large Language Model Agents (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing theories of Spiral of Silence do not apply to large language models . |
| Approach: | They propose an evaluation framework for examining SoS in large language models . they consider four controlled conditions that vary the availability of "History" and "Persona" signals . |
| Outcome: | The proposed framework examines the SoS-like dynamics in large language models . it shows that history and persona together produce strong majority dominance . |
NOVER: Incentive Training for Language Models via Verifier-Free Reinforcement Learning (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in reinforcement learning, such as DeepSeek R1-Zero, highlight the effectiveness of incentive training, but these methods rely on external verifiers, which limits their applicability to domains like mathematics and coding, where such verifier is readily available. |
| Approach: | They propose a general reinforcement learning framework that requires only standard supervised fine-tuning data with no need for an external verifier. |
| Outcome: | The proposed framework outperforms the model of the same size distilled from large reasoning models such as DeepSeek R1 671B by 7.7%. |
Perceiving the World: Question-guided Reinforcement Learning for Text-based Games (2022.acl-long)
Copied to clipboard
| Challenge: | Text-based games provide an interactive way to study natural language processing. |
| Approach: | They propose a two-phase training framework to decouple language learning from reinforcement learning and improve the sample efficiency. |
| Outcome: | The proposed method significantly improves performance and sample efficiency against compound error and limited pre-training data. |
ATLAS: Agent Tuning via Learning Critical Steps (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing agent tuning approaches employ supervised finetuning on entire expert trajectories, but behavior-cloning of full traitories introduces expert bias and weakens generalization to states not covered by the expert data. |
| Approach: | They propose a method that finetunes LLMs on critical steps in expert trajectories and identifies and finetuns them on these steps with reduced costs. |
| Outcome: | The proposed method outperforms existing methods and open-source LLM agents on only 30% critical steps in extensive experiments. |
Entropy-Aware Reshaping of Reinforcement Signals for Multi-Answer Reasoning (2026.findings-acl)
Copied to clipboard
| Challenge: | Reinforcement learning with verifiable rewards (RLVR) is a standard post-training paradigm for large language models. |
| Approach: | They propose a framework that reshapes how learning signals are normalized and aggregated. |
| Outcome: | Experiments on MCTACO and MMLU-Multi show that the proposed framework improves accuracy, training stability and cross-dataset transfer performance. |
VLP: Vision-Language Preference Learning for Embodied Manipulation (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to reward engineering are time-consuming and expensive to collect human preference labels. |
| Approach: | They propose a vision-language preference learning framework which learns from human feedback . they define three types of language-conditioned preferences and construct a visual preference dataset . |
| Outcome: | The proposed framework outperforms baselines on embodied manipulation tasks and can be applied to other tasks. |