Papers by Xiaoteng Zhang
Graph-GRPO: Stabilizing Multi-Agent Topology Learning via Group Relative Policy Optimization (2026.findings-acl)
Copied to clipboard
Yueyang Cang, Xiaoteng Zhang, Erlu Zhao, Zehua Ji, Yuhang Liu, Yuchen He, Zhiyuan Ning, Chen Yijun, Wenge Que, Li Shi
| Challenge: | Recent approaches to optimize communication topology rely on single-sample policy gradients with absolute rewards. |
| Approach: | They propose a topology optimization framework that integrates Group Relative Policy Optimization. |
| Outcome: | The proposed topology optimization framework outperforms state-of-the-art methods on reasoning and code generation benchmarks. |
From Word to World: Can Large Language Models be Implicit Text-based World Models? (2026.acl-long)
Copied to clipboard
Yixia Li, Hongru Wang, Jiahao Qiu, Zhenfei Yin, Dongdong Zhang, Cheng Qian, Zeping Li, Xiaoteng Ma, Guanhua Chen, Heng Ji
| Challenge: | Agentic learning increasingly hinges on interaction, yet real-world experience is expensive, limited, and often irreversible at inference time. |
| Approach: | They propose a framework that reframes language modeling as next-state prediction under interaction. |
| Outcome: | The proposed framework evaluates world models in text-based environments . it shows that sufficiently trained models capture coherent environment dynamics . |
Mistake Notebook Learning: Batch-Clustered Failures for Training-Free Agent Adaptation (2026.findings-acl)
Copied to clipboard
| Challenge: | Mistake Notebook Learning (MNL) is a new memory framework for large language model agents . it allows agents to distill shared error patterns into structured "mistake notes" |
| Approach: | They propose a new memory framework that enables agents to self-curate generalizable guidance from batch-clustered failures. |
| Outcome: | The proposed framework achieves competitive performance compared to existing memory mechanisms. |
DINT Transformer (2025.emnlp-main)
Copied to clipboard
| Challenge: | Experimental results show that the DINT Transformer improves accuracy and robustness across practical applications. |
| Approach: | They propose a differential attention mechanism that suppresses the impact of irrelevant contexts by computing DIF-Ference between two independent attention distributions. |
| Outcome: | The proposed architecture improves numerical stability and ability to capture global dependencies. |
TEXTOIR: An Integrated and Visualized Platform for Text Open Intent Recognition (2021.acl-demo)
Copied to clipboard
| Challenge: | TEXTOIR is the first integrated platform for text open intent recognition . currently, many dialogue systems are limited to handle the uncertain open intents . |
| Approach: | TEXTOIR is the first integrated platform for text open intent recognition . it is composed of two main modules: open intent detection and open intent discovery . authors propose a framework to implement a complete process to identify known intents and discover open intents . |
| Outcome: | TEXTOIR is the first integrated and visualized platform for text open intent recognition . it integrates state-of-the-art algorithms and benchmark intent datasets . however, there are still some issues, which bring difficulties for future research . |