Papers by Ruiming Tang
NL-Debugging: Exploiting Natural Language as an Intermediate Representation for Code Debugging (2025.emnlp-main)
Copied to clipboard
Weiming Zhang, Qingyao Li, Xinyi Dai, Jizheng Chen, Kounianhua Du, Weiwen Liu, Yasheng Wang, Ruiming Tang, Yong Yu, Weinan Zhang
| Challenge: | Early debugging efforts focused on code-level analysis, which often fails when addressing complex programming errors. |
| Approach: | They propose a framework that employs natural language as an intermediate representation to improve code debugging by debuggating at a natural language level. |
| Outcome: | The proposed framework outperforms traditional debugging methods and enables a broader modification space through direct refinement guided by execution feedback. |
DebateCoder: Towards Collective Intelligence of LLMs via Test Case Driven LLM Debate for Code Generation (2025.acl-long)
Copied to clipboard
Jizheng Chen, Kounianhua Du, Xinyi Dai, Weiming Zhang, Xihuai Wang, Yasheng Wang, Ruiming Tang, Weinan Zhang, Yong Yu
| Challenge: | Existing debate-based approaches to code generation are limited due to several reasons: 1) Reliance on different instances of the same LLM for debate, 2) under-utilization of test cases, and 3) reliance on third-party moderators for result consolidation and decision-making. |
| Approach: | They propose to use test cases to analyze code and identify bugs while opposing models generate test cases for each other to challenge each other's code during the debate process. |
| Outcome: | The proposed model collects intelligence of LLMs via test case-driven debate for code generation. |
DARL: Encouraging Diverse Answers for General Reasoning without Verifiers (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent efforts such as RLPR have extended RLVR to general domains, enabling training on broader datasets and achieving improvements over RL PR. |
| Approach: | They propose a framework that encourages the generation of diverse answers within a controlled deviation range from the reference while preserving alignment with it. |
| Outcome: | Extensive experiments on 13 benchmarks show that DARL surpasses RLPR in both reasoning accuracy and output diversity. |
Bridging Relevance and Reasoning: Rationale Distillation in Retrieval-Augmented Generation (2025.findings-acl)
Copied to clipboard
Pengyue Jia, Derong Xu, Xiaopeng Li, Zhaocheng Du, Xiangyang Li, Yichao Wang, Yuhao Wang, Qidong Liu, Maolin Wang, Huifeng Guo, Ruiming Tang, Xiangyu Zhao
| Challenge: | Existing approaches to rerank and align documents based on reasoning capabilities of large language models (LLMs) . prior work shows that LLMs have exceptional reasoning and text generation capabilities . |
| Approach: | They propose a rationale extraction method that leverages reasoning capabilities of large language models to extract the rationales necessary for answering a query. |
| Outcome: | The proposed method is compared with baseline methods on two tasks across three datasets. |
LLMTreeRec: Unleashing the Power of Large Language Models for Cold-Start Recommendations (2025.coling-main)
Copied to clipboard
Wenlin Zhang, Chuhan Wu, Xiangyang Li, Yuhao Wang, Kuicai Dong, Yichao Wang, Xinyi Dai, Xiangyu Zhao, Huifeng Guo, Ruiming Tang
| Challenge: | Lack of training data leads to the system cold-start problem in recommendation systems, making them struggle to provide effective recommendations. |
| Approach: | They propose a tree-based LLM recommendation framework which structures all items into an item tree to improve the efficiency of LLM’s item retrieval. |
| Outcome: | The proposed framework outperforms the baseline model in the A/B test on Huawei industrial system. |
DeepSynth-Eval: Objectively Evaluating Information Consolidation in Deep Survey Writing (2026.findings-acl)
Copied to clipboard
Hongzhi Zhang, Yuanze Hu, Tinghai Zhang, Jia Fu, Tao Wang, Junwei Jing, Zhaoxin Fan, Wei Bi, Ruiming Tang, Han Li, Guorui Zhou, Kun Gai
| Challenge: | Large Language Models (LLMs) are evolving towards autonomous agents . retrieval capabilities are well-benchmarked, but post-retrieval synthesis is under-evaluated due to open-ended writing. |
| Approach: | They propose a benchmark to evaluate information consolidation capabilities using survey papers as gold standards. |
| Outcome: | The proposed benchmark analyzes the post-retrieval synthesis stage of large language models . it leverages high-quality survey papers as gold standards and reverse-engineers research requests . the proposed benchmark outperforms single-turn generation and reduces hallucinations . |
Humanity’s Last Code Exam: Can Advanced LLMs Conquer Human’s Hardest Code Competition? (2025.findings-emnlp)
Copied to clipboard
Xiangyang Li, Xiaopeng Li, Kuicai Dong, null Zhangquanhu, Rongju Ruan, Xinyi Dai, Yasheng Wang, Ruiming Tang
| Challenge: | o4-mini(high) and Gemini-2.5 Pro achieve pass@1 rates of only 15.9% and 11.4%, respectively. |
| Approach: | They propose a harmonized online–offline sandbox that guarantees fully reproducible evaluation. |
| Outcome: | The proposed test reflects the advanced reasoning and code generation ability of large language models. |
CoIR: A Comprehensive Benchmark for Code Information Retrieval Models (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods and benchmarks for information retrieval are inadequately representing the diversity of code in various domains and tasks. |
| Approach: | They propose a benchmark specifically designed to assess code retrieval capabilities. |
| Outcome: | The proposed benchmark aims to invigorate research in the code retrieval domain . it shares the same data schema as other popular benchmarks like MTEB and BEIR . |
Boost, Disentangle, and Customize: A Robust System2-to-System1 Pipeline for Code Generation (2025.findings-acl)
Copied to clipboard
Kounianhua Du, Hanjing Wang, Jianxing Liu, Jizheng Chen, Xinyi Dai, Yasheng Wang, Ruiming Tang, Yong Yu, Jun Wang, Weinan Zhang
| Challenge: | Existing systems 2 methods for code generation are difficult to implement due to the complex hidden reasoning process and heterogeneous data distribution. |
| Approach: | They propose a framework that Boosts reasoning exploration via multi-agent collaboration and Disentangles heterogeneous data into specialized experts. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on APPS and CodeContest benchmarks and achieves 73.8% accuracy on hard problems. |
Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-Judge (2025.acl-long)
Copied to clipboard
Qiyuan Zhang, Yufei Wang, Yuxin Jiang, Liangyou Li, Chuhan Wu, Yasheng Wang, Xin Jiang, Lifeng Shang, Ruiming Tang, Fuyuan Lyu, Chen Ma
| Challenge: | Existing methods rely on majority voting or criteria expansion to capture detailed and detailed details, often leading to incomplete outcomes. |
| Approach: | They propose a method which introduces additional crowd responses to compare with the candidate responses, thereby exposing deeper and more comprehensive details within the candidate answers. |
| Outcome: | Experiments show that the proposed method improves evaluation reliability and achieves an average gain of 6.7% across five benchmarks. |
RethinkMCTS: Refining Erroneous Thoughts in Monte Carlo Tree Search for Code Generation (2025.emnlp-main)
Copied to clipboard
Qingyao Li, Wei Xia, Xinyi Dai, Kounianhua Du, Weiwen Liu, Yasheng Wang, Ruiming Tang, Yong Yu, Weinan Zhang
| Challenge: | Existing tree search methods neglect the underlying reasoning process, resulting in poor search quality. |
| Approach: | They propose a framework that systematically explores and refines the reasoning process for code generation by using a tree search engine and a reflection mechanism. |
| Outcome: | The proposed framework outperforms existing methods in the code generation domain. |
MMDocIR: Benchmarking Multimodal Retrieval for Long Documents (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing benchmarks for multimodal document retrieval are lacking for evaluating performance of systems. |
| Approach: | They propose a benchmark that evaluates page-level and layout-level retrieval tasks . they use a rich dataset featuring 1,685 questions annotated by experts . |
| Outcome: | The proposed benchmark outperforms existing benchmarks in page-level and layout-level retrieval tasks. |
Learning to Edit: Aligning LLMs with Knowledge Editing (2024.acl-long)
Copied to clipboard
Yuxin Jiang, Yufei Wang, Chuhan Wu, Wanjun Zhong, Xingshan Zeng, Jiahui Gao, Liangyou Li, Xin Jiang, Lifeng Shang, Ruiming Tang, Qun Liu, Wei Wang
| Challenge: | Existing knowledge editing techniques rely on memorizing updated knowledge, impeding LLMs from effectively combining the new knowledge with their inherent knowledge when answering questions. |
| Approach: | They propose a Learning to Edit framework that equips LLMs with the ability to apply updated knowledge to input questions through a two-phase process . |
| Outcome: | The proposed framework outperforms existing methods in knowledge editing tasks and compares it with four benchmarks and two LLM architectures. |
CodePRM: Execution Feedback-enhanced Process Reward Model for Code Generation (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in code generation focus on optimizing the thought process, but lack effective process supervision, making it difficult to optimize the thoughts. |
| Approach: | They propose a method that leverages the code execution feedback to build a code PRM by collecting a large dataset of thought traces and then training it to take both the reasoning process and code execution as input. |
| Outcome: | The proposed approach outperforms baselines and strong LLMs in the inference stage. |
Adaptive Tool Use in Large Language Models with Meta-Cognition Trigger (2025.acl-long)
Copied to clipboard
Wenjun Li, Dexun Li, Kuicai Dong, Cong Zhang, Hao Zhang, Weiwen Liu, Yasheng Wang, Ruiming Tang, Yong Liu
| Challenge: | Existing research expands the tool arrays of large language models (LLMs), but the necessity of using these tools is often overlooked, leading to indiscriminate tool invocation. |
| Approach: | They propose a meta-cognition proxy proxy for LLMs self-assessment of their capabilities, reflecting the model’s awareness of its own limitations. |
| Outcome: | The proposed strategy is fine-tuned-free and costs minimal. |
RecBase: Generative Foundation Model Pretraining for Zero-Shot Recommendation (2025.emnlp-main)
Copied to clipboard
Sashuai Zhou, Weinan Gan, Qijiong Liu, Ke Lei, Jieming Zhu, Hai Huang, Yan Xia, Ruiming Tang, Zhenhua Dong, Zhou Zhao
| Challenge: | Existing methods for addressing item-level user interests are lacking in cross-domain generalization . RecBase model is domain-agnostic and can be used to enhance recommender systems' effectiveness . |
| Approach: | They propose a domain-agnostic foundational model pretrained with a recommendation-oriented objective that leverages a large-scale, heterogeneous, cross-domain corpus with unified textual representations and feature mappings to enhance cross- domain generalization. |
| Outcome: | The proposed model matches or surpasses baselines in zero-shot and cross-domain recommendation tasks on eight real-world datasets. |
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning (2026.findings-acl)
Copied to clipboard
Zhenpeng Su, Leiyu Pan, Minxuan Lv, Tiehua Mei, Zijia Lin, Yuntao Li, Wenping Hu, Ruiming Tang, Kun Gai, Guorui Zhou
| Challenge: | Large language model post-training often adopts an off-policy training paradigm . however, the off-poliicy training model introduces distribution shifts that push the policy beyond the trust region. |
| Approach: | They propose to use the entropy ratio as a global metric to measure the relative change in policy exploration throughout updates. |
| Outcome: | Experiments show that the proposed metric improves performance across multiple benchmarks. |
Unleashing the Native Recommendation Potential: LLM-Based Generative Recommendation via Structured Term Identifiers (2026.findings-acl)
Copied to clipboard
Zhiyang Zhang, Junda She, Kuo Cai, Bo Chen, Shiyao Wang, Xinchen Luo, Qiang Luo, Ruiming Tang, Han Li, Kun Gai, Guorui Zhou
| Challenge: | Existing methods for constructing item identifiers face bottlenecks due to their large output space and expensive vocabulary expansion and alignment training. |
| Approach: | They propose to use Large Language Models to develop general-purpose, semantically-aware recommender systems that can be generalized and reusable. |
| Outcome: | Experiments on real-world datasets show that GRAM outperforms baselines and significantly outperformed baselines. |
Instruction-Tuning Data Synthesis from Scratch via Web Reconstruction (2025.findings-acl)
Copied to clipboard
Yuxin Jiang, Yufei Wang, Chuhan Wu, Xinyi Dai, Yan Xu, Weinan Gan, Yasheng Wang, Xin Jiang, Lifeng Shang, Ruiming Tang, Wei Wang
| Challenge: | Existing methods for generating and curating high-quality instruction-tuning data rely heavily on the quality of seed data or strong assumptions about the structure and content of web documents. |
| Approach: | They propose a fully automated framework for synthesizing high-quality instruction-tuning (IT) data directly from raw web documents with minimal assumptions. |
| Outcome: | The proposed framework outperforms state-of-the-art baselines by 16.65% across four instruction-following benchmarks. |
OneRec-Think: In-Text Reasoning for Generative Recommendation (2026.acl-long)
Copied to clipboard
Zhanyu Liu, Shiyao Wang, Xingmei Wang, Rongzhou Zhang, Jiaxin Deng, Honghui Bao, Jinghao Zhang, Wuchao Li, PengFei Zheng, Xiangyu Wu, Yifei Hu, Qigen Hu, Xinchen Luo, Lejian Ren, Zhang Zixing, Qianqian Wang, Kuo Cai, Yunfan Wu, Hongtao Cheng, Zexuan Cheng, Lu Ren, Huanjie Wang, Yi Su, Ruiming Tang, Kun Gai, Guorui Zhou
| Challenge: | Existing generative models lack the capacity for explicit and controllable reasoning, a key advantage of LLMs. |
| Approach: | They propose a framework that integrates dialogue, reasoning, and personalized recommendation. |
| Outcome: | Experiments across public benchmarks show state-of-the-art performance. |
DPWriter: Reinforcement Learning with Diverse Planning Branching for Creative Writing (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods for enhancing large language models (LLMs) lack explicit mechanisms for guiding diverse exploration and instead prioritize efficiency and performance over diversity. |
| Approach: | They propose a reinforcement learning-based framework that decomposes the generation process into explicitly planned intermediate steps and introduces divergence at the planning phase based on diversity variation. |
| Outcome: | The proposed method significantly outperforms existing baselines on creative writing benchmarks on a semi-structured long chain-of-thought (CoT) it introduces divergence at the planning phase based on diversity variation, alongside a group-aware diversity reward to encourage distinct trajectories. |
An Effective Post-training Embedding Binarization Approach for Fast Online Top-K Passage Matching (2022.aacl-short)
Copied to clipboard
| Challenge: | Existing models that learn semantic representations of passages are prone to performance degradation . embedding binarization is a promising branch of model compression . |
| Approach: | They propose an embedding binarization approach that can be used to optimize for online inference. |
| Outcome: | The proposed model can perform query-passage matching acceleration. |