Papers by Pengcheng Wen
Omni-RewardBench: Toward a Comprehensive Evaluation of Generative Reward Models Across Modalities (2026.acl-long)
Copied to clipboard
Chi-Min Chan, Yujin Zhou, Pengcheng Wen, Boqin Yin, Jiaming Ji, Juntao Dai, Wei Xue, Sirui Han, Yike Guo
| Challenge: | Existing evaluation benchmarks for ORMs are largely text-centric or limited to bimodal tasks . a new study examines the effectiveness of Omni-RewardBench for ORms across modalities . |
| Approach: | They propose a hybrid automatic-annotation and human-verification pipeline to construct high-quality evaluation data. |
| Outcome: | The proposed model is the first benchmark for comprehensive evaluation of ORMs across modalities. |
SafeMT: Multi-turn Safety for Multimodal Language Models (2026.acl-long)
Copied to clipboard
Han Zhu, Juntao Dai, Jiaming Ji, Haoran Li, Chengkun Cai, Pengcheng Wen, Chi-Min Chan, Boyuan Chen, Yaodong Yang, Sirui Han, Yike Guo
| Challenge: | Multi-turn dialogues pose a greater risk than single prompts, but existing safety benchmarks do not account for this situation. |
| Approach: | They propose a benchmark that features dialogues of varying lengths generated from harmful queries accompanied by images. |
| Outcome: | The proposed model reduces multi-turn Attack Success Rate (ASR) compared to existing guard models. |
Glance-or-Gaze: Incentivizing LMMs to Adaptively Focus Search via Reinforcement Learning (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing search-augmented approaches rely on indiscriminate whole-image retrieval and lack deep iterative reflection, limiting their effectiveness on complex visual queries. |
| Approach: | They propose a fully autonomous framework that shifts from passive perception to active visual planning and introduces a Selective Gaze mechanism that dynamically chooses whether to glance at global context or gaze into high-value regions. |
| Outcome: | Experiments across six benchmarks demonstrate state-of-the-art performance. |
Boosting Policy and Process Reward Models with Monte Carlo Tree Search in Open-Domain QA (2025.findings-acl)
Copied to clipboard
Chi-Min Chan, Chunpu Xu, Junqi Zhu, Jiaming Ji, Donghai Hong, Pengcheng Wen, Chunyang Jiang, Zhen Ye, Yaodong Yang, Wei Xue, Sirui Han, Yike Guo
| Challenge: | Experimental results show that our approach can effectively improve the performance of both the policy model and the reward model. |
| Approach: | They propose to use Monte Carlo Tree Search for both policy model improvement and reward model improvement to bridge it to more subtle open-domain question answering. |
| Outcome: | The proposed approach surpasses existing methods for annotation and training data with fewer data points and achieves better performance in test-time scaling strategies. |
Natural Language to Code Generation in Interactive Data Science Notebooks (2023.acl-long)
Copied to clipboard
Pengcheng Yin, Wen-Ding Li, Kefan Xiao, Abhishek Rao, Yeming Wen, Kensen Shi, Joshua Howland, Paige Bailey, Michele Catasta, Henryk Michalewski, Oleksandr Polozov, Charles Sutton
| Challenge: | Data scientists use computational notebooks to perform data wrangling and analytic tasks. |
| Approach: | They build a benchmark program that synthesizes programs given NL intents from users by using a Python code language model. |
| Outcome: | The proposed model outperforms public code LMs in a dataset of 1078 code generation problems using the pandas data analysis framework in data science notebooks. |
Personalized Abstractive Summarization by Tri-agent Generation Pipeline (2024.findings-eacl)
Copied to clipboard
| Challenge: | Existing research shows that large language models do not consistently satisfy users' preferences or expectations. |
| Approach: | They propose a tri-agent generation pipeline that includes a generator, an instructor, and an editor to enhance output personalization. |
| Outcome: | The proposed pipeline generates outputs that better meet user expectations on two abstractive summarization datasets. |