Papers by Pengcheng Wen

6 papers
Omni-RewardBench: Toward a Comprehensive Evaluation of Generative Reward Models Across Modalities (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluation benchmarks for ORMs are largely text-centric or limited to bimodal tasks . a new study examines the effectiveness of Omni-RewardBench for ORms across modalities .
Approach: They propose a hybrid automatic-annotation and human-verification pipeline to construct high-quality evaluation data.
Outcome: The proposed model is the first benchmark for comprehensive evaluation of ORMs across modalities.
SafeMT: Multi-turn Safety for Multimodal Language Models (2026.acl-long)

Copied to clipboard

Challenge: Multi-turn dialogues pose a greater risk than single prompts, but existing safety benchmarks do not account for this situation.
Approach: They propose a benchmark that features dialogues of varying lengths generated from harmful queries accompanied by images.
Outcome: The proposed model reduces multi-turn Attack Success Rate (ASR) compared to existing guard models.
Glance-or-Gaze: Incentivizing LMMs to Adaptively Focus Search via Reinforcement Learning (2026.findings-acl)

Copied to clipboard

Challenge: Existing search-augmented approaches rely on indiscriminate whole-image retrieval and lack deep iterative reflection, limiting their effectiveness on complex visual queries.
Approach: They propose a fully autonomous framework that shifts from passive perception to active visual planning and introduces a Selective Gaze mechanism that dynamically chooses whether to glance at global context or gaze into high-value regions.
Outcome: Experiments across six benchmarks demonstrate state-of-the-art performance.
Boosting Policy and Process Reward Models with Monte Carlo Tree Search in Open-Domain QA (2025.findings-acl)

Copied to clipboard

Challenge: Experimental results show that our approach can effectively improve the performance of both the policy model and the reward model.
Approach: They propose to use Monte Carlo Tree Search for both policy model improvement and reward model improvement to bridge it to more subtle open-domain question answering.
Outcome: The proposed approach surpasses existing methods for annotation and training data with fewer data points and achieves better performance in test-time scaling strategies.
Natural Language to Code Generation in Interactive Data Science Notebooks (2023.acl-long)

Copied to clipboard

Challenge: Data scientists use computational notebooks to perform data wrangling and analytic tasks.
Approach: They build a benchmark program that synthesizes programs given NL intents from users by using a Python code language model.
Outcome: The proposed model outperforms public code LMs in a dataset of 1078 code generation problems using the pandas data analysis framework in data science notebooks.
Personalized Abstractive Summarization by Tri-agent Generation Pipeline (2024.findings-eacl)

Copied to clipboard

Challenge: Existing research shows that large language models do not consistently satisfy users' preferences or expectations.
Approach: They propose a tri-agent generation pipeline that includes a generator, an instructor, and an editor to enhance output personalization.
Outcome: The proposed pipeline generates outputs that better meet user expectations on two abstractive summarization datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations