Papers by Wenkai Yu
Multistage Fusion with Forget Gate for Multimodal Summarization in Open-Domain Videos (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for multimodal summarization for open-domain videos lack fine-grained interactions between multisource inputs. |
| Approach: | They propose a multistage fusion network with a forget gate module to integrate multimodal information into a fluent textual summary. |
| Outcome: | The proposed model achieves state-of-the-art on multiple encoder-decoder architectures and low noise transcripts. |
Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding (2026.acl-long)
Copied to clipboard
| Challenge: | Graphical User Interface (GUI) grounding requires mapping natural language instructions to precise pixel coordinates due to visually homogeneous elements and dense layouts. |
| Approach: | They propose to replace static consistency strategies with a learnable selection mechanism that selects the optimal target by critiquing its own proposals rendered on the screenshot. |
| Outcome: | The proposed model significantly improves both grounding and critiquing capabilities over 6 benchmarks. |
AgenticRAGTracer: A Hop-Aware Benchmark for Diagnosing Multi-Step Retrieval Reasoning in Agentic RAG (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks provide only final questions and answers, while lacking intermediate hop-level questions that gradually connect atomic questions to the final multi-hop query. |
| Approach: | They propose to build a multi-hop reasoning model that is primarily constructed automatically by large language models and designed to support step-by-step validation. |
| Outcome: | The proposed benchmark spans multiple domains, contains 1,305 data points, and has no overlap with existing mainstream benchmarks. |
OpenRLHF: A Ray-based Easy-to-use, Scalable and High-performance RLHF Framework (2025.emnlp-demos)
Copied to clipboard
Jian Hu, Xibin Wu, Wei Shen, Jason Klein Liu, Weixun Wang, Songlin Jiang, Haoran Wang, Hao Chen, Bin Chen, Wenkai Fang, null Xianyu, Yu Cao, Haotian Xu, Yiming Liu
| Challenge: | Existing RLHF frameworks face inference bottlenecks and complexity barriers restricting their accessibility for newcomers. |
| Approach: | They propose an open-source RLHF framework that can be used to train large language models. |
| Outcome: | The proposed framework achieves superior training efficiency with speedups ranging from 1.22 to 1.68 across different model sizes compared to state-of-the-art frameworks, while requiring significantly fewer lines of code for implementation. |