Papers by Wentao Wu

24 papers
Direct Multi-Turn Preference Optimization for Language Agents (2024.emnlp-main)

Copied to clipboard

Challenge: Extensive experiments on three multi-turn agent task datasets confirm the effectiveness and superiority of the DMPO loss function.
Approach: They propose a novel loss function for multi-turn agent tasks that replaces the policy constraint with the state-action occupancy measure constraint and adds length normalization to the Bradley-Terry model.
Outcome: Experiments on three multi-turn agent task datasets confirm the effectiveness and superiority of the proposed loss function.
ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch (2026.acl-long)

Copied to clipboard

Challenge: Existing open-source vision language models lack high-quality training data for chart reasoning . current models are simplistic and repetitive, while associated QA pairs are prone to hallucinations .
Approach: They propose a framework to synthesize complex charts and reliable reasoning data from scratch.
Outcome: Experimental results show that ChartVerse-8B surpasses existing models in QA and difficulty . lack of high-quality training data hampers development of open-source models .
Rethinking Text-to-SQL: Dynamic Multi-turn SQL Interaction for Real-world Database Exploration (2026.findings-acl)

Copied to clipboard

Challenge: Structured Query Language (SQL) is the cornerstone for data-driven decision-making.
Approach: They propose a benchmark to rigorously evaluate Large Language Models within a dynamic interaction framework.
Outcome: The proposed benchmark aims to rigorously evaluate LLMs within a dynamic interaction framework.
Towards General Agentic Intelligence via Environment Scaling (2026.findings-acl)

Copied to clipboard

Challenge: Diverse real-world APIs require precise, robust function-calling intelligence, which needs agents to develop these capabilities through interaction in varied environments.
Approach: They propose a framework that scales up environments to enable agentic intelligence . they use a two-phase agent fine-tuning strategy to first endow agents with basic agentic capabilities, then specializing them for domain-specific contexts.
Outcome: Experiments on -bench, -Bench, and ACEBench show that the model significantly enhances the models’ function-calling capability.
AGSC: Adaptive Granularity and Semantic Clustering for Uncertainty Quantification in Long-text Generation (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for aggregating large-form outputs overlook the nuance of neutral information and suffer from the high computational cost of fine-grained decomposition.
Approach: They propose a UQ framework that uses NLI neutral probabilities as triggers to distinguish irrelevance from uncertainty, reducing computation costs.
Outcome: Experiments on BIO and LongFact show that the proposed framework reduces inference time by 60% compared to full atomic decomposition.
Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration (2025.acl-long)

Copied to clipboard

Challenge: Efficient data selection is crucial to accelerate the pretraining of language models . limited research has addressed the inherent conflicts between data selection methods .
Approach: They propose a multi-actor collaborative data selection mechanism that prioritizes data based on its specific criterion and updates prioritization rules using the current state of the model.
Outcome: The proposed model accelerates convergence in LM pretraining and achieves an average relative performance gain of 10.5% across multiple language model benchmarks.
RealHiTBench: A Comprehensive Realistic Hierarchical Table Benchmark for Evaluating LLM-Based Table Analysis (2025.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for large language models focus on simple, flat table structures.
Approach: They propose a benchmark to evaluate the performance of both Large Language Models and Multimodal LLMs across a variety of input formats for complex tabular data, including LaTeX, HTML, and PNG.
Outcome: The proposed benchmark evaluates the performance of LLMs and Multimodal LLM models across a variety of input formats for complex tabular data, including LaTeX, HTML, and PNG.
DTCRS: Dynamic Tree Construction for Recursive Summarization (2025.acl-long)

Copied to clipboard

Challenge: Recursive summarization (RAG) is an important method for mitigating large model hallucinations and enhancing answer interpretability.
Approach: They propose a method that dynamically generates summary trees based on document structure and query semantics.
Outcome: The proposed method significantly reduces summary tree construction time and achieves substantial improvements across three QA tasks.
Speech-Text Pre-training for Spoken Dialog Understanding with Explicit Cross-Modal Alignment (2023.acl-long)

Copied to clipboard

Challenge: Existing speech-text pre-training methods are limited to one or two specific tasks, despite their success in speech-language processing tasks.
Approach: They propose a temporal position prediction task to capture the speech-text alignment . they use a textual dialog pre-training task to generalize a response selection task .
Outcome: The proposed model is superior in learning speech-text alignment and multi-turn dialog context.
The Data Frontier for Large Language Models: Selection, Synthesis, and Tools (2026.acl-tutorials)

Copied to clipboard

Challenge: acquiring and curating high-quality training data remains a significant bottleneck . acquiring such high-quality data is a key challenge for researchers and practitioners .
Approach: This tutorial provides a comprehensive and practical guide to the state-of-the-art in data research directions for LLMs.
Outcome: The tutorial covers methods for curating the most valuable information from vast, noisy datasets and the synthetic data revolution.
MorphoBench: A Benchmark with Difficulty Adaptive to Model Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks designed to evaluate the reasoning capabilities of large models are limited in scope and lack flexibility to adapt difficulty according to evolving reasoning capacities of models.
Approach: They propose a benchmark that incorporates multidisciplinary questions to evaluate the reasoning capabilities of large models and can adjust and update question difficulty based on the reasoning abilities of advanced models.
Outcome: The proposed benchmark incorporates multidisciplinary questions to evaluate the reasoning capabilities of large models and can adjust and update question difficulty based on the reasoning abilities of advanced models.
GCoT-Decoding: Unlocking Deep Reasoning Paths for Universal Question Answering (2026.findings-acl)

Copied to clipboard

Challenge: Recent work on Chain-of-Thought reasoning requires manual prompts to guide the model.
Approach: They propose a general decoding strategy that generates CoT-style reasoning paths without prompts.
Outcome: The proposed method maintains strong performance on fixed and free QA tasks and achieves significant improvements on free qa.
Nested Browser-Use Learning for Agentic Information Seeking (2026.acl-long)

Copied to clipboard

Challenge: Existing information-seeking (IS) agents rely on the web for their information acquisition.
Approach: They propose a browser-action framework that decouples interaction control from page exploration through a nested structure.
Outcome: Empirical results show that NestBrowse offers clear benefits in practice.
EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement Learning (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for strategic reasoning face challenges in adaptability, scalability, and transferring strategies to new contexts.
Approach: They propose an explicit policy optimization model that provides strategies in open-ended action space and can be plugged into arbitrary LLM agents to motivate goal-directed behavior.
Outcome: The proposed model provides strategies in open-ended action space and can be plugged into arbitrary LLM agents to motivate goal-directed behavior.
Multimodal Chemical Structure-Text Coreference in Intellectual Property via Rule-guided Reinforcement Learning (2026.findings-acl)

Copied to clipboard

Challenge: Existing tools for identifying chemical structures and textual referents are inadequate for this multimodal task.
Approach: They propose a RULE-guided multimodal Reinforcement learning framework for chemical structure-text coreference . RULER is a rule-driven reinforcement learning framework that uses rule-based reward functions to obtain the correct domain knowledge.
Outcome: The proposed framework improves on the baseline framework and shows superior efficacy.
SDPO: Segment-Level Direct Preference Optimization for Social Agents (2025.acl-long)

Copied to clipboard

Challenge: Direct Preference Optimization (DPO) has proven effective in aligning LLM behavior with human preferences across various tasks, but is limited in multi-turn social interactions.
Approach: They propose a method which dynamically selects key segments within interactions to optimize multi-turn agent behavior.
Outcome: The proposed methods outperform existing methods and proprietary LLMs on the SOTOPIA benchmark and show that they can improve social intelligence.
CoMave: Contrastive Pre-training with Multi-scale Masking for Attribute Value Extraction (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to extract product features from unstructured text still suffer from problems . e-commerce platforms are focusing on multi-scale values, which can be confusing .
Approach: They propose a pre-training technique to automatically obtain attribute value pairs from product descriptions to aid e-commerce.
Outcome: The proposed method improves on the existing token-level masking strategy and achieves state-of-the-art on four benchmarks.
CRAB: A Benchmark for Evaluating Curation of Retrieval-Augmented LLMs in Biomedicine (2025.emnlp-industry)

Copied to clipboard

Challenge: Recent development in Retrieval-Augmented Large Language Models (LLMs) have shown great promise in biomedical applications.
Approach: They propose a multilingual benchmark to evaluate retrieval-augmented large language models' curation ability.
Outcome: The proposed benchmark is available in English, French, German and Chinese.
Self-Explanation Prompting Improves Dialogue Understanding in Large Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have achieved great success in various NLP tasks, but the vast model parameters pose challenges in downstream fine-tuning.
Approach: They propose a task-agnostic prompting strategy that analyzes each dialogue utterance before task execution to enhance LLMs' comprehension in multi-turn dialogues.
Outcome: The proposed strategy outperforms other zero-shot prompts and matches or exceeds efficacy of few-shot ones.
Trust Within? Seek Beyond? Knowledge Boundary Aware Policy Optimization for Agentic Search (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to augment large language models with external knowledge suffer from a lack of calibration regarding the model’s knowledge boundary.
Approach: They propose a reinforcement learning framework that explicitly aligns retrieval decisions with quantified knowledge states.
Outcome: The proposed framework outperforms strong baselines while exhibiting reduced hallucination rates.
The Dominance of Text Space: Unveiling the Asymmetric Nature of Cross-Modal Alignment in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for cross-modal alignment assume a symmetric interaction between visual and textual modalities, implying that both spaces adapt to each other.
Approach: They propose a method that regularizes the projector to maintain the geometric structure of the text embedding space via spectral filtering.
Outcome: The proposed method preserves the LLM’s inherent linguistic capabilities and reduces object hallucination significantly better than standard fine-tuning methods.
FlowBench: Revisiting and Benchmarking Workflow-Guided Planning for LLM-based Agents (2024.findings-emnlp)

Copied to clipboard

Challenge: LLM-based agents are susceptible to undesired planning hallucinations when lacking specific knowledge for expertise-intensive tasks.
Approach: They propose a benchmark to evaluate the efficacy of workflow-guided agent planning by formalizing different formats of workflow knowledge.
Outcome: The proposed benchmark aims to improve the planning reliability of LLM-based agents by incorporating external workflow-related knowledge.
UniPCM: Universal Pre-trained Conversation Model with Task-aware Automatic Prompt (2024.lrec-main)

Copied to clipboard

Challenge: Recent studies have shown that multi-task instruction tuning after pre-training greatly improves the model’s robustness and transfer ability, which is crucial for building a high-quality dialog system.
Approach: They propose to use Task-aware Automatic Prompt generation (TAP) to automatically generate high-quality prompts from 15 dialog-related tasks.
Outcome: The proposed model is robust to input prompts and capable of various dialog-related tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations