Papers by Ze Wang

13 papers
Self-Taught Agentic Long Context Understanding (2025.acl-long)

Copied to clipboard

Challenge: Extensive experiments across seven long-context tasks demonstrate that AgenticLU significantly outperforms state-of-the-art prompting methods and specialized long-consumer LLMs.
Approach: They propose a framework to enhance an LLM's understanding of long-context questions by integrating targeted self-clarification with contextual grounding within an agentic workflow.
Outcome: The proposed framework outperforms state-of-the-art prompting methods and specialized long-context LLMs in seven long-constitut tasks.
Do VLMs Have a Moral Backbone? A Study on the Fragile Morality of Vision-Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Vision-Language Models (VLMs) have advanced multimodal learning, driving progress in cross-modal reasoning.
Approach: They propose to examine moral robustness of vision-language models by analyzing their moral stances under multimodal perturbations.
Outcome: The proposed model-agnostic multimodal perturbations expose VLMs to a variety of moral vulnerabilities, including a sycophancy trade-off where stronger instruction-following models are more susceptible to persuasion.
Agent Laboratory: Using LLM Agents as Research Assistants (2025.findings-emnlp)

Copied to clipboard

Challenge: Agent Laboratory is an autonomous LLM-based framework that can complete the entire research process.
Approach: Agent Laboratory is an autonomous LLM-based framework that can complete the entire research process.
Outcome: Agent Laboratory is an autonomous LLM-based framework that can complete the entire research process.
Reliable Use of Lemmas via Eligibility Reasoning and Section-Aware Reinforcement Learning (2026.acl-short)

Copied to clipboard

Challenge: Recent large language models (LLMs) perform strongly on mathematical benchmarks but often import conclusions without validating assumptions.
Approach: They propose a model that encodes a lemma specification and trains with reinforcement learning and section-aware loss masking to assign penalty to the section responsible for errors.
Outcome: The proposed model performs well on benchmarks but often misapplyes lemmas . the model is able to encode the specification and train with reinforcement learning .
MegaPairs: Massive Data Synthesis for Universal Multimodal Retrieval (2025.acl-long)

Copied to clipboard

Challenge: despite the growing demand for multimodal retrieval, there is a lack of training data.
Approach: They propose a data synthesis method that leverages vision language models and open-domain images to generate high-quality data.
Outcome: The proposed method outperforms baseline models on 70 more datasets and can scale up.
FedNLP: Benchmarking Federated Learning Methods for Natural Language Processing Tasks (2022.findings-naacl)

Copied to clipboard

Challenge: Increasing concerns and regulations about data privacy necessitate the study of privacy-preserving, decentralized learning methods for natural language processing tasks.
Approach: They propose a framework for evaluating federated learning methods on four different tasks . they propose federation between Transformer-based language models and FL methods .
Outcome: The proposed framework compares FL methods on four different tasks under non-IID partitioning strategies.
How do Language Models Reshape Entity Alignment? A Survey of LM-Driven EA Methods: Advances, Benchmarks, and Future (2025.emnlp-main)

Copied to clipboard

Challenge: Entity alignment (EA) is critical for knowledge graph (KG) integration.
Approach: They propose a taxonomy that categorizes methods in three stages: data preparation, feature embedding, and alignment.
Outcome: The proposed taxonomy categorizes methods in three key stages: data preparation, feature embedding, and alignment.
LibVulnWatch: A Deep Assessment Agent System and Leaderboard for Uncovering Hidden Vulnerabilities in Open-Source AI Libraries (2025.acl-srw)

Copied to clipboard

Challenge: Open-source AI libraries present significant, underexamined risks spanning security, licensing, maintenance, supply chain integrity, and regulatory compliance.
Approach: They propose a system that leverages large language models and agentic workflows to perform deep, evidence-based evaluations of open-source AI libraries.
Outcome: The proposed system covers up to 88% of OpenSSF Scorecard checks and uncovers 19 additional risks per library.
SAGED: A Holistic Bias-Benchmarking Pipeline for Language Models with Customisable Fairness Calibration (2025.coling-main)

Copied to clipboard

Challenge: Existing benchmarks for large language models fail to detect bias due to limited scope, contamination, and lack of a fairness baseline.
Approach: They propose a benchmarking pipeline to detect biases in large language models . they use metrics for max disparity, impact ratio, and bias concentration to analyze disparity .
Outcome: SAGED(bias) is the first holistic benchmarking pipeline to address biases in large language models.
DRIFT: Transferring Reasoning Priors for Efficient MLLM Fine-Tuning (2026.findings-acl)

Copied to clipboard

Challenge: Multimodal large language models (MLLMs) have made rapid progress in perception and alignment, but their reasoning ability often lags behind strong text-only LLMs.
Approach: They propose a method that transfers reasoning knowledge in the gradient space while preserving multimodal alignment.
Outcome: Experiments on multimodal reasoning benchmarks show that DRIFT outperforms naive merging and standard SFT.
StyleDGPT: Stylized Response Generation with Pre-trained Language Models (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for generating responses following a desired style are lacking of parallel data for training.
Approach: They propose a KL loss and a style classifier to fine-tune response generation . they show that their model can significantly outperform state-of-the-art methods .
Outcome: The proposed model outperforms state-of-the-art models in style consistency and contextual coherence with two public datasets.
JobFair: A Framework for Benchmarking Gender Hiring Bias in Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: a framework for benchmarking hierarchical gender hiring bias in Large Language Models (LLMs) is developed to protect vulnerable demographic groups.
Approach: They propose a framework for benchmarking hierarchical gender hiring bias in Large Language Models for resume scoring.
Outcome: The proposed framework reveals significant issues of reverse gender hiring bias and overdebiasing in ten state-of-the-art LLMs.
TANet: Thread-Aware Pretraining for Abstractive Conversational Summarization (2022.findings-naacl)

Copied to clipboard

Challenge: Existing pre-trained language models are difficult to apply to abstractive conversational summarization tasks.
Approach: They propose a thread-aware Transformer-based network that incorporates contextual dependency into the conversational summarization model.
Outcome: The proposed model can be applied to real conversations using a large-scale pretraining dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations