Papers by Zhi Wang

39 papers
CODERL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models excel at code generation by learning from vast code corpora, but a fundamental semantic gap remains between training on textual patterns and the goal of functional correctness . reinforcement learning with verifiable rewards (RLVR) approaches are inefficient for establishing a well-aligned connection between the textual representation of code and its execution semantics.
Approach: They propose a novel approach that integrates execution semantics alignment into the RLVR training pipeline for code generation.
Outcome: The proposed model outperforms baseline training and RLVR and shows strong applicability across RL and LLMs.
Unsupervised Sentence Representation Learning with Syntactically Aligned Negative Samples (2025.findings-naacl)

Copied to clipboard

Challenge: Existing approaches to sentence representation learning often encounter semantic inconsistencies and feature suppression.
Approach: They propose a method for generating syntactically aligned negative (SAN) samples using a semantic importance-aware Masked Language Model (MLM) approach.
Outcome: The proposed method produces negative samples with substantial textual overlap with the original sentences while conveying different meanings.
Learning from Contrasts: Synthesizing Reasoning Paths from Diverse Search Trajectories (2026.acl-long)

Copied to clipboard

Challenge: MCTS methods retain only the single highest-reward trajectory, discarding comparative signals present in the many explored paths.
Approach: They propose a framework that transforms supervision extraction into a synthesis procedure.
Outcome: The proposed framework matches or exceeds baselines on 60K CRPS-synthesized examples on out-of-domain benchmarks.
I-AM-G: Interest Augmented Multimodal Generator for Item Personalization (2024.emnlp-main)

Copied to clipboard

Challenge: e-commerce and recommender systems lack a framework for personalized generation . a new framework extracts tags from multimodal information of items that the user has interacted with .
Approach: They propose a framework that extracts tags from multimodal information and rewrites item description . they then use a decoupled text-to-text and image-to image retriever to search for similar item text .
Outcome: The proposed framework can generate results aligned with user preferences . it can be used in e-commerce and recommender systems to win over diverse user base .
DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories (2024.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks are poorly aligned with real-world code repositories and are insufficient to evaluate the coding abilities of Large Language Models (LLMs).
Approach: They propose a repository-level benchmark named DevEval to evaluate LLMs' coding abilities in real-world code repositories.
Outcome: The proposed benchmarks show that the LLMs perform better in real-world code repositories than existing benchmarks.
SPARD: Self-Paced Curriculum for RL Alignment via Integrating Reward Dynamics and Data Utility (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are shifting the focus from single verifiable tasks toward complex, open-ended real-world scenarios.
Approach: They propose a framework that automatically adjusts reward weights and data importance to synchronize learning intent with data utility for optimal performance.
Outcome: The proposed framework improves model capabilities across all domains and scales.
Granular Entity Mapper: Advancing Fine-grained Multimodal Named Entity Recognition and Grounding (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for fine-grained content extraction are limited by long-tailed distribution of textual entity categories and performance of object detectors.
Approach: They propose a multi-granularity entity recognition module and a reranking module to integrate hierarchical information of entity categories, visual cues, and external textual resources collectively.
Outcome: The proposed framework achieves state-of-the-art on the fine-grained content extraction task.
UnClE: Explicitly Leveraging Semantic Similarity to Reduce the Parameters of Word Embeddings (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to reduce word embedding parameters ignore semantic information . existing methods do not consider semantic information, allowing for performance degradation .
Approach: They propose a method that leverages semantic similarity with weight sharing to reduce dimensionality of word embeddings.
Outcome: The proposed method reduces word embedding parameters by more than 11x on a standard English-German dataset.
ExploraCoder: Advancing Code Generation for Multiple Unseen APIs via Planning and Chained Exploration (2025.acl-long)

Copied to clipboard

Challenge: Large language models face intrinsic limitations in coding with unseen APIs in training corpora.
Approach: They propose a training-free framework that empowers LLMs to invoke multiple unseen APIs in code solution by planning a complex problem into several API invocation subtasks and experimenting with correct API usage at intermediate steps.
Outcome: The proposed framework significantly improves performance for models lacking prior API knowledge, achieving 11.99% over retrieval-based approaches and 17.28% over pretraining-based methods in pass@10.
Time-for-Accuracy: Formalizing Chain-of-Thought as an Expansion of Logical Depth (2026.findings-acl)

Copied to clipboard

Challenge: Chain-of-thought (CoT) prompting can improve multi-step reasoning, but it is unclear what kind of additional sequential computation longer traces actually enable.
Approach: They propose a deletion-based measure of step necessity under a specified inference interface to operationalize realized depth beyond raw length.
Outcome: The proposed method combines effective logical depth with Bennett's logical depth to show that it is more efficient than a linear model.
From What Is Said to Why It Is Framed: Intent-Aware News Video Understanding (2026.findings-acl)

Copied to clipboard

Challenge: Existing verification methods for short-form news videos neglect communicative intent . stylistic presentation and factual manipulation are often intertwined, resulting in shortcut learning .
Approach: They propose a theory-grounded representation of communicative intent that captures creator stance, audience need activation, and communication strategy.
Outcome: The proposed framework captures creator stance, audience need activation, and communication strategy.
Focus-Constrained Attention Mechanism for CVAE-based Response Generation (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing models generate high-frequency but trivial responses such as "I don't know" or "I'm ok" due to the discrepancy in discourse-level information, standard models generate one-to-many relationships.
Approach: They propose to transform coarse-grained discourse-level information into fine-grounded word-level knowledge by introducing a fine-grain focus signal and a focus-constrained attention mechanism to take full advantage of focus.
Outcome: The proposed model can generate more diverse and informative responses compared with state-of-the-art models.
On Safety Risks in Experience-Driven Self-Evolving Agents (2026.findings-acl)

Copied to clipboard

Challenge: Experience-driven self-evolution has emerged as a promising paradigm for improving the autonomy of large language model agents, yet its reliance on self-curated experience introduces underexplored safety risks.
Approach: They investigate how experience accumulation and utilization in self-evolving agents affect safety performance across web-based and embodied environments.
Outcome: The findings expose inherent limitations of current self-evolving agents and call for more principled strategies to ensure safe and reliable adaptation.
SelfAug: Mitigating Catastrophic Forgetting in Retrieval-Augmented Generation via Distribution Self-Alignment (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing solutions for supervised fine-tuning often lead to catastrophic forgetting, where models lose their previously acquired knowledge and general capabilities.
Approach: They propose a self-distribution alignment method that aligns input sequence logits to preserve the model’s semantic distribution, thereby mitigating catastrophic forgetting and improving downstream performance.
Outcome: The proposed method achieves a superior balance between downstream learning and general capability retention.
Registering Source Tokens to Target Language Spaces in Multilingual Neural Machine Translation (2025.acl-long)

Copied to clipboard

Challenge: Multilingual neural machine translation (MNMT) aims for arbitrary translations across multiple languages.
Approach: They propose a method that inserts a set of tokens specifying the target language into the input sequence between the source and target tokens.
Outcome: The proposed method outperforms existing models on a large-scale benchmark.
NeuroSym-Cal: Bridging the Reasoning-Execution Gap in Code Generation via Hierarchical Calibration (2026.findings-acl)

Copied to clipboard

Challenge: Existing calibration methods rely on the assumption that consensus implies correctness . Existing methods fail under systematic errors, leading to miscalibrated high-confidence predictions.
Approach: They propose a hierarchical calibration framework that measures confidence at two levels . they propose sensitivity analysis to measure local curvature of deductive process .
Outcome: The proposed framework de-saturates overconfident errors and improves selective generation performance on OOD benchmarks.
PATIENT-𝜓: Using Large Language Models to Simulate Patients for Training Mental Health Professionals (2024.emnlp-main)

Copied to clipboard

Challenge: Mental illness remains one of the most critical public health issues.
Approach: They propose a patient simulation framework for cognitive behavior therapy training that uses large language models to act as a simulated therapy patient.
Outcome: The proposed framework improves the skill acquisition and confidence of mental health trainees beyond textbooks, videos, and role-play with non-patients.
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models (2025.findings-acl)

Copied to clipboard

Challenge: Existing multi-objective preference alignment methods for large language models face limitations such as auxiliary reward/reference models and computational complexity.
Approach: They propose a framework that achieves dynamic balance across preference dimensions by using dimension-aware generation metrics as implicit rewards.
Outcome: Empirical results show that AMoPO outperforms state-of-the-art methods by 28.5% .
Double-Checker: Large Language Model as a Checker for Few-shot Named Entity Recognition (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent studies have demonstrated remarkable performance on few-shot Named Entity Recognition tasks due to the high cost of obtaining high-quality labeled data.
Approach: They propose to decompose the task into entity span detection and entity type classification using a type-independent entity span detector and then classify the detected spans based on their types.
Outcome: The proposed method consistently yields improvements over two baseline approaches.
MLeVLM: Improve Multi-level Progressive Capabilities based on Multimodal Large Language Model for Medical Visual Question Answering (2024.findings-acl)

Copied to clipboard

Challenge: Existing MVQA models ignore multi-level progressive capabilities due to unspecific data and plain architecture.
Approach: They propose a multi-level visual language model for medical visual question answering (MVQA) which covers multi- level questions and answers as well as reasoning processes from visual clues to semantic cognition.
Outcome: The proposed model outperforms existing medical multimodal large language models on a multi-level instruction dataset and a feature alignment module.
LLMs can Perform Multi-Dimensional Analytic Writing Assessments: A Case Study of L2 Graduate-Level Academic English Writing (2025.acl-long)

Copied to clipboard

Challenge: a growing number of studies have indicated the general usefulness of LLMs for automated writing assessments.
Approach: They propose a framework that evaluates LLMs' ability to provide scores and comments based on multiple assessment criteria.
Outcome: The proposed framework is interpretable, cost-efficient, scalable, and reproducible . it is compared to existing methods that rely on manual judgments .
Semantic-Space Exploration and Exploitation in RLVR for LLM Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have substantially improved the reasoning abilities of Large Language Models (LLMs).
Approach: They propose a method that balances exploration and exploitation in the hidden-state space of response trajectories.
Outcome: The proposed model yields consistent improvements across models, algorithms and reasoning benchmarks.
AlphaContext: An Evolutionary Tree-based Psychometric Context Generator for Creativity Assessment (2026.acl-long)

Copied to clipboard

Challenge: Existing LLM-based tools struggle with insufficient assessment cues, weak narrative coherence, limited stylistic diversity, and poor support for creative thinking.
Approach: They propose an evolutionary tree-based psychometric context generator that integrates rule-guided outline planning, sentence-level MCTS generation, MAP-Elites quality-diversity optimization and assessment-guide refiner simulation.
Outcome: The proposed tool outperforms strong LLMs and structured frameworks on 7 evaluation dimensions and shows higher alignment with expert-designed contexts.
Think and Recall: Layer-Level Prompting for Lifelong Model Editing (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for lifelong model editing suffer from limitations in usability, such as requiring additional training corpora or lacking support for reversible and detachable edits.
Approach: They propose a plug-and-play method for knowledge retrieval and storage, i.e., Layer-Level Prompting, which enables seamless and efficient lifelong model editing.
Outcome: The proposed method outperforms existing methods on question answering and hallucination benchmarks across different LLMs.
Infusing Sequential Information into Conditional Masked Translation Model with Self-Review Mechanism (2020.coling-main)

Copied to clipboard

Challenge: Existing non-autoregressive models generate target words in parallel, but with a large latency due to the left-to-right dependency.
Approach: They propose to train a conditional masked translation model and refine results within several iterations to remedy a flawed translation by non-autoregressive models.
Outcome: The proposed model outperforms state-of-the-art models by over 1 BLEU while using less training computations.
TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text (2025.emnlp-main)

Copied to clipboard

Challenge: Current RAG systems concatenate and process numerous retrieved document chunks for prefill . this leads to significant latency in time-to-first-token (TTFT) Experimental results demonstrate that TurboRAG reduces TTFT by up to 9.4x compared to the conventional RAG system.
Approach: They propose a hybrid offline-online paradigm that precomputes chunk-level key-value caches and stitches them together at inference time using independent–attention and reorderedRoPE techniques.
Outcome: Experimental results show that TurboRAG reduces TTFT by 9.4x compared to the conventional RAG systems . long concatenated contexts consume disproportionate GPU memory, limiting throughput .
Reasoning Over Space: Enabling Geographic Reasoning for LLM-Based Generative Next POI Recommendation (2026.acl-long)

Copied to clipboard

Challenge: Existing LLM-based recommenders lack explicit modeling of geographic signals . without explicit modeling geographic signals, recommenders struggle to capture core mobility patterns .
Approach: They propose a framework that utilizes geography as a decision variable within the reasoning process.
Outcome: The proposed framework achieves over 10% relative gains in hit rate over strongest LLM-based baselines and improves cross-city transfer.
bert2BERT: Towards Reusable Pretrained Language Models (2022.acl-long)

Copied to clipboard

Challenge: Pre-training large language models can be expensive and wasteful.
Approach: They propose a method which can transfer the knowledge of an existing smaller pre-trained model to a large model through parameter initialization and a two-stage learning method to further accelerate the pre-training.
Outcome: The proposed method can transfer the knowledge of an existing smaller pre-trained model to a large model through parameter initialization and significantly improve the pre-training efficiency of the large model.
Beyond A Single AI Cluster: A Survey of Decentralized LLM Training (2025.emnlp-main)

Copied to clipboard

Challenge: Decentralized LLM training leverages dispersed resources at varying scales.
Approach: They propose a resource-driven paradigm that leverages dispersed resources across clusters, datacenters and even regions.
Outcome: The proposed model scales are 175 billion to 660 billion parameters, and the exponential growth in computational requirements poses significant challenges.
KoCo-Bench: Can Large Language Models Leverage Domain Knowledge in Software Development? (2026.acl-long)

Copied to clipboard

Challenge: Existing domain-specific code benchmarks focus on assessing what knowledge LLMs possess rather than how they acquire and apply new knowledge.
Approach: They propose a benchmark to evaluate domain specialization methods in real-world software development.
Outcome: KOCO-bench is a new benchmark for evaluating domain specialization methods in real-world software development.
MusTQ: A Temporal Knowledge Graph Question Answering Dataset for Multi-Step Temporal Reasoning (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies focus on fact-centered reasoning with limited attention to temporal reasoning.
Approach: They propose a new TKGQA dataset, MusTQ, which contains 666K multi-step temporal reasoning questions and a TKG.
Outcome: The proposed model achieves state-of-the-art multi-step temporal reasoning ability with entity-time attention mechanism and optimized temporal knowledge graph representation.
GUI0: Self-Evolving Foundational GUI Agents in Super App Ecosystems (2026.acl-long)

Copied to clipboard

Challenge: Automated interaction with graphical user interfaces (GUIs) is central to general artificial intelligence, but remains challenging within Super App ecosystems.
Approach: They propose a framework synergizing autonomous data synthesis with dual-agent co-evolution . GUI0 establishes a domain-aware foundation model via synthesized corpora and employs curriculum-driven reinforcement learning .
Outcome: The proposed framework outperforms Gemini-2.5-Pro and Claude-4-Sonnet in the SuperAPP benchmark and has universal efficacy across base models.
VulAgent: Hypothesis-Validation Driven Multi-Agent Architecture for Vulnerability Detection (2026.findings-acl)

Copied to clipboard

Challenge: Recent reports indicate that software vulnerabilities caused by insecure coding practices remain a major security threat.
Approach: They propose a multi-agent vulnerability detection framework based on hypothesis validation . they use multi-view analyzers to localize and localize security-sensitive operations .
Outcome: The proposed framework reduces false positives and increases accuracy by 6.6 percentage points on PrimeVul and SVEN.
SEAG: Structure-Aware Event Causality Generation (2023.findings-acl)

Copied to clipboard

Challenge: Current methods for extracting event causality are limited by the lack of cross-task dependencies and may cause error propagation.
Approach: They propose an approach for Structure-Aware Event Causality Generation (SEAG) they generate the ECG structure using a pre-trained language model and perform structural discriminative training alongside auto-regressive generation.
Outcome: The proposed method is effective in extracting event causality from text.
FinEval: A Chinese Financial Domain Knowledge Evaluation Benchmark for Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Large language models have demonstrated outstanding performance in various natural language processing tasks, but their security capabilities in the financial domain have not been explored.
Approach: They propose to use a benchmark to evaluate large language models' financial domain knowledge and practical abilities.
Outcome: The proposed benchmark evaluates large language models' financial domain knowledge and practical abilities.
Your Inference Request Will Become a Black Box: Confidential Inference for Cloud-based Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches fail to ensure privacy, maintain model performance, and preserve computational efficiency simultaneously.
Approach: They propose a confidential inference framework that partitions the LLM pipeline between a client-verified Confidential Virtual Machine (CVM) and the public cloud to protect client data without compromising the cloud’s model intellectual property or inference quality.
Outcome: The proposed framework can defend against state-of-the-art token inference attacks while preserving model privacy, performance, and efficiency.
QueryAttack: Jailbreaking Aligned Large Language Models Using Structured Non-natural Query Language (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods to bypass security defenses of large language models (LLMs) are not effective, but QueryAttack can be jailbroken.
Approach: They propose a framework to examine generalizability of safety alignment by translating malicious queries into structured non-natural query languages.
Outcome: The proposed framework can achieve high attack success rates and jailbreak various defense methods on mainstream LLMs.
Bridging Kernel Drivers and Virtual Device Models with LLM-Powered Automation (2026.acl-demo)

Copied to clipboard

Challenge: Linux kernel device drivers are tightly coupled with hardware, making them difficult to execute and test without physical devices.
Approach: They present a tool that generates QEMU-based virtual devices directly from Linux driver source code.
Outcome: The proposed tool generates QEMU-based virtual devices directly from Linux driver source code.
CLEVA: Chinese Language Models EVAluation Platform (2023.emnlp-demo)

Copied to clipboard

Challenge: Large language models (LLMs) have revolutionized natural language processing.
Approach: They propose a Chinese-based platform that assesses Chinese LLMs using a standardized workflow and a unique sampling strategy.
Outcome: CLEVA evaluates Chinese LLMs on a standardized workflow and a competitive leaderboard with minimal coding.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations