Papers by Jun Zhuang

10 papers
Experience-driven Multi-turn Reinforcement Learning for GUI Agents (2026.acl-long)

Copied to clipboard

Challenge: GUI agents have demonstrated remarkable progress in automating complex user interface interactions . training such agents for long-horizon tasks remains challenging due to limited rewards and prohibitive costs.
Approach: They propose a method that leverages expert trajectories as environment experiences for on-policy multi-turn training.
Outcome: The proposed method achieves significant gains over the base model with 1K public trajectories as RL experiences . it achieves competitive performance against strong baselines such as UI-TARS-7B and GPT-4o .
De-Biased Court’s View Generation with Causality (2020.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to court’s view generation can be used to address this problem, but neglecting the confounding bias in data can limit the model performance and pollute learning outcomes.
Approach: They propose a novel Attentional and Counterfactual based Natural Language Generation method consisting of an attentional encoder and a pair of innovative counterfactual decoders to generate judgment-discriminative court's views.
Outcome: The proposed method is able to generate judgment-discriminative court's views (both supportive and non-supportive views) under both quantitative and qualitative evaluation metrics.
Video Dialog via Progressive Inference and Cross-Transformer (D19-1)

Copied to clipboard

Challenge: Existing visual dialog methods use RNN to encode the dialog history as a vector representation . a new method for video dialog is proposed, which progressively updates query information based on dialog history and video content until the agent think the information is sufficient and unambiguous.
Approach: They propose a method which progressively updates query information based on dialog history and video content until the agent thinks it is sufficient and unambiguous.
Outcome: The proposed method can be used to infer video dialog answers on large-scale datasets.
Large Language Models Can Help Mitigate Barren Plateaus in Quantum Neural Networks (2026.findings-acl)

Copied to clipboard

Challenge: Quantum Neural Networks (QNNs) are often hindered by barren plateaus (BPs) barren peaks are where gradient variance vanishes exponentially as qubit size increases .
Approach: They propose a framework that leverages large language models with the submartingale property to iteratively synthesize initial parameters for QNNs that yield non-negligible gradient variance.
Outcome: The proposed framework outperforms existing initialization methods in maintaining higher gradient variance across various QNN scales.
Natural Language Video Localization with Learnable Moment Proposals (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for video moment localization have poor performance due to predefined rules.
Approach: They propose a model with a fixed set of learnable moment proposals with 'border-aware loss' they propose to localize the video moment corresponding to the query by locating the start and end timestamps in an untrimmed video.
Outcome: The proposed model outperforms state-of-the-art models on two challenging benchmarks.
UI-Copilot: Advancing Long-Horizon GUI Automation via Tool-Integrated Policy Optimization (2026.acl-long)

Copied to clipboard

Challenge: Experimental results show that UI-Copilot-7B achieves state-of-the-art performance on challenging MemGUI-Bench, outperforming strong 7B-scale GUI agents such as GUI-Owl-7B and UITARS-1.5-7B.
Approach: They propose a collaborative framework where the GUI agent focuses on task execution while a lightweight copilot provides on-demand assistance for memory retrieval and numerical computation.
Outcome: The proposed framework outperforms GUI-Owl-7B and UI-TARS-1.5-7B on MemGUI-Bench and delivers 17.1% improvement on AndroidWorld over the base Qwen model.
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation (2025.findings-emnlp)

Copied to clipboard

Challenge: Prior work has shown that intent detection enhances LLMs’ moderation guardrails, but the robustness of these guardrail mechanisms under malicious manipulations remains under-explored.
Approach: They propose a two-stage intent-based prompt-refinement framework that first transforms harmful inquiries into structured outlines and further reframes them into declarative-style narratives.
Outcome: The proposed framework outperforms several cutting-edge jailbreak methods and evades even advanced Intent Analysis (IA) and Chain-of-Thought (CoT)-based defenses.
Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly deployed in high-stakes domains such as finance where unsafe behavior can lead to serious regulatory risks.
Approach: They propose a black-box multi-turn risk-concealed redteaming framework that progressively conceals surface-level risk while exploiting regulatory-violating behaviors.
Outcome: Experiments on nine widely used LLMs show that the proposed framework achieves 93.19% average attack success rate (ASR) and improves the average ASR to 95.00%.
Self-Contrast: Better Reflection Through Inconsistent Solving Perspectives (2024.acl-long)

Copied to clipboard

Challenge: Recent research indicates without external feedback, LLM’s intrinsic reflection is unstable.
Approach: They propose a method that combines self-evaluated and external feedback to improve LLM's reflection.
Outcome: The proposed method improves the quality of self-evaluated feedback and can catalyze more accurate and stable reflection.
Now You Hear Me: Audio Narrative Attacks Against Large Audio–Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Existing jailbreaks against large audio-language models fall into two categories . early work converted text-based prompts into synthetic speech, while subsequent work introduced minor acoustic variations such as accent shifts, phonetic spellings, or stress patterns.
Approach: They propose a text-to-audio jailbreak that embeds disallowed directives within a narrative-style audio stream.
Outcome: The proposed attack exploits structural and acoustic properties of a text-to-audio model . it achieves 98.26% success rate, significantly exceeding baselines for text-based models .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations