Papers by Yanzhe Zhang
Generative Interfaces for Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Large language models are increasingly seen as assistants, copilots, and consultants . however, their linear request-response format often makes interactions inefficient in multi-turn tasks . |
| Approach: | They propose a paradigm in which large language models respond to user queries by generating user interfaces that enable more adaptive and interactive engagement. |
| Outcome: | The proposed paradigm outperforms traditional chat-based interfaces in many tasks and interaction patterns. |
Continual Sequence Generation with Adaptive Compositional Modules (2022.acl-long)
Copied to clipboard
| Challenge: | Existing work on continual sequence generation relies on reuse of existing parameters to learn new tasks or blindly adds new parameters for each new task. |
| Approach: | They propose to use continual sequence generation to reuse existing parameters for new tasks . current continuous learning models often forget knowledge of dissimilar tasks if data distributions shift . |
| Outcome: | The proposed framework outperforms state-of-the-art models on various sequences of generation tasks. |
Leveraging Expert Guided Adversarial Augmentation For Improving Generalization in Named Entity Recognition (2022.findings-acl)
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) systems perform well on in-distribution data, but perform poorly on examples drawn from a shifted distribution. |
| Approach: | They propose to use expert-guided heuristics to change entity tokens and their contexts to alter their entity types as adversarial attacks. |
| Outcome: | The proposed model significantly improves performance on the challenging set and out-of-domain generalization. |
AudioPrivacy: Parallel Audio Dataset for Speaker Profiling with Diverse Audio Types and Rich Attributes (2026.findings-acl)
Copied to clipboard
| Challenge: | Speech signals convey abundant speaker-related metadata, yet current privacy research focuses on identity-centric voiceprint protection, leaving sensitive Speaker Attribute Privacy (SAP) underexplored. |
| Approach: | They propose a large-scale Chinese dataset to evaluate speaker-related privacy leakage . the dataset includes 227.3 hours of audio from 1,000 speakers . |
| Outcome: | The proposed model systematically evaluates speaker-related privacy leakage in everyday scenarios. |
Continual Learning for Text Classification with Information Disentanglement Based Regularization (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing continual learning methods focus on preserving knowledge from previous tasks . Continual learning is a useful tool for learning over time, but it is not always possible to generalize to new tasks. |
| Approach: | They propose a disentanglement-based regularization method for continual learning on text classification that disentangles text hidden spaces into generic representations and regularizes them differently to constrain knowledge required to generalize. |
| Outcome: | The proposed method disentangles text hidden spaces into representations that are generic to all tasks and representations specific to each individual task. |
Bounding the Capabilities of Large Language Models in Open Text Generation with Prompt Constraints (2023.findings-eacl)
Copied to clipboard
| Challenge: | Existing and potential applications of open-ended text generation are farreaching, spanning domains such as QA, story generation, open-end dialogue, and ChatGPT 1 . |
| Approach: | They propose a prompt-centric approach to analyzing and bounding the abilities of open-ended generative models by a set of structural and stylistic prompts. |
| Outcome: | The proposed method can be generalized to other large models like BLOOM and OPT. |
Attacking Vision-Language Computer Agents via Pop-ups (2025.acl-long)
Copied to clipboard
| Challenge: | Existing tools for analyzing and testing VLMs are lacking in understanding what types of attacks are possible and what types are effective. |
| Approach: | They propose to integrate pop-ups into existing agent testing environments to attack VLM agents by ignoring them. |
| Outcome: | The proposed attack success rate is 86% and decreases by 47% when integrating pop-ups into existing agent testing environments. |
Distilling an End-to-End Voice Assistant Without Instruction Training Data (2025.acl-long)
Copied to clipboard
| Challenge: | Recent efforts to train speech-only LLMs have led to models “forging” speech information from text-only models. |
| Approach: | They propose a paradigm for training Speech Large Language Models without instruction data by using the response of a text-only LLM to transcripts as self-supervision. |
| Outcome: | The proposed model generalizes to Spoken Question Answering, Classification, and Translation and achieves a 72% win rate compared with state-of-the-art models like Qwen 2 Audio . |
Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering (2025.naacl-long)
Copied to clipboard
| Challenge: | Generative AI has made rapid advances in multimodal understanding and code generation. |
| Approach: | They construct a first real-world benchmark for multimodal large language models that directly convert visual designs into code implementations by manually curating 484 diverse real-life webpages as test cases. |
| Outcome: | The proposed model can generate code implementations that directly render into the given reference webpages, given the screenshots as input. |
EgoNormia: Benchmarking Physical-Social Norm Understanding (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing VLMs lack robust grounded norm understanding, a new study finds . current VLM models lack robust grounding, despite a high score for safety and privacy . |
| Approach: | They propose a pipeline to generate grounded MCQs from ego-centric videos of human interactions. |
| Outcome: | The proposed pipeline can generate grounded MCQs from egocentric video . it shows that current VLMs lack robust grounded norm understanding . |
Robustness of Demonstration-based Learning Under Limited Data Scenario (2022.emnlp-main)
Copied to clipboard
| Challenge: | Current large pretrained language models struggle to learn NLP tasks under limited data scenarios. |
| Approach: | They propose to augment input with some demonstrations to improve model performance under limited data scenarios. |
| Outcome: | The proposed demonstrations improve performance on few-shot NER tasks and show that the length of demonstrations and relevance of random tokens are the main factors affecting the model's performance. |
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation (2026.acl-long)
Copied to clipboard
Hui Wang, Jinghua Zhao, Yifan Yang, Shujie Liu, Junyang Chen, Yanzhe Zhang, Shiwan Zhao, Jinyu Li, Jiaming Zhou, Haoqin Sun, Yan Lu, Yong Qin
| Challenge: | Existing methods for evaluating the perceptual quality of synthetic speech are limited due to the complexity of perceptual quality factors and the diversity of speech generation tasks. |
| Approach: | They propose a new paradigm for enabling large language models to conduct structured speech quality evaluation using a large-scale dataset. |
| Outcome: | The proposed model performs well across tasks and languages. |
Sketch2Code: Evaluating Vision-Language Models for Interactive Web Design Prototyping (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing research on UI/UX automation often requires high-fidelity inputs like Figma designs or detailed screenshots, limiting accessibility and impeding efficient design iteration. |
| Approach: | They propose a benchmark that evaluates state-of-the-art Vision Language Models on converting sketches into webpage prototypes. |
| Outcome: | The benchmark evaluates state-of-the-art Vision Language Models on automating the conversion of rudimentary sketches into webpage prototypes. |