Papers by Ziyu Yang

11 papers
Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) can solve reasoning and mathematical problems using the Chain-of-Thought technique, but require costly and long CoT data and fine-tuning.
Approach: They propose a method that uses Sparse Autoencoders to extract interpretable features from vanilla CoT and use them to steer the LLM's internal states.
Outcome: The proposed method uses Sparse Autoencoders (SAEs) to extract interpretable features from vanilla CoT and steer the LLM's internal states during generation.
Progra: Progress-Aware Reinforcement Learning for Multi-Turn Function Calling (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for multi-turn function calling are limited by redundancy and lack explicit integration of progress awareness into training.
Approach: They propose a framework that explicitly integrates progress awareness into LLM training for multi-turn function calling.
Outcome: Empirical results show that Progra outperforms existing methods on two public benchmarks.
Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Sparse Mixture-of-Experts (SMoE) architectures require loading all expert parameters . previous work focused on expert pruning and merging but focused on neuron-level structure .
Approach: They propose a task-agnostic framework for expert pruning and reconstruction . it prunes redundant experts using router statistics, then decomposes them into neuron-level expert segments .
Outcome: The proposed framework reduces the number of experts and memory usage, making it easier to deploy.
LoraRetriever: Input-Aware LoRA Retrieval and Composition for Mixed Tasks in the Wild (2024.findings-acl)

Copied to clipboard

Challenge: Low-Rank Adaptation (LoRA) is an effective yet efficient solution for fine-tuning large language models.
Approach: They propose a low-rank Adaptation framework that retrieves and composes multiple LoRAs according to input prompts.
Outcome: Experimental results show that LoraRetriever outperforms baselines in terms of performance and versatility.
CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution (2026.acl-long)

Copied to clipboard

Challenge: Extensive experiments on AppWorld and BFCL demonstrate consistent and significant improvements over strong base models, yielding absolute gains of 19.43%, 15.58%, and 18.14%, respectively.
Approach: They propose a framework that extracts feedback signals such as forgetting and uncertainty from rollout trajectories and utilizes them to guide LLM-based task synthesis.
Outcome: Extensive experiments on AppWorld and BFCL show that the proposed framework improves over strong base models.
Two-Pronged Human Evaluation of ChatGPT Self-Correction in Radiology Report Simplification (2024.findings-acl)

Copied to clipboard

Challenge: Radiology reports are highly technical documents aimed primarily at doctor-doctor communication.
Approach: They propose a new evaluation protocol that employs radiologists and laypeople to produce high-quality simplifications.
Outcome: The proposed evaluation protocol combines radiologists and laypeople to produce high-quality simplifications.
LaoPLM: Pre-trained Language Models for Lao (2022.lrec-1)

Copied to clipboard

Challenge: Pre-trained language models (PLMs) can capture different levels of concepts in context . previous work on Lao has been hampered by the lack of annotated datasets .
Approach: They construct a text classification dataset to alleviate the resource-scarce situation of Lao . they evaluate them on two downstream tasks: part-of-speech tagging and text classification .
Outcome: The proposed model can capture different levels of concepts in context and generate universal language representations.
H-LegalKI: A Hierarchical Legal Knowledge Integration Framework for Legal Community Question Answering (2024.findings-emnlp)

Copied to clipboard

Challenge: Legal question answering (LQA) aims to bridge the gap between limited availability of legal professionals and the extensive volume of legal issues.
Approach: They propose a legal knowledge retriever and a hierarchical legal knowledge integration framework to address multiple user-specific circumstances.
Outcome: The proposed framework outperforms baselines on the legal community question-answering dataset.
Data Augmentation for Radiology Report Simplification (2023.findings-eacl)

Copied to clipboard

Challenge: Existing approaches to improve radiology reports are limited due to the high cost of manual simplification.
Approach: They propose a data augmentation approach to generate simplifications of unlabeled radiology sentences using a pre-trained language model and paraphrasing of labeled radiologists sentences.
Outcome: The proposed model generates simplifications of unlabeled radiology sentences and paraphrases labeled radiologists sentences.
MidPO: Dual Preference Optimization for Safety and Helpfulness in Large Language Models via a Mixture of Experts Framework (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent studies address safety-constrained online and offline preferences optimizations, but offline methods perform poorly in adaptively balancing safety and helpfulness.
Approach: They propose a mixture of experts framework for safety-helpfulness dual Preference Optimization . they combine a single-preference enhanced direct preference optimization approach with a dynamic routing mechanism .
Outcome: The proposed framework outperforms state-of-the-art methods in safety and helpfulness.
OS Agents: A Survey on MLLM-based Agents for Computer, Phone and Browser Use (2025.acl-long)

Copied to clipboard

Challenge: a new generation of (M)LLMs is enabling the creation of superintelligent AI assistants . OS Agents can complete tasks autonomously and have the potential to significantly enhance the lives of billions of users worldwide.
Approach: They propose to build OS Agents that operate within operating systems' GUIs and GUIs . they examine evaluation metrics and benchmarks to identify promising directions .
Outcome: The proposed agents are based on operating systems (OS) and operating systems frameworks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations