Papers by Shuyang Yu

11 papers
Miner: Mining Intrinsic Mastery for Data-Efficient RL in Large Reasoning Models (2026.acl-long)

Copied to clipboard

Challenge: Current critic-free RL methods for large reasoning models suffer from severe inefficiency when training on positive homogeneous prompts.
Approach: They propose a method that repurposes the policy’s intrinsic uncertainty as a self-supervised reward signal, with no external supervision, auxiliary models, or additional inference cost.
Outcome: Evaluated across six reasoning benchmarks on Qwen3-4B and Qwend3-8B base models, the proposed method achieves state-of-the-art performance among the other four methods.
Dynamic Uncertainty Ranking: Enhancing Retrieval-Augmented In-Context Learning for Long-Tail Knowledge in LLMs (2025.naacl-long)

Copied to clipboard

Challenge: Prior work has shown that in-context learning (ICL) with retriever augmentation can help LLMs better capture long-tail knowledge, reducing their reliance on pre-trained data.
Approach: They propose a reinforcement learning-based dynamic uncertainty ranking method that accounts for the varying impact of each retrieved sample on LLM predictions.
Outcome: The proposed method outperforms baseline models on question-answering datasets by 2.76% and 5.96% on long-tail questions that elude zero-shot inference.
Self-Improvement of Non-autoregressive Model via Sequence-Level Distillation (2023.emnlp-main)

Copied to clipboard

Challenge: Existing non-autoregressive Transformers (NAT) models generate the entire sequence in parallel, but the multimodality problem limits their performance.
Approach: They propose a method to generate distilled data by the NAT model itself, eliminating the need for additional teacher networks.
Outcome: The proposed method can generate distilled data by the NAT model without teacher networks and adapt to different NAT models without precise adjustments.
Towards Omni-RAG: Comprehensive Retrieval-Augmented Generation for Large Language Models in Medical Applications (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to source planning fail to achieve this due to misalignment between the model’s expectation of the sources and their actual content.
Approach: They propose a method to optimise large-scale medical knowledge models by combining multiple medical knowledge sources into one query.
Outcome: The proposed method significantly improves multi-source planning performance while training a smaller model to learn source alignment.
Neural Parameter Search for Slimmer Fine-Tuned Models and Better Transfer (2025.acl-long)

Copied to clipboard

Challenge: Foundational models and their checkpoints have advanced deep learning, boosting performance across applications.
Approach: They propose a method for pruning fine-tuned models by calculating differences between them and original model.
Outcome: The proposed method can improve performance across vision, NLP, and multi-modal benchmarks.
EvoR: Evolving Retrieval for Code Generation (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing pipelines for retrieval-augmented code generation (RACG) use static knowledge bases with a single source, limiting adaptation capabilities of Large Language Models (LLMs) Extensive experiments demonstrate that EVOR achieves two to four times of execution accuracy compared to other methods such as Reflexion.
Approach: They propose a retrieval-augmented code generation pipeline that employs the synchronous evolution of queries and diverse knowledge bases.
Outcome: The proposed pipeline achieves two to four times of execution accuracy compared to other methods.
MedCare: Advancing Medical LLMs through Decoupling Clinical Alignment and Knowledge Aggregation (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have made significant progress in natural language understanding and generation, proving valuable especially in the medical field.
Approach: They propose a medical LLM through decoupling Clinical Alignment and Knowledge Aggregation which uses a and a to encode diverse knowledge in the first stage and filter out detrimental information.
Outcome: The proposed model achieves promising performance on over 20 medical tasks and specific medical alignment tasks.
ReflecTool: Towards Reflection-Aware Tool-Augmented Clinical Agents (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown promising potential in the medical domain, assisting with tasks like clinical note generation and patient communication.
Approach: They propose a framework that excels at utilizing domain-specific tools within two stages.
Outcome: The proposed framework surpasses the pure LLMs with more than 10 points and the well-established agent-based methods with 3 points.
To See a World in a Spark of Neuron: Disentangling Multi-Task Interference for Training-Free Model Merging (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to model merging ignore the fundamental roles of neurons, connectivity and activation.
Approach: They propose a framework that relies on neuronal mechanisms to mitigate task interference . they decomposed task-specific representations into two complementary subspaces . their results offer new insights into mitigating task interference and improving knowledge fusion .
Outcome: The proposed framework reduces task interference within neurons and improves knowledge fusion.
APo-VAE: Text Generation in Hyperbolic Space (2021.naacl-main)

Copied to clipboard

Challenge: Existing models that learn embeddings only in Euclidean vector space do not account for such structural property of language.
Approach: They propose a Poincare Variational Autoencoder to capture latent hierarchies in hyperbolic space . they propose enabling adversarial learning procedures to empower robust model training .
Outcome: The proposed model outperforms existing models in a hyperbolic latent space . it captures latent language hierarchies in hyperbolical space and is robust to training .
Knowledge Fusion By Evolving Weights of Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Experimental results on mainstream language models show that Evolver outperforms previous state-of-the-art models by large margins due to the high training costs of large language models.
Approach: They propose a method to integrate multiple models from diverse training scenarios into a unified model.
Outcome: The proposed method outperforms state-of-the-art models on mainstream language models by large margins.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations