Papers by Yunhao Yao

2 papers
RemoteRAG: A Privacy-Preserving LLM Cloud RAG Service (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have a tendency to generate factually incorrect or purely fictional responses, a phenomenon known as hallucination.
Approach: They propose to use remote RAG to protect user query from privacy leakage . they introduce (n,)-DistanceDP to characterize privacy leakages of user query .
Outcome: The proposed solution can resist embedding inversion attacks while achieving no loss in retrieval under various settings.
LightMoE: Task-Aware Expert Availability Management for Memory-Efficient MoE-LLM Inference (2026.findings-acl)

Copied to clipboard

Challenge: Existing solutions for balancing model accuracy with inference latency are limited due to memory constraints.
Approach: They propose a framework for memory-efficient MoE inference that exploits the functional redundancy and temporal locality of expert activation.
Outcome: The proposed framework improves accuracy-efficiency trade-off by 4.3% over pruned models and 2.4% over dynamic swapping methods while maintaining inference latency comparable to pruned model.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations