Papers by Xiaoyuan Wu

2 papers
User Perceptions vs. Proxy LLM Judges: Privacy and Helpfulness in LLM Responses to Privacy-Sensitive Scenarios (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are rapidly being adopted for tasks like drafting emails, summarizing meetings, and answering health questions.
Approach: They conducted a scenario-based evaluation of Large language models (LLMs) using 90 PrivacyLens scenarios.
Outcome: The proposed models can leak private information in complex scenarios, but they do not measure user perceptions directly.
Estimating LLM Consistency: A User Baseline vs Surrogate Metrics (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are prone to hallucinations and sensitive to prompt perturbations, resulting in inconsistent or unreliable generated text.
Approach: They propose a logit-based ensemble method to measure LLM consistency and propose to use it to evaluate human ratings of LLM reliability.
Outcome: The proposed method matches the best-performing existing metric in estimating human ratings of LLM consistency.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations