Papers by Chenxu Zhao

2 papers
Lil: Less is Less When Applying Post-Training Sparse-Attention Algorithms in Long-Decode Stage (2026.findings-acl)

Copied to clipboard

Challenge: Prior work typically decomposes inference into prefill and decode stages, with the decode stage dominating total latency.
Approach: They propose an algorithm that detects threshold where information loss exceeds information gain during sparse decoding to reduce token consumption by up to 90% and a marginal accuracy degradation of less than 2%.
Outcome: The proposed algorithm reduces token consumption by 90% with a marginal accuracy degradation of less than 2% across reasoning-intensive benchmarks.
Quantifying and Understanding Uncertainty in Large Reasoning Models (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for estimating generation uncertainty do not provide finite-sample guarantees for reasoning-answer generation.
Approach: They propose a method that provides the uncertainty of the reasoning-answer structure with statistical guarantees.
Outcome: The proposed method disentangles reasoning quality from answer correctness while establishing theoretical guarantees for efficient explanation methods.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations