Papers by Wenxiang Geng
The Paradox of Outcome Optimization: A Causal Information-Theoretic Bound on Reasoning Shortcuts in LLMs (2026.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) aligned via outcome-based Reinforcement Learning (RL) exhibit a critical failure mode: they exhibit brittle reasoning capabilities on out-of-distribution tasks. |
| Approach: | They propose a framework bridging Structural Causal Models and the Information Bottleneck principle to explain this paradox. |
| Outcome: | The proposed framework bridges the framework between SCM and IB principles to explain the problem. |