Papers by Yonghong He
Bridging SFT and RL: Dynamic Policy Optimization for Robust Reasoning (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing unified optimization strategies overlook the statistical conflict between these distinct gradient signals. |
| Approach: | They propose a framework to reduce bias-variance trade-offs in Large Language Models . they propose DYPO, which leverages intrinsic group dynamics to significantly reduce RL gradient variance . |
| Outcome: | The proposed framework outperforms traditional pipelines on reasoning benchmarks and out-of-distribution tasks. |
Discriminating between Similar Languages on Imbalanced Conversational Texts (L18-1)
Copied to clipboard
| Challenge: | Empirical results suggest that our system achieves an accuracy of 95.7% on our Uyghur and Kazakh dataset, which is higher than that of the CNN classifier. |
| Approach: | They propose to build a balanced Uyghur and Kazakh corpus and build morphological classifiers to discriminate between the two languages. |
| Outcome: | The proposed system outperforms the champions on both test sets B1 and B2. |