Papers by Hongao Zhu
A Systematic Assessment of Language Models with Linguistic Minimal Pairs in Chinese (2026.tacl-1)
Copied to clipboard
Yikang Liu, Yeting Shen, Hongao Zhu, Lilong Xu, Zhiheng Qian, Siyuan Song, Kejia Zhang, Jialong Tang, Pei Zhang, Baosong Yang, Rui Wang, Hai Hu
| Challenge: | Using sub-linear length normalized log-probabilities (SLLN-LP), we find unequal lengths of sentences in minimal pairs difficult for LMs even up to 32B parameters. |
| Approach: | They propose to use ZhoBLiMP as a linguistic minimal pair benchmark for Chinese language models to mitigate biases. |
| Outcome: | The proposed metric mitigates biases in Chinese language models with over 100 paradigms . Anaphor, Quantifiers, and Ellipsis are difficult for LMs even up to 32B parameters . |
The Inverse Scaling Effect of Pre-Trained Language Model Surprisal Is Not Due to Data Leakage (2025.findings-acl)
Copied to clipboard
| Challenge: | Language models (LMs) have been shown to flexibly capture many linguistic regularities from raw text, but the source stimuli of reading time datasets are often naturalistic text that are available online. |
| Approach: | They propose to replicate the negative relationship between language model size and the fit of surprisal to reading times using models trained on ‘leakage-free’ data that overlaps only minimally with the reading time corpora. |
| Outcome: | The proposed models show that language models trained on 'leakage-free' data are not driven by data leakage. |