Papers by Huibin Ge
Chinese WPLC: A Chinese Dataset for Evaluating Pretrained Language Models on Word Prediction Given Long-Range Context (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing datasets for word prediction with long-range context have not been tested. |
| Approach: | They propose automatic and manual selection strategies tailored to Chinese to ensure that target words can only be predicted with long-term context. |
| Outcome: | The proposed model is 45 points behind human in terms of top-1 word prediction accuracy. |