Papers by Yunqi Ba
DecorateLM: Data Engineering through Corpus Rating, Tagging, and Editing with Language Models (2024.emnlp-main)
Copied to clipboard
Ranchi Zhao, Zhen Thai, Yifan Zhang, Shengding Hu, Jie Zhou, Yunqi Ba, Jie Cai, Zhiyuan Liu, Maosong Sun
| Challenge: | Large Language Models (LLMs) are pre-trained on vast datasets composed of billions of tokens harvested from diverse text sources. |
| Approach: | They propose a data engineering method to refine the pretraining corpus through data rating, tagging and editing. |
| Outcome: | The proposed method improves the quality of the pretraining corpus by enhancing 100 billion tokens of the training corpus. |