Papers by Yanjin He

1 papers
Pre-trained Models Perform the Best When Token Distributions Follow Zipf’s Law (2025.emnlp-main)

Copied to clipboard

Challenge: Existing large language models typically fix a vocabulary size in advance, then use Byte Pair Encoding (BPE) to construct the tokenizer.
Approach: They propose a method for determining the vocabulary size by analyzing token frequency distributions through Zipf’s law and propose to use it to optimize model performance.
Outcome: The proposed method improves model efficiency and effectiveness across NLP, genomics, and chemistry.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations