Challenge: Existing studies have reported that an appropriate tokenization depends on each downstream task.
Approach: They propose a method to find an appropriate tokenization to a downstream task by optimizing a tokenizer and a model.
Outcome: The proposed method improves on text classification and machine translation tasks.

Similar Papers

Optimizing Word Segmentation for Downstream Task (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to optimize tokenizations for downstream tasks are not suitable for traditional NLP.
Approach: They propose a method to explore a tokenization appropriate for a downstream task . they train a model to assign a high probability to such appropriate tokenization based on the downstream task loss .
Outcome: The proposed method improves sentiment analysis and textual entailment tasks . it is also integrated into state-of-the-art contextualized embeddings and reports a positive effect .
Where are we Still Split on Tokenization? (2024.findings-eacl)

Copied to clipboard

Challenge: Identifying tokens is a crucial first step for many tasks in Natural Language Processing (NLP) gold tokenization is often assumed, but some work on token-level tasks is more challenging.
Approach: They propose an efficient method for tokenization with subword-based language models and evaluate it on 122 languages in 20 scripts.
Outcome: The proposed method performs on par with the state-of-the-art on 122 languages in 20 scripts.
Tokenizer Choice For LLM Training: Negligible or Crucial? (2024.findings-naacl)

Copied to clipboard

Challenge: Recent success of large language models has been driven by curating the training dataset composition, scaling of model architectures and advancements in pretraining objectives, leaving tokenizer influence as a blind spot.
Approach: They conduct a comprehensive study on the influence of tokenizer choice on LLM downstream performance by training 24 mono- and multilingual LLMs at a 2.6B parameter scale.
Outcome: The proposed model can significantly impact the model's downstream performance and training costs.
Tokenization is Sensitive to Language Variation (2025.findings-acl)

Copied to clipboard

Challenge: Variation in language is often linked to regional, social, and contextual factors.
Approach: They propose a method to estimate tokenizer impact on downstream LLM performance . they pre-train BERT models with the popular Byte-Pair Encoding algorithm .
Outcome: The proposed model improves on Rényi efficiency and other metrics on language variation.
How do different tokenizers perform on downstream tasks in scriptio continua languages?: A case study in Japanese (2023.acl-srw)

Copied to clipboard

Challenge: Existing studies on scriptio continua languages lack comprehensiveness of tokenizers . authors use Byte-Pair-Encoding or Unigram instead of WordPiece for subword tokenizer .
Approach: They investigate the effect of tokenizers on the downstream performance of pretrained language models in scriptio continua languages where no explicit spaces exist between words.
Outcome: The proposed tokenizers perform better on a wide range of tasks compared with other tokenizer methods . the results show that each task has an optimal morphological analyzer .
How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models (2021.acl-long)

Copied to clipboard

Challenge: Using pretraining data, we find that a designated monolingual tokenizer plays an equally important role in the downstream performance of the model.
Approach: They propose to compare pretrained multilingual models with their monolingual counterparts on a set of five diverse monolingual downstream tasks.
Outcome: The proposed models offer previously unmatched performance in all NLP tasks.
Unpacking Tokenization: Evaluating Text Compression and its Correlation with Model Performance (2024.findings-acl)

Copied to clipboard

Challenge: Despite being the cornerstone of BPE, the importance of compression in the tokenization process is still unclear.
Approach: They argue for the theoretical importance of compression in the tokenization process . they also demonstrate the empirical importance of compressing tokenizers for downstream success of pre-trained language models.
Outcome: The proposed method can be viewed as 0-gram language modeling where equal probability is assigned to all tokens.
BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training (2024.emnlp-main)

Copied to clipboard

Challenge: Tokenization is a relatively understudied area, but it can greatly impact model performance and efficiency.
Approach: They propose a modified BPE tokenizer that removes merges that leave intermediate "junk" tokens from the vocabulary.
Outcome: The proposed method improves vocabulary efficiency, eliminates under-trained tokens, and does not compromise text compression.
Stop Taking Tokenizers for Granted: They Are Core Design Decisions in Large Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Subword tokenization approaches misalign with linguistic structure and waste capacity across languages and domains.
Approach: They argue for a context-aware framework that integrates tokenizer and model co-design . they argue that tokenization should be treated as a core design problem, not an afterthought .
Outcome: The proposed framework integrates tokenizer and model co-design, guided by linguistic, domain, and deployment considerations.
Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization (2026.acl-long)

Copied to clipboard

Challenge: Tokenization is the first step of most NLP pipelines.
Approach: They propose a parity-aware byte pair encoder that maximizes the compression gain of the currently worst-compressed language for cross-lingual parity.
Outcome: a new algorithm reduces tokenization inequality by 89% compared to classical BPE . the proposed algorithm is based on a fair-max rule that maximizes the compression gain of the currently worst-compressed language .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations