Challenge: Large Language Models have proven successful at modelling tasks, but they are expensive and slow to scale.
Approach: They propose a Multi-Word Tokenizer that represents frequent multi-word expressions as single tokens.
Outcome: The proposed tokenizer is more robust across shorter sequence lengths, allowing for major speedups via early sequence truncation.

Similar Papers

Tokenization Is More Than Compression (2024.emnlp-main)

Copied to clipboard

Challenge: Existing tokenization approaches like Byte-Pair Encoding (BPE) have been suggested that their effectiveness stems from their ability to condense text into a relatively small number of tokens.
Approach: They propose a tokenizer that segments a document’s text into the minimum number of tokens for a given vocabulary and propose fewer tokens to improve downstream performance.
Outcome: The proposed tokenizers can initialize vocabulary construction and pre-tokenization, and the results show that fewer tokens lead to better performance.
LoPT: Lossless Parallel Tokenization Acceleration for Long Context Inference of Large Language Model (2026.acl-long)

Copied to clipboard

Challenge: Existing parallel tokenization methods suffer from inconsistent results due to boundary artifacts that occur after merging.
Approach: They propose a Lossless Parallel Tokenization framework that ensures output identical to standard sequential tokenization.
Outcome: The proposed method achieves significant speedup while guaranteeing lossless tokenization.
Learn Your Tokens: Word-Pooled Tokenization for Language Modeling (2023.findings-emnlp)

Copied to clipboard

Challenge: Language models typically tokenize text into subwords, using a deterministic, hand-engineered heuristic of combining characters into longer surface-level strings such as ‘ing’ or whole words.
Approach: They propose a 'learn your tokens' scheme which pooles bytes/characters into word representations and decodes individual characters/bytes per word in parallel.
Outcome: The proposed tokenizer outperforms subword models and byte/character models over the word boundary and outperformed on rare words by a factor of 30!
Pre-tokenization of Multi-word Expressions in Cross-lingual Word Embeddings (2020.emnlp-main)

Copied to clipboard

Challenge: Multi-Word Expressions (MWEs) are common in every language, but they are not translated by cross-lingual word embeddings.
Approach: They propose a method for word translation of Multi-Word Expressions (MWEs) they compile lists of MWEs in each language and tokenize them as single tokens before training word embeddings.
Outcome: The proposed method can translate multi-word expressions to and from English in 10 languages.
Investigating the Effectiveness of BPE: The Power of Shorter Sequences (D19-1)

Copied to clipboard

Challenge: Byte-Pair Encoding (BPE) is an unsupervised sub-word tokenization technique, but its reasons for its effectiveness are not well understood.
Approach: They link BPE to the broader family of dictionary-based compression algorithms and compare it with other members of this family.
Outcome: The proposed method is compared with dictionary-based compression algorithms and improves on a fixed vocabulary size budget.
Trainable, Multiword-aware Linguistic Tokenization Using Modern Neural Networks (2026.eacl-srw)

Copied to clipboard

Challenge: Tokenization is a fundamental task in natural language processing that forms the first step of many pipelines.
Approach: They propose to use a standard tokenizer trained without MWE-awareness as a baseline and a character-level SRN+CRF model to train token-level models.
Outcome: The proposed tokenizers are based on a character-level and token-level sequence labeling problem and are consistent with the proposed pipelines.
Tokenization Falling Short: On Subword Robustness in Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Language models typically tokenize raw text into sequences of subword identifiers from a predefined vocabulary.
Approach: They propose to tokenize raw text into sequences of subword identifiers from a predefined vocabulary . they also investigate the challenges and their impact on large language models .
Outcome: The proposed model can mitigate tokenization issues, but still suffer from typos and other variations.
Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem? (2025.findings-acl)

Copied to clipboard

Challenge: Multimodal large language models have shown remarkable performance for cross-modal understanding and generation, yet suffer from severe inference costs.
Approach: They propose to prune redundant tokens in MLLMs to reduce computation and storage costs.
Outcome: The proposed method reduces the computational and storage costs of MLLMs by identifying redundant tokens and pruning them.
Which Pieces Does Unigram Tokenization Really Need? (2026.findings-acl)

Copied to clipboard

Challenge: Despite its theoretical elegance, its implementation in practice is complex, limiting its adoption to SentencePiece.
Approach: They propose a Unigram-based probabilistic alternative to the greedy heuristics of Byte-Pair Encoding that is based on C++.
Outcome: The proposed algorithm is remarkably robust to hyperparameter choices and can be simplified to reduce computational costs.
Unpacking Tokenization: Evaluating Text Compression and its Correlation with Model Performance (2024.findings-acl)

Copied to clipboard

Challenge: Despite being the cornerstone of BPE, the importance of compression in the tokenization process is still unclear.
Approach: They argue for the theoretical importance of compression in the tokenization process . they also demonstrate the empirical importance of compressing tokenizers for downstream success of pre-trained language models.
Outcome: The proposed method can be viewed as 0-gram language modeling where equal probability is assigned to all tokens.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations