Challenge: Currently, there is no dataset containing compound and non-compound words across languages . however, current LLMs perform poorly on words tokenized unfavorably by subword tokenization.
Approach: They propose to use a Wiktionary dataset to evaluate large language models on decompounding . they find that current LLMs perform poorly on words tokenized unfavorably .
Outcome: The proposed model outperforms the best unsupervised models by 13.9% accuracy on average.

Similar Papers

Massively Translingual Compound Analysis and Translation Discovery (L18-1)

Copied to clipboard

Challenge: Morphological compounding is one of the most common and productive methods of word formation across the world's languages.
Approach: They propose a model for compounding using bilingual dictionaries and no annotated training data . they also release a massively multilingual dataset of compound words and their decompositions .
Outcome: The proposed model generates novel translations of English concepts on a multilingual dataset . the model can be applied to a wide range of languages and is highly reproducible.
Learn Your Tokens: Word-Pooled Tokenization for Language Modeling (2023.findings-emnlp)

Copied to clipboard

Challenge: Language models typically tokenize text into subwords, using a deterministic, hand-engineered heuristic of combining characters into longer surface-level strings such as ‘ing’ or whole words.
Approach: They propose a 'learn your tokens' scheme which pooles bytes/characters into word representations and decodes individual characters/bytes per word in parallel.
Outcome: The proposed tokenizer outperforms subword models and byte/character models over the word boundary and outperformed on rare words by a factor of 30!
To Split or Not to Split: Composing Compounds in Contextual Vector Spaces (2023.emnlp-main)

Copied to clipboard

Challenge: Contextual word embedding models rely on sub-word tokenization to represent single orthographic words but are often suboptimal in under-resourced contexts.
Approach: They propose to use a masked language modelling task to evaluate the model's performance . they use re-trained tokenizers to pre-split compounds into constituents .
Outcome: The proposed models improve on the masked language modelling task and compositionality prediction by pre-splitting compounds into constituents.
How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models (2021.acl-long)

Copied to clipboard

Challenge: Using pretraining data, we find that a designated monolingual tokenizer plays an equally important role in the downstream performance of the model.
Approach: They propose to compare pretrained multilingual models with their monolingual counterparts on a set of five diverse monolingual downstream tasks.
Outcome: The proposed models offer previously unmatched performance in all NLP tasks.
How do different tokenizers perform on downstream tasks in scriptio continua languages?: A case study in Japanese (2023.acl-srw)

Copied to clipboard

Challenge: Existing studies on scriptio continua languages lack comprehensiveness of tokenizers . authors use Byte-Pair-Encoding or Unigram instead of WordPiece for subword tokenizer .
Approach: They investigate the effect of tokenizers on the downstream performance of pretrained language models in scriptio continua languages where no explicit spaces exist between words.
Outcome: The proposed tokenizers perform better on a wide range of tasks compared with other tokenizer methods . the results show that each task has an optimal morphological analyzer .
How Important Is Tokenization in French Medical Masked Language Models? (2024.lrec-main)

Copied to clipboard

Challenge: Word tokenization into subword units has become the prevailing standard in the field of natural language processing (NLP) over recent years . the precise factors contributing to its success remain unclear .
Approach: They propose a tokenization strategy that integrates morpheme-enriched word segmentation into existing tokenization methods.
Outcome: The proposed tokenization strategy outperforms character and word tokenization but the precise factors contributing to its success remain unclear.
Beyond Text Compression: Evaluating Tokenizers Across Scales (2025.acl-long)

Copied to clipboard

Challenge: Language models rely on tokenizers to convert text into machine-interpretable tokens, which shape the statistical patterns that language models learn to estimate.
Approach: They propose to use Zipf's law to measure tokenizer performance by combining several metrics to capture multiple aspects of tokenizer behavior.
Outcome: The proposed metrics correlate more strongly with downstream performance than text compression when modeling unseen languages.
Where are we Still Split on Tokenization? (2024.findings-eacl)

Copied to clipboard

Challenge: Identifying tokens is a crucial first step for many tasks in Natural Language Processing (NLP) gold tokenization is often assumed, but some work on token-level tasks is more challenging.
Approach: They propose an efficient method for tokenization with subword-based language models and evaluate it on 122 languages in 20 scripts.
Outcome: The proposed method performs on par with the state-of-the-art on 122 languages in 20 scripts.
A Multi-dimensional Evaluation of Tokenizer-free Multilingual Pretrained Models (2023.findings-eacl)

Copied to clipboard

Challenge: Recent work on tokenizer-free models shows promising results in cross-lingual transfer . previous work focused on reporting accuracy on a limited set of tasks and data settings .
Approach: They compare tokenizer-free and subword-based models using various dimensions . they find subword models are still the most practical choice in many settings .
Outcome: The proposed model improves cross-lingual transfer and reduces engineering overhead.
Synthetic Pre-Training Tasks for Neural Machine Translation (2023.findings-acl)

Copied to clipboard

Challenge: toxicity and bias can be addressed by pre-training with synthetic resources . BLEU scores are used to compare methods with real-world data .
Approach: They propose several ways to generate obfuscated data from large parallel corpus and concatenating phrase pairs from small word-aligned corpus with synthetic parallel data without real human language corpora.
Outcome: The proposed methods can be used to generate obfuscated data or synthetic parallel data without real human language corpora even with high levels of oblication.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations