Challenge: Existing subword segmentation tools assume input is pre-tokenized into word sequences, but SentencePiece can train subword models directly from raw sentences.
Approach: They propose a language-independent subword tokenizer and detokenizer for Neural-based text processing.
Outcome: The proposed system achieves comparable accuracy to training from raw sentences.

Similar Papers

Bilingual Subword Segmentation for Neural Machine Translation (2020.coling-main)

Copied to clipboard

Challenge: Existing subword segmentation methods tokenize sentences without considering translation . proposed method could be more favorable to machine translation if it uses bilingual sentences .
Approach: They propose a subword segmentation method that tokenizes sentences by using subword units induced from bilingual sentences.
Outcome: The proposed method improves translation performance on translation tasks up to +0.81 BLEU.
Neural Machine Translation without Embeddings (2021.naacl-main)

Copied to clipboard

Challenge: Existing models operate over subword tokens, but byte-based models employ a different approach . a one-hot representation of each byte does not hurt performance, but it improves BLEU scores .
Approach: They propose to represent every computerized text as a sequence of bytes via UTF-8 . this eliminates the need for an embedding layer and improves performance .
Outcome: The proposed model improves BLEU scores on byte-to-byte translation models compared to character-level models . the proposed model does not require an embedding layer and does not drop out of the decoder .
MaxMatch-Dropout: Subword Regularization for WordPiece (2022.coling-1)

Copied to clipboard

Challenge: Existing subword regularization methods are specialized to a particular tokenizer type.
Approach: They propose a subword regularization method for WordPiece that uses a maximum matching algorithm for tokenization.
Outcome: The proposed method improves the performance of text classification and machine translation tasks as well as other subword regularization methods.
How Important Is Tokenization in French Medical Masked Language Models? (2024.lrec-main)

Copied to clipboard

Challenge: Word tokenization into subword units has become the prevailing standard in the field of natural language processing (NLP) over recent years . the precise factors contributing to its success remain unclear .
Approach: They propose a tokenization strategy that integrates morpheme-enriched word segmentation into existing tokenization methods.
Outcome: The proposed tokenization strategy outperforms character and word tokenization but the precise factors contributing to its success remain unclear.
Local Byte Fusion for Neural Machine Translation (2023.acl-long)

Copied to clipboard

Challenge: Existing NLP models rely on a pre-built subword tokenizer to tokenize a sentence . this can be rigid and subwords from low-resource languages are under-represented .
Approach: They propose a method for byte-based machine translation that aggregates local semantic information.
Outcome: The proposed method improves on multilingual translation and cross-lingual transfer . it is parameter-efficient and performs competitively to subword models, it is shown .
Fast WordPiece Tokenization (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for tokenization of text are not efficient, but they are based on Aho-Corasick's algorithm.
Approach: They propose an efficient algorithm for WordPiece tokenization using a longest-match-first strategy . they propose an algorithm whose tokenization complexity is strictly O(n)
Outcome: The proposed method is 8.2x faster than HuggingFace Tokenizers and 5.1x faster on average for general text tokenization.
Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates (P18-1)

Copied to clipboard

Challenge: Subword units are an effective way to alleviate the open vocabulary problems in neural machine translation.
Approach: They propose a method to regularize subword segmentations probabilistically by sampling subwords . they also propose 'unigram' language model to be used for better subword sampling .
Outcome: The proposed method improves on low resource and out-of-domain settings with multiple corpora.
Treepiece: Faster Semantic Parsing via Tree Tokenization (2023.findings-emnlp)

Copied to clipboard

Challenge: Autoregressive (AR) encoder-decoder neural networks are slow in sequential prediction of natural language to machine-readable parse trees.
Approach: They propose a technique that tokenizes a parse tree into subtrees and generates one subtrea per decoding step.
Outcome: The proposed approach shows 4.6 times faster decoding speed and comparable speed but significantly higher accuracy compared to non-autoregressive (NAR) models.
CompoundPiece: Evaluating and Improving Decompounding Performance of Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Currently, there is no dataset containing compound and non-compound words across languages . however, current LLMs perform poorly on words tokenized unfavorably by subword tokenization.
Approach: They propose to use a Wiktionary dataset to evaluate large language models on decompounding . they find that current LLMs perform poorly on words tokenized unfavorably .
Outcome: The proposed model outperforms the best unsupervised models by 13.9% accuracy on average.
Should we find another model?: Improving Neural Machine Translation Performance with ONE-Piece Tokenization Method without Model Modification (2021.naacl-industry)

Copied to clipboard

Challenge: Recent studies using pretrain-finetuning approach have achieved state-of-the-art (SOTA) performance in many natural language processing tasks.
Approach: They propose a new tokenization method that combines morphology-considered subword tokenization and vocabulary methods to address this limitation.
Outcome: The proposed method can be used without modifying the model structure.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations