Challenge: Pretrained sequence-to-sequence (seq2sequ) models have been widely used to solve extractive tasks, where parts of the input are extracted to form the desired output.
Approach: They propose a simple fix to tokenization inconsistency that damages extractive nature of generative models by causing performance drop and hallucination.
Outcome: The proposed model performs better in both in-domain and out-of-domain datasets with a notable average of +1.7 F1 gain when a BART model is trained on SQuAD and evaluated on 8 QA datasets.

Similar Papers

Token Alignment via Character Matching for Subword Completion (2024.findings-acl)

Copied to clipboard

Challenge: Generative models struggle with prompts corresponding to partial tokens due to tokenization, where partial token is out-of-distribution during inference.
Approach: They propose a method to alleviate tokenization artifact on text completion by backtracking to the last complete tokens and aligning subsequent generations to match with the prompt.
Outcome: The proposed method shows that it improves on partial token scenarios with only a minor time increase.
Where are we Still Split on Tokenization? (2024.findings-eacl)

Copied to clipboard

Challenge: Identifying tokens is a crucial first step for many tasks in Natural Language Processing (NLP) gold tokenization is often assumed, but some work on token-level tasks is more challenging.
Approach: They propose an efficient method for tokenization with subword-based language models and evaluate it on 122 languages in 20 scripts.
Outcome: The proposed method performs on par with the state-of-the-art on 122 languages in 20 scripts.
Tokenization Falling Short: On Subword Robustness in Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Language models typically tokenize raw text into sequences of subword identifiers from a predefined vocabulary.
Approach: They propose to tokenize raw text into sequences of subword identifiers from a predefined vocabulary . they also investigate the challenges and their impact on large language models .
Outcome: The proposed model can mitigate tokenization issues, but still suffer from typos and other variations.
Striking a Balance: Alleviating Inconsistency in Pre-trained Models for Symmetric Classification Tasks (2022.findings-acl)

Copied to clipboard

Challenge: Inconsistency is observed in symmetric classification tasks that take two inputs and require the output to be invariant of the order of the inputs.
Approach: They propose a consistency loss function to alleviate inconsistency in symmetric classification tasks that take two inputs and require the output to be invariant of the order of the inputs.
Outcome: The proposed model improves consistency in predictions for three paraphrase detection datasets without significant drop in accuracy scores.
Tokenization and Representation Biases in Multilingual Models on Dialectal NLP Tasks (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) pre-trained on massive text data in many languages are preferred solution for various Natural Language processing tasks.
Approach: They compare tokenization parity and information parity as representational biases in pre-trained models . they find TP is better predictor of performance on tasks reliant on syntactic and morphological cues .
Outcome: The proposed model improves on dialect classification, topic classification, and extractive question answering tasks.
From Where Words Come: Efficient Regularization of Code Tokenizers Through Source Attribution (2026.acl-long)

Copied to clipboard

Challenge: Currently, subword tokenization is the most common approach for vocabulary building in large models.
Approach: They propose to regularize training and minimize overfitting by using source-attributed BPE . they find that undertrained tokens are prone to producing unused, unusable tokens .
Outcome: The proposed techniques reduce the number of under-trained tokens while maintaining the same inference procedure as with regular BPE.
Improving Consistency in LLM Inference using Probabilistic Tokenization (2025.findings-naacl)

Copied to clipboard

Challenge: Prior work has shown that probabilistic tokenizations can generate multiple tokenization of the same input string.
Approach: They propose a method to leverage the multiple tokenization capabilities of modern LLM tokenizers.
Outcome: The proposed method improves the self-consistency of large language models by generating multiple tokenizations.
Separating Retention from Extraction in the Evaluation of End-to-end Relation Extraction (2021.emnlp-main)

Copied to clipboard

Challenge: State-of-the-art NLP models adopt shallow heuristics that limit their generalization capability.
Approach: They propose to use heuristics that limit their generalization capability to model lexical overlap with the training set in Named-Entity Recognition and Event or Type heuristic in Relation Extraction to test their models.
Outcome: The proposed model can perform better on the two key tasks, while the retention of training relation triples.
Consistency Regularization Training for Compositional Generalization (2023.acl-long)

Copied to clipboard

Challenge: Existing neural models have difficulty generalizing to unseen combinations of seen components.
Approach: They propose to improve the capability of Transformer on compositional generalization by consistency regularization training without modifying model architectures.
Outcome: The proposed model performs well on semantic parsing and machine translation benchmarks.
FLEXITOKENS: Flexible Tokenization for Evolving Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Widely used subword tokenizers overfragment sequences in unseen domains, languages, and scripts . inefficient tokenizer models can cause overfragments in out-of-distribution domains if not trained properly .
Approach: They propose a byte-level LM with learnable tokenizers to make tokenization adaptive . they propose 'flexitoken' which enables significantly greater flexibility during adaptation .
Outcome: The proposed method significantly reduces token overfragmentation and improves on multilingual benchmarks and domains.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations