Challenge: Using attention-based models, certain tokens are less ambiguous than others, and they require fewer refinements for disambiguation.
Approach: They propose a lazy transition mechanism to adjust the significance of iterative refinements for each token representation.
Outcome: The proposed model outperforms baseline models on several tasks with the same number of parameters.

Similar Papers

Deep Equilibrium Non-Autoregressive Sequence Learning (2023.findings-acl)

Copied to clipboard

Challenge: et al., 2017) is the most prevailing neural architecture for sequence-to-sequence learning.
Approach: They propose to solve for the equilibrium state of NAR models with black-box root-finding solvers and back-propagate through the equilibrium point via implicit differentiation with constant memory.
Outcome: The proposed framework can converge to a more accurate prediction on four WMT benchmarks.
FLEXITOKENS: Flexible Tokenization for Evolving Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Widely used subword tokenizers overfragment sequences in unseen domains, languages, and scripts . inefficient tokenizer models can cause overfragments in out-of-distribution domains if not trained properly .
Approach: They propose a byte-level LM with learnable tokenizers to make tokenization adaptive . they propose 'flexitoken' which enables significantly greater flexibility during adaptation .
Outcome: The proposed method significantly reduces token overfragmentation and improves on multilingual benchmarks and domains.
Exact Hard Monotonic Attention for Character-Level Transduction (P19-1)

Copied to clipboard

Challenge: Neural sequence-to-sequence models with soft attention outperform monotonic models . current dominant method is the neural sequenceto-Sequency model with soft focus .
Approach: They develop a hard attention sequence-to-sequence model that enforces strict monotonicity and learns alignment jointly.
Outcome: The proposed model achieves state-of-the-art on grapheme-to-phoneme conversion and morphological inflection generation.
Lazy-k Decoding: Constrained Decoding for Information Extraction (2023.emnlp-main)

Copied to clipboard

Challenge: Specifically, we combine probabilistic models with constrained decoding approaches in structured prediction tasks.
Approach: They propose a constrained decoding method called Lazy-k to combine probabilistic models with constrained methods in structured prediction.
Outcome: The proposed method allows for more flexibility between decoding time and accuracy.
Prefix Propagation: Parameter-Efficient Tuning for Long Sequences (2023.acl-short)

Copied to clipboard

Challenge: Prefix-tuning prepends trainable tokens to sequences while freezing the rest of the model’s parameters.
Approach: They propose a method that prefixes on previous hidden states to improve model performance.
Outcome: The proposed architecture outperforms prefix-tuning on long-document tasks while using 50% fewer parameters.
Token-Level Self-Evolution Training for Sequence-to-Sequence Learning (2023.acl-short)

Copied to clipboard

Challenge: Adaptive training approaches do not consider the variation of learning difficulty in different training steps, making the learning deterministic and sub-optimal.
Approach: They propose a dynamic token-level self-evolution training method that reweighs the training losses of different target tokens based on priors.
Outcome: Empirically, the proposed method yields significant improvements on three translation tasks.
One Token Is Enough: Improving Diffusion Language Models with a Sink Token (2026.findings-acl)

Copied to clipboard

Challenge: Existing Diffusion Language Models lack a structural constraint to stabilize attention sinks.
Approach: They propose a simple but effective extra sink token that is constrained to attend to itself while remaining globally visible to all other tokens.
Outcome: The proposed token is able to stabilize attention sinks and improve model performance.
Learning to Insert [PAUSE] Tokens for Better Reasoning (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies have explored incorporating special-purpose tokens into the training process to enhance reasoning capabilities.
Approach: They propose a method for inserting dummy tokens consecutively just before reasoning steps to increase model effectiveness.
Outcome: The proposed method outperforms fine-tuning and previous token insertion methods on multiple datasets and models.
Stability Implies Redundancy: Delta Attention Selective Halting for Efficient Long-Context Prefilling (2026.acl-long)

Copied to clipboard

Challenge: Existing methods to reduce sequence length rely on heuristics that break compatibility with hardware-efficient kernels like FlashAttention.
Approach: They propose a method that selectively halts stabilized tokens by monitoring layer-wise update dynamics of the self-attention mechanism.
Outcome: The proposed method can reduce prefill complexity while preserving model accuracy and hardware efficiency.
SepSeq: A Training-Free Framework for Long Numerical Sequence Processing in LLMs (2026.findings-acl)

Copied to clipboard

Challenge: Existing large-scale large-context models suffer from performance degradation when processing long numerical sequences.
Approach: They propose a framework to mitigate attention dispersion by strategically inserting separator tokens into the model to recalibrat attention to local segments while preserving global context.
Outcome: The proposed framework improves accuracy and reduces inference token consumption by 16.4% on 9 widely-adopted LLMs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations