Challenge: Length extrapolation allows training a transformer language model on short sequences that preserves perplexities when tested on substantially longer sequences.
Approach: They propose a relative positional embedding design that uses longer than the training sequence to create sandwich.
Outcome: The proposed model can extrapolate to L ex L tr much better than other models.

Similar Papers

Attention Alignment and Flexible Positional Embeddings Improve Transformer Length Extrapolation (2024.findings-naacl)

Copied to clipboard

Challenge: Existing methods for length extrapolation are tailored for natural language modeling, a task known to have strong recency bias.
Approach: They propose two attention alignment strategies to improve T5's long-context utilization capability without fine-tuning.
Outcome: The proposed methods improve the long-context utilization capability of T5 on language modeling, retrieval, multi-document question answering, and code completion tasks without any fine-tuning.
Length Extrapolation of Transformers: A Survey from the Perspective of Positional Encoding (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to enhance length extrapolation of large language models have been developed, but a systematic survey is lacking.
Approach: They propose to examine the effects of positional encoding on length extrapolation.
Outcome: The proposed methods improve the extrapolation of large language models, but they are still lacking a systematic survey.
A Length-Extrapolatable Transformer (2023.acl-long)

Copied to clipboard

Challenge: Existing Transformers can only deal with the in-distribution size of inputs.
Approach: They propose a relative position embedding to explicitly maximize attention resolution . they also use blockwise causal attention during inference for better resolution a .
Outcome: The proposed model achieves strong performance in interpolation and extrapolation settings.
Context-aware Biases for Length Extrapolation (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for Relative Positional Encoding (RPE) lack the capacity to adapt to different input contexts.
Approach: They propose an additive RPE method that learns token-specific, context-aware biases for each attention head in transformers by dynamically adjusting positional biase based on the input sequence.
Outcome: The proposed method significantly improves the extrapolation performance of existing RPE methods on the fineWeb-Edu-10B and WikiText-103 datasets.
Context Length Extension via Generalized Extrapolation Scale (2024.findings-acl)

Copied to clipboard

Challenge: Existing work on extrapolating positional embedding (RoPE) has limited results in the application of long context language models.
Approach: They propose a set of parameterized extrapolation functions applied to each layer and attention head to adaptively adjust its extrapolations scales.
Outcome: The proposed model achieves stable extrapolation on 64k contexts by training on 16k length text.
Position Encoding with Random Float Sampling Enhances Length Generalization of Transformers (2026.findings-eacl)

Copied to clipboard

Challenge: Length generalization is the ability of language models to maintain performance on inputs longer than those seen during pretraining.
Approach: They propose a position encoding strategy that uses random float sampling to generalize to unseen lengths.
Outcome: The proposed strategy can generalize to lengths unseen during training and in benchmarks.
ETC: Encoding Long and Structured Inputs in Transformers (2020.emnlp-main)

Copied to clipboard

Challenge: Existing models for natural language processing (NLP) have been challenging to scale attention to longer inputs.
Approach: They propose an extended Transformer construction architecture that scales attention to longer inputs by combining global-local attention with relative position encodings and a "Contrastive Predictive Coding" objective.
Outcome: The proposed architecture scales attention to longer inputs and encodes structured inputs.
Shortformer: Better Language Modeling using Shorter Inputs (2021.acl-long)

Copied to clipboard

Challenge: Existing methods require computationally expensive relative position embeddings.
Approach: They propose two methods that decrease input length to improve perplexity and perplexability.
Outcome: The proposed methods speed up training by a factor of 1.65 and reduce memory usage.
Improve Transformer Models with Better Relative Position Embeddings (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for generating position embeddings are not fully utilized in NLP tasks.
Approach: They propose to generalize the absolute position embedding to a generalized relative position embedded method . they also propose to use the relative embeddable method to improve the accuracy of large models .
Outcome: The proposed method improves accuracy on the SQuAD1.1 dataset compared to previous methods . it can be easily adopted as a drop-in replacement for improving accuracy of large models .
Explore Better Relative Position Embeddings from Encoding Perspective for Transformer Models (2021.emnlp-main)

Copied to clipboard

Challenge: Experimental results on nine authoritative datasets demonstrate the effectiveness of our methods empirically.
Approach: They propose to use a method to explicitly encode position information into Transformer models by using absolute position embedding.
Outcome: The proposed methods improve Shaw-RPE and XL-R PE and achieve the best overall performance among five different RPEs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations