Dissecting Transformer Length Extrapolation via the Lens of Receptive Field Analysis (2023.acl-long)
Copied to clipboard
| Challenge: | Length extrapolation allows training a transformer language model on short sequences that preserves perplexities when tested on substantially longer sequences. |
| Approach: | They propose a relative positional embedding design that uses longer than the training sequence to create sandwich. |
| Outcome: | The proposed model can extrapolate to L ex L tr much better than other models. |
Similar Papers
Attention Alignment and Flexible Positional Embeddings Improve Transformer Length Extrapolation (2024.findings-naacl)
Copied to clipboard
| Challenge: | Existing methods for length extrapolation are tailored for natural language modeling, a task known to have strong recency bias. |
| Approach: | They propose two attention alignment strategies to improve T5's long-context utilization capability without fine-tuning. |
| Outcome: | The proposed methods improve the long-context utilization capability of T5 on language modeling, retrieval, multi-document question answering, and code completion tasks without any fine-tuning. |
Length Extrapolation of Transformers: A Survey from the Perspective of Positional Encoding (2024.findings-emnlp)
Copied to clipboard
Liang Zhao, Xiachong Feng, Xiaocheng Feng, Weihong Zhong, Dongliang Xu, Qing Yang, Hongtao Liu, Bing Qin, Ting Liu
| Challenge: | Existing methods to enhance length extrapolation of large language models have been developed, but a systematic survey is lacking. |
| Approach: | They propose to examine the effects of positional encoding on length extrapolation. |
| Outcome: | The proposed methods improve the extrapolation of large language models, but they are still lacking a systematic survey. |
A Length-Extrapolatable Transformer (2023.acl-long)
Copied to clipboard
Yutao Sun, Li Dong, Barun Patra, Shuming Ma, Shaohan Huang, Alon Benhaim, Vishrav Chaudhary, Xia Song, Furu Wei
| Challenge: | Existing Transformers can only deal with the in-distribution size of inputs. |
| Approach: | They propose a relative position embedding to explicitly maximize attention resolution . they also use blockwise causal attention during inference for better resolution a . |
| Outcome: | The proposed model achieves strong performance in interpolation and extrapolation settings. |
Context-aware Biases for Length Extrapolation (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for Relative Positional Encoding (RPE) lack the capacity to adapt to different input contexts. |
| Approach: | They propose an additive RPE method that learns token-specific, context-aware biases for each attention head in transformers by dynamically adjusting positional biase based on the input sequence. |
| Outcome: | The proposed method significantly improves the extrapolation performance of existing RPE methods on the fineWeb-Edu-10B and WikiText-103 datasets. |
Context Length Extension via Generalized Extrapolation Scale (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing work on extrapolating positional embedding (RoPE) has limited results in the application of long context language models. |
| Approach: | They propose a set of parameterized extrapolation functions applied to each layer and attention head to adaptively adjust its extrapolations scales. |
| Outcome: | The proposed model achieves stable extrapolation on 64k contexts by training on 16k length text. |
Position Encoding with Random Float Sampling Enhances Length Generalization of Transformers (2026.findings-eacl)
Copied to clipboard
| Challenge: | Length generalization is the ability of language models to maintain performance on inputs longer than those seen during pretraining. |
| Approach: | They propose a position encoding strategy that uses random float sampling to generalize to unseen lengths. |
| Outcome: | The proposed strategy can generalize to lengths unseen during training and in benchmarks. |
ETC: Encoding Long and Structured Inputs in Transformers (2020.emnlp-main)
Copied to clipboard
Joshua Ainslie, Santiago Ontanon, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, Li Yang
| Challenge: | Existing models for natural language processing (NLP) have been challenging to scale attention to longer inputs. |
| Approach: | They propose an extended Transformer construction architecture that scales attention to longer inputs by combining global-local attention with relative position encodings and a "Contrastive Predictive Coding" objective. |
| Outcome: | The proposed architecture scales attention to longer inputs and encodes structured inputs. |
Shortformer: Better Language Modeling using Shorter Inputs (2021.acl-long)
Copied to clipboard
| Challenge: | Existing methods require computationally expensive relative position embeddings. |
| Approach: | They propose two methods that decrease input length to improve perplexity and perplexability. |
| Outcome: | The proposed methods speed up training by a factor of 1.65 and reduce memory usage. |
Improve Transformer Models with Better Relative Position Embeddings (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for generating position embeddings are not fully utilized in NLP tasks. |
| Approach: | They propose to generalize the absolute position embedding to a generalized relative position embedded method . they also propose to use the relative embeddable method to improve the accuracy of large models . |
| Outcome: | The proposed method improves accuracy on the SQuAD1.1 dataset compared to previous methods . it can be easily adopted as a drop-in replacement for improving accuracy of large models . |
Explore Better Relative Position Embeddings from Encoding Perspective for Transformer Models (2021.emnlp-main)
Copied to clipboard
| Challenge: | Experimental results on nine authoritative datasets demonstrate the effectiveness of our methods empirically. |
| Approach: | They propose to use a method to explicitly encode position information into Transformer models by using absolute position embedding. |
| Outcome: | The proposed methods improve Shaw-RPE and XL-R PE and achieve the best overall performance among five different RPEs. |