Papers by Tim Genewein
Randomized Positional Encodings Boost Length Generalization of Transformers (2023.acl-short)
Copied to clipboard
Anian Ruoss, Grégoire Delétang, Tim Genewein, Jordi Grau-Moya, Róbert Csordás, Mehdi Bennani, Shane Legg, Joel Veness
| Challenge: | Moreover, simply training on longer sequences is inefficient due to the quadratic computation complexity of the global attention mechanism. |
| Approach: | They propose a randomized positional encoding scheme that randomly selects an ordered subset to fit the sequence’s length. |
| Outcome: | The proposed method allows Transformers to generalize to sequences of unseen length (increasing test accuracy by 12.0% on average). |