| Challenge: | Sequence-to-sequence models are dense and assigning nonzero probability to implausible outputs. |
| Approach: | They propose a new family of -entmax transformations that includes softmax and sparsemax as particular cases and is sparser for any > 1 . they provide fast algorithms to evaluate these transformations and their gradients, which scale well for large vocabulary sizes. |
| Outcome: | The proposed models are able to produce sparse alignments and assign nonzero probability to short list of plausible outputs, sometimes rendering beam search exact. |
Similar Papers
Smoothing and Shrinking the Sparse Seq2Seq Search Space (2021.naacl-main)
Copied to clipboard
| Challenge: | entmax-based sparse sequence-to-sequence models give high scores to short hypotheses . ent max models can shrink the search space by assigning zero probability to bad hypothese . |
| Approach: | They propose entmax-based sparse sequence-to-sequence models that minimize cross-entropy and use softmax to compute local normalized probabilities over target sequences. |
| Outcome: | The proposed models remove a major source of model error for word-level tasks . the proposed models improve cross-lingual morphological inflection and machine translation . |
Speeding Up Entmax (2022.findings-naacl)
Copied to clipboard
| Challenge: | Recent studies suggest that sparsity is a problem when the trained model is used for inference. |
| Approach: | They propose an alternative to softmax that produces a dense probability distribution but is slower than softmax. |
| Outcome: | The proposed method keeps its virtuous characteristics but is slower than softmax and achieves on par or better performance in machine translation task. |
Surprisingly Easy Hard-Attention for Sequence to Sequence Learning (D18-1)
Copied to clipboard
| Challenge: | Existing attention mechanisms are hard and hard, but they are more accurate when trained. |
| Approach: | They propose to use a beam approximation of the joint distribution between attention and output to train sequence to sequence learning. |
| Outcome: | The proposed method is compared to existing attention mechanisms on five translation tasks and shows consistent gains on the same tasks. |
SEQˆ3: Differentiable Sequence-to-Sequence-to-Sequence Autoencoder for Unsupervised Abstractive Sentence Compression (N19-1)
Copied to clipboard
| Challenge: | Neural sequence-to-sequence models are currently the dominant approach in natural language processing tasks, but require massive parallel corpora. |
| Approach: | They propose a sequence-to-sequence-tosequnce autoencoder with words as latent variables . they apply the model to unsupervised abstractive sentence compression . |
| Outcome: | The proposed model achieves promising results in unsupervised sentence compression on benchmark datasets. |
Sparse-to-Dense: A Free Lunch for Lossless Acceleration of Video Understanding in LLMs (2025.acl-short)
Copied to clipboard
| Challenge: | Recent advances in Video Large Language Models (Video-LLMs) have achieved exceptional performance on tasks like video question answering and captioning. |
| Approach: | They propose a decoding strategy that leverages sparse top-K attention and dense full attention to accelerate Video-LLMs without loss. |
| Outcome: | The proposed approach achieves a 1.94 walltime speedup in video processing. |
Pre-Trained Multilingual Sequence-to-Sequence Models: A Hope for Low-Resource Language Translation? (2022.findings-acl)
Copied to clipboard
En-Shiun Lee, Sarubi Thillainathan, Shravan Nayak, Surangika Ranathunga, David Adelani, Ruisi Su, Arya McCarthy
| Challenge: | Pre-trained multilingual sequence-to-sequence models like mBART and mT5 can be used to translate low-resource languages, but their practical application is unclear. |
| Approach: | They conduct an empirical experiment in 10 languages to determine what can pre-trained multilingual sequence-to-sequence models like mBART do to translate low-resource languages? |
| Outcome: | The proposed models are robust to domain differences, but translations for unseen and typologically distant languages remain below 3.0 BLEU. |
On Sparsifying Encoder Outputs in Sequence-to-Sequence Models (2021.findings-acl)
Copied to clipboard
| Challenge: | Using sequence-to-sequence models, encoder outputs are usually transferred to the decoder for generation, but in this study, encoded outputs can be compressed to shorten the sequence for decoding. |
| Approach: | They propose to use a stochastic gate-based algorithm to mask encoder outputs to shorten the sequence delivered for decoding. |
| Outcome: | The proposed model can be used to shorten encoder outputs to short a sequence . the proposed model yields a speedup of up to 1.65 on document summarization and 1.20 on character-based machine translation tasks. |
A Detailed Evaluation of Neural Sequence-to-Sequence Models for In-domain and Cross-domain Text Simplification (L18-1)
Copied to clipboard
| Challenge: | Xu et al., 2016) show that a simple neural architecture can be efficiently used for in-domain and cross-domain text simplification. |
| Approach: | They evaluate neural sequence-to-sequence models for text simplification on Wikipedia and Newsela datasets. |
| Outcome: | The proposed model can generalize across corpora and overcome challenges when tested on Wikipedia and Newsela datasets. |
Sparse Text Generation (2020.emnlp-main)
Copied to clipboard
| Challenge: | Current text generators require sampling from a modified softmax to avoid degenerate text . entmax sampling creates a mismatch between training and testing conditions . |
| Approach: | They propose to use entmax transformation to train and sample from a sparse language model to avoid degenerate text. |
| Outcome: | The proposed model improves fluency and consistency, fewer repetitions, and n-gram diversity closer to human text. |
Efficient Out-of-Domain Detection for Sequence to Sequence Models (2023.findings-acl)
Copied to clipboard
Artem Vazhentsev, Akim Tsvigun, Roman Vashurin, Sergey Petrakov, Daniil Vasilev, Maxim Panov, Alexander Panchenko, Artem Shelmanov
| Challenge: | Sequence-to-sequence (seq2sequ) models are a ubiquitous tool for text generation but they are not suitable for many other tasks. |
| Approach: | They propose to use UE techniques to identify out-of-domain (OOD) inputs where the model is susceptible to errors. |
| Outcome: | The proposed methods outperform heavyweight ensembles on the task of OOD detection. |