Papers with self-attention
Copied to clipboard
| Challenge: | Modern pre-trained language models are mostly built upon stereotyped development sets . LV-BERT model obtained by our method outperforms BERT on various downstream tasks . |
| Approach: | They propose to exploit layer variety from the layer type set and the layer order to improve pre-trained models. |
| Outcome: | The proposed model outperforms BERT and its variants on various downstream tasks. |
Copied to clipboard
| Challenge: | Self-attention suffers from quadratic computation and memory requirements with respect to sequence length . despite its effectiveness, self-attention models suffer from quadratic computation and a limited set of locations . |
| Approach: | They propose to learn dynamic sparse attention patterns that avoid allocating computation and memory to attend to content unrelated to the query of interest. |
| Outcome: | The proposed model outperforms similar sparse attention models on language modeling and image generation on Wikitext-103 . |
Copied to clipboard
| Challenge: | Existing approaches to extend transformers to source-side trees are linearized into sequences, but they are limited by positional encodings. |
| Approach: | They propose a method for extending transformers to source-side trees by using masks based on tree positions . they define a number of masks that limit self-attention based upon relationships among tree nodes . |
| Outcome: | The proposed method improves on translations from English to germany and English to english and germany by +2.1 BLEU. |
Copied to clipboard
| Challenge: | Recent studies suggest that self-attention implements kernel principal component analysis (KPCA) Across 10 transformer architectures, we conclude that the KPCA interpretation of self- attention lacks empirical support. |
| Approach: | They revisit claims that self-attention implements kernel principal component analysis . they argue that self attention projects queries onto principal component axes of key matrix K . |
| Outcome: | The proposed kernel principal component analysis does not match the proposed kernel . the proposed method is not able to detect the eigenvalues of the gram matrix . |
Copied to clipboard
| Challenge: | Existing work suggests that the computational capabilities of self-attention to model hierarchical structures are limited. |
| Approach: | They investigate the computational power of self-attention to model formal languages . they show strong theoretical limitations of self attention to model periodic finite-state languages unless the number of layers or heads increases with input length. |
| Outcome: | The proposed models can model periodic finite-state languages, nor hierarchical structure unless the number of layers or heads increases with input length. |
Copied to clipboard
| Challenge: | Existing approaches to read comprehension using gated-attention have been effective . collaborative gating and self-belief aggregation are proposed to address these assumptions . |
| Approach: | They propose to use a document-to-query attention system to gate token encodings of a query . they conjecture that query tokens other than the cloze token may be informative . |
| Outcome: | The proposed approaches advance the state-of-the-art results in CNN, Daily Mail, and Who Did What public test sets. |
Copied to clipboard
| Challenge: | Existing methods for named entity recognition (NER) do not exploit word boundary information from CWS or cannot filter the specific information of CWS. |
| Approach: | They propose to exploit task-shared boundary information to make full use of Chinese NER task and Chinese word segmentation (CWS) . |
| Outcome: | The proposed model significantly outperforms other state-of-the-art methods on two widely used datasets. |
Copied to clipboard
| Challenge: | Existing methods for duplicate classification require manual review and assigning bugs to the correct teams. |
| Approach: | They propose a loss function that can detect duplicate bug reports and aggregate them into latent topics without supervision. |
| Outcome: | The proposed model outperforms state-of-the-art methods for duplicate classification on both cases and can learn meaningful latent clusters without supervision. |
Copied to clipboard
| Challenge: | Existing attention mechanisms are data-driven, but most are data driven. |
| Approach: | They propose a knowledge-attention encoder which integrates prior knowledge from external lexical resources into deep neural networks for relation extraction task. |
| Outcome: | The proposed system outperforms existing CNN, RNN, and self-attention based models on a large-scale relation extraction dataset. |
Copied to clipboard
| Challenge: | Recent advances in QA models focus on the targeted area in the passage. |
| Approach: | They propose a model which uses BERT and hierarchical attention to locate a continuous span of the passage that is the answer to the question. |
| Outcome: | The proposed model is based on a BERT embedding and a hierarchical attention model . it can locate a continuous span of the passage that is the answer to the question . |
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is a task of locating and classifying entities mentioned in unstructured text into predefined categories. |
| Approach: | They propose to use a BERT-based multi-question MRC task where multiple questions (one question per entity) are considered at the same time for a single text. |
| Outcome: | The proposed architecture leads to 2.5 times faster training and 2.3 times faster inference on three NER datasets. |
Copied to clipboard
| Challenge: | Pretrained transformer-based language models have produced state-of-the-art performance in most natural language understanding tasks. |
| Approach: | They propose two hybrid architectures that combine self-attention and additive attention mechanisms with sub-layer normalization to achieve double the pretraining accuracy of a vanilla-BERT baseline. |
| Outcome: | The proposed architectures outperform BERT-base on two downstream tasks while accelerating inference. |
Copied to clipboard
| Challenge: | Recent advances in NLP have created problems with the complexity of the self-attention layer. |
| Approach: | They propose to substitute standard self-attention with a local efficient one to avoid the computation of attention weights. |
| Outcome: | The proposed model matches the baseline performance and improves efficiency by skipping the computation of weights that standard attention discards. |
Copied to clipboard
| Challenge: | Named-entity recognition (NER) models are highly dependent on large amounts of labeled data. |
| Approach: | They propose a method that finds translations based on bilingual word embeddings . they also propose 'self-attention' which allows for a degree of flexibility with respect to word order . |
| Outcome: | The proposed method achieves state-of-the-art or competitive performance on common languages with lower resource requirements than previous approaches. |
Copied to clipboard
| Challenge: | Using textual features, our proposed HiGRU models achieve at least 8.7%, 7.5%, 6.0% improvement over the state-of-the-art methods on each dataset. |
| Approach: | They propose a hierarchical gated recurrent unit framework to model word-level inputs and an upper-level GRU to capture contexts of utterance-level embeddings. |
| Outcome: | The proposed framework achieves 8.7%, 7.5%, 6.0% improvement over state-of-the-art methods on three datasets. |
Copied to clipboard
| Challenge: | Existing classification and regression models that only extract finer-grained information from magnetic resonance imaging (MRI) may not be effective for Alzheimer's disease (AD). |
| Approach: | They propose to use a 3D Adapter in a Vision Transformer to extract the patient's EHR information and questions related to the disease as text prompts. |
| Outcome: | The proposed model can discriminate and predict the corresponding MMSE score based on the extracted brain structural information and textual content . |
Copied to clipboard
| Challenge: | Existing methods to transform contextualised representations weaken excessive effects of contextual information. |
| Approach: | They propose a self-supervised learning method that distils word meaning in context from a pre-trained masked language model. |
| Outcome: | The proposed method outperforms the state-of-the-art method for lexical semantics and STS estimation. |
Copied to clipboard
| Challenge: | Language models are generally trained on short, truncated input sequences, which limits their ability to use discourse-level information present in long-range context to improve their predictions. |
| Approach: | They analyze two long-range Transformer language models that accept 8K token inputs . they find that providing long-term context only improves their predictions on a small set of tokens - not sentence-level ones . |
| Outcome: | The proposed model improves on PG-19 with only 2K tokens and does not help at all for sentence-level prediction tasks. |
Copied to clipboard
| Challenge: | Recent work shows that attention mechanisms provide arguably explainable attention distributions that can help to interpret predictions. |
| Approach: | They propose a new self-attention layer where attention heads represent labels. |
| Outcome: | The proposed model obtains state-of-the-art results on the Penn Treebank and Chinese Treebank. |
Copied to clipboard
| Challenge: | Existing “attribute-as-feature” approaches do not create new aggregation pathways for sparsely connected entities. |
| Approach: | They propose an “attribute-as-structure” approach specifically designed for heterogeneous hypergraphs that integrates attributes directly into the hypergraph topology as distinct node types. |
| Outcome: | The proposed approach integrates attributes directly into the hypergraph topology as distinct node types, creating new structural pathways to enrich sparsely connected entities while preserving semantic distinctiveness within complex complex interactions. |
Copied to clipboard
| Challenge: | Existing neural machine translation models use a deep multi-head self-attention network with no explicit phrase information. |
| Approach: | They propose a neural network that combines multi-head self-attention and phrase modeling to train attention heads to attend to phrases in either n-gram or syntactic formalisms. |
| Outcome: | The proposed approach improves on English-to-German and NIST Chinese-to English translation tasks. |
Copied to clipboard
| Challenge: | Tabular data with text fields can be used in financial risk assessment and diagnosis prediction. |
| Approach: | They propose a tabular/text dual-stream Transformer network with numerical embedding schemes and an overall attention module to estimate whether a prediction is uncertain. |
| Outcome: | The proposed model can estimate whether a prediction is uncertain or not based on two well-informed modality streams . |
Copied to clipboard
| Challenge: | Existing word representation models for morphologically rich languages use subword-level information, but their systematic comparative analysis across typologically diverse languages and tasks is still missing. |
| Approach: | They propose a framework for learning subword-informed word representations that allows for easy experimentation with different segmentation and composition components. |
| Outcome: | The proposed framework allows for easy experimentation with different segmentation and composition components, as well as advanced techniques based on position embeddings and self-attention. |
Copied to clipboard
| Challenge: | Existing work on hierarchical structure in neural networks has not captured human intuitions about hierarchic structures. |
| Approach: | They propose to add an extra constraint to attention heads of the bidirectional Transformer encoder to encourage attention heads to follow tree structures. |
| Outcome: | The proposed model improves language modeling and learning more explainable attention scores. |
Copied to clipboard
| Challenge: | Existing methods to improve entity translation in Neural machine translation still suffer from inaccurate translation of entities due to the lack of entity training instances. |
| Approach: | They propose an extract-and-tend approach to enhance entity translation in NMT by extracting entities from a dictionary and attending to them with a prefix. |
| Outcome: | Experiments on En-Zh and En-Ru show that the proposed approach improves translation accuracy and translation quality. |
Copied to clipboard
| Challenge: | Existing work has extended recurrent neural networks to model lattice inputs but these models suffer from slow computation speeds. |
| Approach: | They propose to extend the paradigm of self-attention to handle lattice inputs by adding probabilistic reachability masks that incorporate latticae structure into the model and support lattics if available. |
| Outcome: | The proposed model outperforms baseline models while being much faster to compute than previous models. |
Copied to clipboard
| Challenge: | Neural networks equipped with self-attention have parallelizable computation and the ability to capture both long-range and local dependencies. |
| Approach: | They propose a novel attention mechanism called "Multi-mask Tensorized Self-Attention" it captures pairwise and global dependencies by a compatibility function composed of dot-product and additive attentions . |
| Outcome: | The proposed model outperforms CNN-/RNN-/attention-based models on nine NLP benchmarks with compelling memory- and time-efficiency. |
Copied to clipboard
| Challenge: | Existing work on memory-efficient parallelisms to reduce time and space complexity focuses on reducing time and complexity from system perspective. |
| Approach: | They propose a memory-efficient parallelism to reduce time and space complexity . they split input sequence into multiple chunks and feed each chunk into GPU . |
| Outcome: | The proposed approach is compatible with most existing parallelisms and makes 4D parallelismal possible. |
Copied to clipboard
| Challenge: | Keyword or keyphrase extraction is to identify words or phrases presenting the main topics of a document. |
| Approach: | They propose a hybrid attention model to identify keyphrases from a document in an unsupervised manner. |
| Outcome: | The proposed model is effective and robust on long and short documents. |
Copied to clipboard
| Challenge: | Neural machine translation models can perform word sense disambiguation (WSD) however, it is unclear which component dominates the process of disambiguating words. |
| Approach: | They evaluate hidden states and investigate distributions of self-attention in NMT encoders and decoders to disambiguate word senses. |
| Outcome: | The proposed model outperforms encoder hidden states on large datasets . the model outpersforms decoders on large data sets . |
Copied to clipboard
| Challenge: | Named entity recognition (NER) is a computationally difficult task in Chinese since there is no natural delimiter between words in sentences. |
| Approach: | They propose a data-driven Adaptive Threshold Selective Self-Attention mechanism to select the most relevant characters to enhance Transformer architecture for Chinese named entity recognition. |
| Outcome: | Experiments on four benchmark Chinese NER datasets show the proposed mechanism improves performance. |
Copied to clipboard
| Challenge: | Using parallelizable attention networks, the neural Transformer is slow to train due to auto-regressive architecture and self-attention in the decoder. |
| Approach: | They propose an average attention network to replace the original self-attention model in the decoder of the neural Transformer. |
| Outcome: | The proposed network can decode sentences over four times faster than the original version with almost no loss in training time and translation performance. |
Copied to clipboard
| Challenge: | Neural Machine Translation (NMT) has been replaced by convolutional or self-attentional approaches. |
| Approach: | They propose an architecture definition language that allows for a flexible combination of common building blocks. |
| Outcome: | The proposed architectures can bring recurrent and convolutional models close to the Transformer architecture, but not using self-attention. |
Copied to clipboard
| Challenge: | Modern automatic speech recognition systems rely on encoder-decoder architectures and their encoders are a critical bottleneck for efficient deployment due to high computational intensity. |
| Approach: | They propose a low-rank compression scheme for ASR encoders that leverages the strong low-ranked properties observed in intermediate activations and approximates linear transformations with a chain of low-Rank matrix multiplications. |
| Outcome: | The proposed method reduces inference costs while maintaining transcription accuracy while preserving low-rank properties observed in intermediate activations. |
Copied to clipboard
| Challenge: | Large-scale pre-trained language models require enormous computational resources and long training time. |
| Approach: | They propose an algorithm to reduce inference time and train large NLP models by slimming the self-attention and fully-connected sub-layers inside a transformer. |
| Outcome: | The proposed algorithm achieves comparable performance to standard BERT with 35 45% less training time. |
Copied to clipboard
| Challenge: | Named entity recognition (MNER) aims at identifying entity spans and recognizing their categories in social media posts with the aid of images. |
| Approach: | They propose to use sentences and general domain words to obtain visual cues to transform the fine-grained semantic representation of vision and text into a unified lattice structure and leverage entity boundary detection as an auxiliary task to alleviate visual bias. |
| Outcome: | The proposed method achieves state-of-the-art on two benchmark datasets. |
Copied to clipboard
| Challenge: | Pushdown Layers model recursive state via stack tape that tracks estimated depths of tokens in incremental parsing . pushdown layers are drop-in replacement for standard self-attention . recursion is a key component of many aspects of intelligent behavior, authors say . |
| Approach: | They propose a self-attention layer that models recursive state via a stack tape . Pushdown Layers is a drop-in replacement for standard self- attention . |
| Outcome: | The proposed self-attention layer improves on parse tasks with a recursive-state model . it can model recursion using a stack tape that tracks estimated depths of tokens . |
Copied to clipboard
| Challenge: | Existing approaches to learn discriminative features using contrastive objective are lacking. |
| Approach: | They propose a self-supervised framework that leverages a contrastive loss directly at the level of self-attention. |
| Outcome: | The proposed framework outperforms all comparable unsupervised approaches while occasionally surpassing supervised ones. |
Copied to clipboard
| Challenge: | Despite its importance, discourse element identification is challenging due to the ambiguity of sentences . the number of elaboration sentences could be 10 times more than the number edna sentences. |
| Approach: | They propose to use sentence positional encodings to explicitly represent sentence positions and inter-sentence attentions to capture sentence interactions and enhance sentence representation. |
| Outcome: | The proposed model improves on a Chinese and English dataset. |
Copied to clipboard
| Challenge: | Existing models that use self-attention and position embedding have anomalous behavior that hinder long context window extrapolation. |
| Approach: | They propose a collinear constraint between Q and K to integrate RoPE and self-attention. |
| Outcome: | The proposed model integrates self-attention and position embedding into LLMs without fine-tuning. |
Copied to clipboard
| Challenge: | Long-context understanding is crucial for many NLP applications, but transformers struggle with efficiency due to quadratic complexity of self-attention. |
| Approach: | They propose a dynamic sparse attention mechanism that assigns adaptive masks at the attention-map level, preserving heterogeneous attention patterns. |
| Outcome: | The proposed method achieves high alignment with full-attention models while reducing memory and compute overhead. |
Copied to clipboard
| Challenge: | ProFormer is a projection based transformer architecture that is faster and lighter making it suitable to deploy to memory constraint devices such as mobile phones, watches and IoT. |
| Approach: | They propose a projection based transformer architecture that generates word representations on-the-fly without embedding lookup tables and a local projection attention layer that transforms the input sequence of N LSH word projections into a sequence of K representations. |
| Outcome: | The proposed architecture reduces memory footprint from 92.16 MB to 1.7 KB and requires 16x less computation overhead making it suitable to deploy to memory constraint devices and preserve user privacy. |
Copied to clipboard
| Challenge: | Recent research questions the importance of dot-product self-attention in Transformer models and shows that most attention heads learn simple positional patterns. |
| Approach: | They propose a novel mechanism to replace dot-product self-attention with a recurrent atteNtion mechanism that directly learns attention weights without token-to-token interaction. |
| Outcome: | The proposed model outperforms the Transformer model on translation tasks with fewer parameters and inference time. |
Copied to clipboard
| Challenge: | Currently, word embeddings are playing a pivotal role in many natural language processing tasks. |
| Approach: | They propose a model to learn Chinese word embeddings via three-level composition . they use convolutional neural network to extract intra-character compositionality from character shape . |
| Outcome: | The proposed model performs better on word similarity, sentiment analysis, named entity recognition and part-of-speech tagging tasks. |
Copied to clipboard
| Challenge: | Neural models for morphological inflection have recently attained very high results, but their interpretation remains challenging. |
| Approach: | They propose a linguistically-motivated variant to the encoder-decoder model with attention that incorporates a character-level cross-attention mechanism and a self-attention module over substrings of the input. |
| Outcome: | The proposed model performs well on three typologically-different languages and is highly interpretable. |
Copied to clipboard
| Challenge: | Existing models for discourse relation recognition use self-attention and interactive-attention mechanisms. |
| Approach: | They develop a propagative attention learning model using a cross-coupled two-channel network. |
| Outcome: | The proposed model improves on the baseline models on a Penn Discourse Treebank. |
Copied to clipboard
| Challenge: | Existing methods for recommending citations suffer from severe information loss . citation recommender methods do not consider the section of the paper for which the user is writing and for which they need to find a citation . |
| Approach: | They propose a novel embedding-based neural network to recommend citations during manuscript preparation. |
| Outcome: | The proposed method can recommend citations during manuscript preparation. |
Copied to clipboard
| Challenge: | We compare attention functions in pre-trained language models to human eye fixation patterns during task-specific reading tasks. |
| Approach: | They compare attention functions in large-scale pre-trained language models to classical cognitive models of human attention by using a dataset with eye-tracking recordings of native speakers of English. |
| Outcome: | The proposed model is as predictive of human eye fixation patterns as classical cognitive models of human attention. |
Copied to clipboard
| Challenge: | Detecting disfluency can be difficult because of the flexible nature of reparandum structure and the lack of a nested structure. |
| Approach: | They propose a semi-supervised approach which extracts hidden features from self-attention without any Recurrent Neural Network (RNN) or Convolutional Neural Net (CNN). |
| Outcome: | The proposed approach improves over baselines by using unlabelled data . identifying and removing non-fluent factors would help to improve spontaneous speech quality . |
Copied to clipboard
| Challenge: | Recent work has shown that convolutions have been successful in natural language learning. |
| Approach: | They propose a convolutional approach to construct relative position embeddings in self-attention layers and propose 'compact attention' they propose multiple ways to integrate convolutions into Transformer self- attention. |
| Outcome: | The proposed composite attention improves performance on multiple downstream tasks, replacing absolute position embeddings, and is more expressive than convolutions in NLP. |
Copied to clipboard
| Challenge: | Recent top-performing models in Answer Sentence Selection use self-attention and transfer learning, but not syntactic structure. |
| Approach: | They propose a recursive, tree-structured self-attention model that can represent all levels of syntactic parse trees with only one additional layer. |
| Outcome: | The proposed model can represent all levels of syntactic parse trees with only one additional layer without transfer learning. |
Copied to clipboard
| Challenge: | Existing sparse self-attention fine-tuning models have been used to improve sentiment analysis, question answering, and natural language inference tasks. |
| Approach: | They propose a Sparse Self-Attention Fine-tuning model which integrates sparsity into self-attention mechanism to enhance the fine-tune performance of BERT. |
| Outcome: | The proposed model outperforms the baseline models on sentiment analysis, question answering, and natural language inference tasks and is able to interpret the input better. |
Copied to clipboard
| Challenge: | Aspect-based Sentiment Analysis (ABSA) aims to predict sentiment polarity towards aspects in sentences . a novel model for ABSA is proposed, but how to harness it is still a challenge . |
| Approach: | They propose a syntactic and semantic enhanced Graph Convolutional Network (SSEGCN) model for ABSA task using aspect-aware attention mechanism and self-attention. |
| Outcome: | The proposed model outperforms state-of-the-art methods on benchmark datasets. |
Copied to clipboard
| Challenge: | Existing models for long sequences are not efficient due to the quadratic space and time complexity of the self-attention modules. |
| Approach: | They propose to reduce the quadratic complexity to linear (modulo logarithmic factors) by low-dimensional projection and row selection. |
| Outcome: | The proposed methods outperform transformer-based models with smaller time/space footprint on the Long Range Arena benchmark. |
Copied to clipboard
| Challenge: | Recent work has explored the generalized Dyck-n (Dn) languages . |
| Approach: | They compare the performance of two variants of self-attention networks for Dyck-n (Dn) languages with a starting symbol. |
| Outcome: | The proposed model can generalize to longer sequences and deeper dependencies. |
Copied to clipboard
| Challenge: | In the Transformer model, “self-attention” combines information from attended embeddings into the representation of the focal embeddable in the next layer. |
| Approach: | They propose two methods to quantify flow of information through self-attention using attention weights as relative relevance of input tokens. |
| Outcome: | The proposed methods give complementary views on the flow of information and yield higher correlations with importance scores of input tokens. |
Copied to clipboard
| Challenge: | Existing transformer-based models can only process long documents with limited computational resources due to their quadratic computation time and space. |
| Approach: | They propose to use state-space models for long document classification tasks instead of using sparse or hierarchical structures to solve this problem. |
| Outcome: | The proposed model performs comparable to self-attention models while being 36% more efficient. |
Copied to clipboard
| Challenge: | Phrase-level self-attention networks (PSAN) can capture context dependencies at the phrase level instead of the sentence level. |
| Approach: | They propose to perform self-attention across words inside a phrase to capture context dependencies at the phrase level and use the gated memory updating mechanism to refine each word’s representation hierarchically with longer-term context dependency captured in a larger phrase. |
| Outcome: | The proposed model can achieve state-of-the-art performance across a plethora of NLP tasks including binary and multi-class classification, natural language inference and sentence similarity. |
Copied to clipboard
| Challenge: | Attention-based language models rely on the softmax function to convert attention logits into probability distributions, but this process can result in attention entropy collapse. |
| Approach: | They propose to use the softmax function to re-weight attention logits to create probability distributions, but this reweighting can lead to attention entropy collapse . they find that entropic-stable attention methods can prevent entrapment and enable more stable training by controlling or insensitive to variance of attention logit variance. |
| Outcome: | The proposed methods prevent attention entropy collapse and enable more stable training. |
Copied to clipboard
| Challenge: | Existing methods for large language models with extended context lengths face significant computational challenges during the prefill phase. |
| Approach: | They propose a difference-aware, dynamic sparse attention mechanism that efficiently identifies critical attention regions at a finer stripe granularity while adapting to global contextual information. |
| Outcome: | The proposed model achieves a speedup of 1.44 while maintaining higher recall rates. |
Copied to clipboard
| Challenge: | Existing models of BERT-based learning systems are lacking specific mechanisms that contribute to its success. |
| Approach: | They propose to use GLUE tasks to analyze the interpretation of self-attention, which is one of the underlying components of BERT. |
| Outcome: | The proposed model outperforms the regular model on GLUE tasks by disabling attention in certain heads. |
Copied to clipboard
| Challenge: | Extensive experiments on ten WMT machine translation tasks show that the proposed model yields an average of 1.35x faster (with almost no decrease in BLEU) |
| Approach: | They propose a weighted residual network which reconstructs attention by reusing the features across layers. |
| Outcome: | The proposed model is 1.35x faster than the state-of-the-art inference model on translation tasks compared to AAN and SAN models with fewer parameter numbers . |
Copied to clipboard
| Challenge: | Existing multimodal DA classification approaches are limited by ineffective audio modeling and late-stage fusion. |
| Approach: | They propose a framework for online multimodal dialog act (DA) classification based on raw audio and ASR-generated transcriptions of current and past utterances. |
| Outcome: | The proposed model achieves a significant increase in the F1 score relative to current state-of-the-art models on two prominent DA classification datasets, MRDA and EMOTyDA. |
Copied to clipboard
| Challenge: | Character-level BERT pre-trained in Chinese suffers from lacking lexicon information, which shows effectiveness for Chinese NER. |
| Approach: | They propose a semi-supervised method to integrate lexicon into pre-trained LMs in Chinese . they extract an entity lexiconal from raw text and integrate it into BERT . |
| Outcome: | The proposed method is highly effective and achieves the best results on a news dataset and two datasets annotated by the authors. |
Copied to clipboard
| Challenge: | Existing models for named entity recognition (NER) use sentence-level labels, which are expensive to obtain, to improve NER. |
| Approach: | They propose a sentence-level named entity recognition model that uses sentence-based labels that are easy to obtain. |
| Outcome: | The proposed model produces 3.78%, 4.20%, 2.08% improvements in F1 over the baseline on e-commerce product titles in Vietnamese, Thai, and Indonesian, respectively. |
Copied to clipboard
| Challenge: | Existing models to incorporate syntactic structures into neural language models have relied heavily on elaborate components for a specific language model, which makes them unwieldy in practice to fit into other models. |
| Approach: | They propose a dependency-based mixture language model that incorporates syntactic structures into neural language models by mixing previous dependency modeling probabilities with self-attention. |
| Outcome: | The proposed method can be easily and effectively applied to different neural language models while improving neural text generation on various tasks. |
Copied to clipboard
| Challenge: | Existing approaches to NLP are sparsifying attention patterns or approximating the attention computation with kernel methods. |
| Approach: | They propose a method for dynamic contextual compression for decoder-only LMs. |
| Outcome: | The proposed method reduces the cost of self-attention to a fraction of typical time and space. |
Copied to clipboard
| Challenge: | Retrieval-augmented generation (RAG) aims to mitigate the hallucination of Large Language Models (LLMs) however, external knowledge may contain noise and conflict with parametric knowledge of LLMs, leading to degraded performance. |
| Approach: | They propose a Dual-Stream Knowledge-Augmented Framework for Shared-Private Semantic Synergy that refines the traditional self-attention into a mixed-attention that distinguishes shared and private semantics for a controlled knowledge integration. |
| Outcome: | Extensive experiments show that the proposed framework achieves a superior performance over baselines. |
Copied to clipboard
| Challenge: | Existing work on pre-trained Transformers has focused on learning the meaning of positions . Embedding the position information in the self-attention mechanism is also an indispensable factor in NLP . |
| Approach: | They propose to use feature-level analysis to examine pre-trained Transformers' position embeddings . they also use empirical experiments to determine the appropriate positional encoding function . |
| Outcome: | The results of the empirical study can guide future work to choose the appropriate positional encoding function for specific tasks. |
Copied to clipboard
| Challenge: | Existing RNN-based LLMs struggle with long-context scenarios due to their quadratic computational complexity and linear memory requirements. |
| Approach: | They propose an efficient scaling method to scale RNN models to match the 2k context length of Transformers with small parameters overhead. |
| Outcome: | The proposed method improves long-context understanding and improves performance on FDA recall-intensive tasks. |
Copied to clipboard
| Challenge: | cloze-style reading comprehension is a task that requires much semantic understanding and reasoning using various clues from texts. |
| Approach: | They propose a multi-choice relational reasoning model that emulates human reading comprehension by combining fusion representations of document, query and candidates. |
| Outcome: | The proposed model outperforms baseline models significantly on four datasets. |
Copied to clipboard
| Challenge: | Existing transformer models are computationally demanding and prohibitively costly for long sequences due to the quadratic complexity of its selfattention module. |
| Approach: | They propose a transformer-based model that inherits weights from large pretrained models by removing redundancies in hidden sequences using the ready-made Fast Fourier Transform operator. |
| Outcome: | The proposed model outperforms the standard BART model on the long-range modeling benchmark LRA with significant improvements in speed and space. |
Copied to clipboard
| Challenge: | Vision-Language Models (VLMs) have gained prominence due to their success in solving complex cross-modal tasks. |
| Approach: | They propose a Gaussian-Noise-free pipeline for mechanistic interpretability in VLMs that introduces Semantic Image Pairs corruption, the first visual counterpart to Symmetric Token Replacement for text. |
| Outcome: | The proposed pipeline identifies a set of “universal attention heads” in BLIP and LLaVA that consistently contribute across different tasks and modalities. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have yielded impressive successes on many language tasks, but efficient processing of long contexts remains a significant challenge. |
| Approach: | They propose a method for processing long contexts in an episodic memory module while holistically attending to semantically-relevant context chunks. |
| Outcome: | The proposed method outperforms baseline decoders on multiple long-context recall and question-answering benchmarks on 16k to 256k tokens. |
Copied to clipboard
| Challenge: | Existing studies focus on multi-hop question answering across multiple documents or paragraphs. |
| Approach: | They propose a graph neural network to deal with graph structure in textual multi-hop reasoning . they propose 'self-attention' and propose removing entire graph structure may not hurt the final results . |
| Outcome: | The proposed model shows that graph-attention or the entire graph structure can be replaced by self-attention . hotpotQA is a widely used benchmark for multi-hop question answering . |
Copied to clipboard
| Challenge: | a new approach to the self-attention mechanism is proposed for integrating data from multiple batches. |
| Approach: | They propose an autoregressive with exogenous inputs approach for the Transformer model . the proposed method transforms the Encoder block into a negative feedback predictive control system . |
| Outcome: | The proposed method is validated through comparative evaluations. |
Copied to clipboard
| Challenge: | Neural architectures based on self-attention have attracted interest from the research community . a recent study examined the performance of Transformers on a task of Neural Question Generation . |
| Approach: | They propose to adapt Transformers to a task of Neural Question Generation without constraining the model to focus on a specific answer passage. |
| Outcome: | The proposed architectures have obtained significant improvements over the state-of-the-art in several tasks. |
Copied to clipboard
| Challenge: | Experiments show that ChunkAttention can speed up the self-attention kernel by 3.2-4.8 compared to the start-of-the-art implementation. |
| Approach: | They propose a prefix-aware self-attention module that can detect matching prompt prefixes across multiple requests and share their key/value tensors in memory at runtime. |
| Outcome: | The proposed module can speed up the self-attention kernel by 3.2-4.8 compared to the start-of-the-art implementation, with the length of the system prompt ranging from 1024 to 4096. |
Copied to clipboard
| Challenge: | Transformer-based large language models are trained to make predictions about the next word by aggregating representations of previous tokens through their self-attention mechanism. |
| Approach: | They propose an entropy-based predictor that quantifies the diffuseness of self-attention and a distance-based one that captures the incremental change in attention patterns across timesteps. |
| Outcome: | The proposed models perform better over a rigorous baseline including GPT-2 surprisal than previous models that used entropy-based predictors and distance-based ones. |
Copied to clipboard
| Challenge: | Neural networks have become indispensable across a variety of natural language processing tasks. |
| Approach: | They propose a theoretical approach based on Neural Tangent Kernels to investigate neural networks' internal mechanisms. |
| Outcome: | The proposed approach can be applied to analyze language modeling tasks . it shows that the choice of activation function can affect feature extraction . |
Copied to clipboard
| Challenge: | Existing frameworks for symptom status recognition in doctor-patient dialogues are inadequate. |
| Approach: | They propose a framework for symptom status recognition that formalizes a natural language inference task . they generate knowledge about the symptom and a hypothesis about its status for each symptom . |
| Outcome: | The proposed framework outperforms baselines and has advantages in cross-disease and cross-symptom scenarios. |
Copied to clipboard
| Challenge: | Recent work has questioned the importance of multi-headed attention in achieving high translation quality. |
| Approach: | They develop a “hard-coded” attention variant without any learned parameters. |
| Outcome: | The proposed model reduces BLEU scores by adding a single learned cross attention head to an otherwise hard-coded Transformer. |
Copied to clipboard
| Challenge: | Named entity recognition (NER) is a fundamental information extraction task that focuses on extracting entities from a given text and classifying them using pre-defined categories. |
| Approach: | They propose to use “entity triggers” to facilitate label-efficient learning of NER models. |
| Outcome: | The proposed model is significantly more cost-effective than the traditional neural NER frameworks. |
Copied to clipboard
| Challenge: | Existing methods for text regression lack local grounding and rely on shared representations. |
| Approach: | They propose a distributional regression model with quantile tokens that insert dedicated quantiles into the input sequence. |
| Outcome: | The proposed method outperforms baseline models on the inside Airbnb and StackSample datasets. |
Copied to clipboard
| Challenge: | Existing Transformers that scale to long sequences are not compatible with relative position encoding. |
| Approach: | They propose a Performer-based model with relative position encoding that scales linearly on long sequences. |
| Outcome: | The proposed model outperforms performer on long sequences with no computational overhead and outperformed vanilla Transformer on most of the tasks. |
Copied to clipboard
| Challenge: | In multivariate long-term time series forecasting, it is widely believed that the effectiveness of self-attention arises from its attention matrix. |
| Approach: | They propose a multi-branch MLP that isolates the ‘multi-brain mapping with element-wise operation’ structure from the Transformer and shows that it achieves competitive performance. |
| Outcome: | The proposed model outperforms three classic and three latest Transformer models and shows that it achieves competitive performance. |
Copied to clipboard
| Challenge: | Structured dropout approaches have been investigated to regularize the multi-head attention mechanism in Transformers. |
| Approach: | They propose a new regularization scheme based on token-level rather than structure-level to reduce overfitting by manipulating the connections between tokens in the multi-head attention via masking. |
| Outcome: | The proposed regularization scheme outperforms attention dropout and DropHead on 18 datasets and can establish a new record on the data-to-text benchmark Rotowire (18.93 BLEU). |
Copied to clipboard
| Challenge: | Large Language Models trained on code corpora have limitations such as suggesting codes with syntactic errors, variable misuse etc. |
| Approach: | They conduct a fine-grained analysis of attention maps and hidden representations of large-scale Large Language Models (cLLMs) trained on a large corpus of code and natural language -programming language pairs. |
| Outcome: | The proposed models encode relations among syntactic tokens and identifiers, but fail to encode relations between syntaktic token and identifier. |
Copied to clipboard
| Challenge: | Existing approaches address these bottlenecks separately: Multi-head Latent Attention (MLA) reduces the KV cache by projecting tokens into a low-dimensional latent space, while sparse attention reduces computation. |
| Approach: | They propose a Latent-Condensed Attention mechanism that performs structured context condensation directly within MLA's latent space. |
| Outcome: | The proposed approach reduces KV cache size and attention cost without adding parameters. |
Copied to clipboard
| Challenge: | Current decoder-only architectures achieve higher performance but lower efficiency . cross-attention-based architectures skip visual token computations . |
| Approach: | They propose a training-free framework for analyzing trained MLLMs to investigate redundancy . they propose 'probe-activated Dynamic FFN and Hollow Attention' algorithms for visual token reductions and a layer ranking algorithm for inference acceleration. |
| Outcome: | The proposed framework achieves comparable performance to or better than state-of-the-art methods while remaining compatible with them. |
Copied to clipboard
| Challenge: | Existing studies focus on manipulating word inputs, but they lack generalization to versatile real-world attacks. |
| Approach: | They propose a powerful perturbation technique which perturbs the attention scores within the SA matrices via meticulously crafted attention masks. |
| Outcome: | The proposed perturbation technique achieves high attack success rate (98%) and low cost. |
Copied to clipboard
| Challenge: | Existing models for depression severity estimations lack uncertainty estimates and temporal interpretability. |
| Approach: | They propose a Probabilistic framework for Depression Detection from clinical interview utterance sequences that predicts PHQ-8 scores while modeling calibrated uncertainty. |
| Outcome: | The proposed framework achieves competitive performance among text-only systems and produces well-calibrated intervals. |
Copied to clipboard
| Challenge: | Existing methods for post-training quantization struggle to support weight–activation joint quantization and extreme low-bit weight quantization. |
| Approach: | They propose a framework that addresses weight–activation joint quantization and extreme weight quantization. |
| Outcome: | The proposed framework achieves superior performance under both W4A4 and highly aggressive W2 settings while incurring negligible additional computational overhead. |