Papers with Transformer
Copied to clipboard
| Challenge: | Existing methods for machine translation require intensive keyboard interaction, which is inconvenient on mobile devices. |
| Approach: | They propose a touch-based editing method that is more flexible than keyboard-mouse-based translation postediting. |
| Outcome: | The proposed method significantly outperforms existing interactive translation methods on translation datasets and on post-editing datasets. |
Copied to clipboard
| Challenge: | Experimental results show that unified model outperforms other models that treat encoding and matching separately. |
| Approach: | They evaluate a unified model with Transformer layers for machine reading comprehension . they find that the model learns different modeling strategies compared with previous models . |
| Outcome: | The unified model outperforms models with Transformer layers on the machine reading comprehension task. |
Copied to clipboard
| Challenge: | Existing attempts to integrate semantic structures into NMT Transformers have failed . |
| Approach: | They propose two parameter-free methods for injecting semantic information into Transformers, using a Scene-Aware Self-Attention (SASA) head and a Scenario-Award Cross-Action (SACrA) head. |
| Outcome: | The proposed methods improve on the vanilla Transformer and syntax-aware models for four language pairs and show an additional gain when using both semantic and syntactic structures in some language pairs. |
Copied to clipboard
| Challenge: | Existing approaches to simultaneous translation are limited by monotonic constraint . a novel architecture for simultaneous translation is proposed . |
| Approach: | They propose a cross attention-augmented transducer for simultaneous translation that optimizes both policies and translation models by expanding target sequences with blank symbols. |
| Outcome: | The proposed architecture achieves better latency-quality trade-offs than state-of-the-art approaches. |
Copied to clipboard
| Challenge: | a new method to prune attention heads is proposed for adversarial detection . attention heads in models such as BERT are over-provisioned and can be pruned . |
| Approach: | They propose a method to construct input-specific attention subnetworks from which three features are extracted to discriminate between authentic and adversarial inputs. |
| Outcome: | The proposed method significantly improves state-of-the-art adversarial detection accuracy on 10 NLU datasets with 11 different adversarials. |
Copied to clipboard
| Challenge: | Existing studies on the effectiveness of different architectures for modeling low resource languages are limited. |
| Approach: | They propose a trainable memory efficient CNN architecture for Bengali and Hindi . they propose two learnable convolutional sub-models that are end to end trainable . |
| Outcome: | The proposed model outperforms pretrained BERT models on Bengali and Hindi with 16X less parameters and achieves much better performance than SOTA LSTMs on multiple real-world datasets. |
Copied to clipboard
| Challenge: | Evaluating translation models is a trade-off between effort and detail. |
| Approach: | They propose to use a neural text classifier to automatically expose systematic differences between human and machine translations to human experts. |
| Outcome: | The proposed method exposes systematic differences between human and machine translations to human experts. |
Copied to clipboard
| Challenge: | Various tools have been developed to visualize attention in NLP models, ranging from attention-matrix heatmaps to bipartite graph representations. |
| Approach: | They propose an open-source tool that visualizes attention at multiple scales and provides a unique perspective on the attention mechanism. |
| Outcome: | The proposed model outperforms OpenAI GPT-2 and BERT on several language modeling benchmarks. |
Copied to clipboard
| Challenge: | Asynchronous stochastic gradient descent (SGD) converges poorly for Transformer models . synchronous SGD is faster at raw training speed since it avoids waiting for synchronization . |
| Approach: | They propose a method to restore convergence by summing several asynchronous updates instead of applying them immediately. |
| Outcome: | The proposed method achieves the same BLEU score 1.36 times faster than asynchronous SGD. |
Copied to clipboard
| Challenge: | Existing models have limitations to generalize to diverse semantic phenomena, and it is unclear whether they can capture compositional meanings. |
| Approach: | They propose a systematic generalization testbed based on Natural language semantics to map natural language sentences to multiple meaning representations. |
| Outcome: | The proposed model can generalize to unseen combinations of quantifiers, negations, and modifiers, but not to the others. |
Copied to clipboard
| Challenge: | In the digital age, large-scale online travel platforms face the challenge of extracting valuable insights from massive volumes of textual data. |
| Approach: | They propose a model that uses a bi-encoder transformer architecture to extract structured information from textual data. |
| Outcome: | The proposed model outperforms state-of-the-art models with 92.9% micro mAP and 75.8% macro mA score compared to baseline models . the proposed model can be used to find hotel facilities and hotel rooms based on positive reviews . |
Copied to clipboard
| Challenge: | Existing methods for length extrapolation are tailored for natural language modeling, a task known to have strong recency bias. |
| Approach: | They propose two attention alignment strategies to improve T5's long-context utilization capability without fine-tuning. |
| Outcome: | The proposed methods improve the long-context utilization capability of T5 on language modeling, retrieval, multi-document question answering, and code completion tasks without any fine-tuning. |
Copied to clipboard
| Challenge: | Emotion recognition in conversation (ERC) is a task arousing increasing interest in many fields. |
| Approach: | They propose a novel GNN-based ERC model that captures speaker and position information. |
| Outcome: | The proposed model captures speaker and position-aware conversation structure information. |
Copied to clipboard
| Challenge: | Existing studies ignore hierarchical structures of sememes in sememe-based semantic description systems. |
| Approach: | They propose a structured sememe prediction problem to predict a sememes tree with hierarchical structures rather than a set of sememas. |
| Outcome: | The proposed model outperforms baseline models and shows its effectiveness . it predicts a sememe tree with hierarchical structures rather than a set of sememes . |
Copied to clipboard
| Challenge: | Existing models that use VTs as their backbone model are based on UTs that share parameters across layers and have better compositional generalization. |
| Approach: | They propose to use Sparse Mixture of Experts to reduce UT's computation complexity while retaining its parameter efficiency and generalization ability. |
| Outcome: | The proposed model achieves strong generalization results on formal language tasks and impressive parameter and computation efficiency on standard natural language benchmarks. |
Copied to clipboard
| Challenge: | Existing inference frameworks for natural language processing are not the best choice for online service of sequence processing problems. |
| Approach: | They propose a highly efficient inference library for Transformer models that includes GPU optimization techniques to streamline computation and reduce memory footprint. |
| Outcome: | The proposed library achieves 14x speedup compared with TensorFlow and 1.4x speed up compared to a concurrent CUDA implementation. |
Copied to clipboard
| Challenge: | a myriad of complex tasks require both prior knowledge and reasoning intelligence. |
| Approach: | They propose a plug-and-play quasi-attention mechanism to integrate multimodal graph information to vanilla self-attention as effective prior. |
| Outcome: | The proposed model is able to perform reasoning across multiple modalities. |
Copied to clipboard
| Challenge: | Currently, a mainstream approach to generate pseudo data is back-translation (BT). |
| Approach: | They propose to use back-translation to generate pseudo data that contains grammatical and ungrammatically produced sentences. |
| Outcome: | The proposed methods improve or interpolate the performance of each error type compared with a single BT model with different seeds. |
Copied to clipboard
| Challenge: | a novel data-augmentation technique for neural machine translation is based on a letter substitution cipher . a bijective ciphered text is in effect invisible to modern NLP techniques because of its invariant distributional features . |
| Approach: | They propose a data-augmentation technique for neural machine translation based on ROT-k ciphertexts. |
| Outcome: | The proposed method outperforms existing methods on several datasets by a significant margin. |
Copied to clipboard
| Challenge: | Recent advances in neural machine translation have been made in the field of multi-head self-attention and there is no explicit mechanism to ensure that different attention heads capture different features. |
| Approach: | They propose a novel multi-head self-attention model which models not only global and local attention but also forward and backward attention in different attention heads. |
| Outcome: | The proposed model improves on WAT17 English-Japanese and IWSLT14 German-English translation tasks without increasing the number of parameters. |
Copied to clipboard
| Challenge: | Existing methods for learning word alignment include statistical word aligners (e.g. GIZA++) Existing word alignment models employ a target-to-source attention mechanism which can provide rough word alignments but with a low accuracy. |
| Approach: | They propose a bidirectional Transformer based alignment model for unsupervised learning of the word alignment task. |
| Outcome: | The proposed model outperforms both previous neural word alignment approaches and the popular statistical word aligner GIZA++ on three word alignment tasks. |
Copied to clipboard
| Challenge: | Neural sequence-to-sequence models are sensitive to architecture and hyperparameter settings. |
| Approach: | They incorporate architecture search into a single training run through auto-sizing . they show that auto-size can improve BLEU scores by up to 3.9 points . |
| Outcome: | The proposed algorithm improves BLEU scores on low-resource language pairs while removing one-third of the parameters from the model. |
Copied to clipboard
| Challenge: | Recent studies have examined how linguistic knowledge of language models (LMs) varies across English phenomena. |
| Approach: | They propose a benchmark to evaluate linguistic knowledge of language models on major grammatical phenomena in English. |
| Outcome: | The proposed benchmark evaluates the linguistic knowledge of language models on major grammatical phenomena in English. |
Copied to clipboard
| Challenge: | generative diversity is a critical yet underexplored issue in natural language generation . previous approaches to enhance diversity of Transformer models have been limited by their latent variables . |
| Approach: | They propose a framework that bridges Transformer with VAE to enhance generative diversity. |
| Outcome: | The proposed framework improves generative diversity while maintaining generative quality. |
Copied to clipboard
| Challenge: | Existing methods to handle out-of-vocabulary identifiers are not suitable for source code processing. |
| Approach: | They propose a method to handle out-of-vocabulary identifiers by identifies anonymization . they show that the method significantly improves the performance of the Transformer . |
| Outcome: | The proposed method significantly improves the performance of the Transformer in two code processing tasks. |
Copied to clipboard
| Challenge: | Using attention heads, we can explore how Transformer language models process semantic knowledge, especially regarding the plausibility of noun-verb relations. |
| Approach: | They propose to investigate how Transformer language models process semantic knowledge, especially regarding the plausibility of noun-verb relations. |
| Outcome: | The proposed model exhibits a higher degree of similarity with humans in plausibility processing compared to other Transformer language models. |
Copied to clipboard
| Challenge: | Transformer model has been a de-facto standard in natural language processing, but it is limited to images, text, and/or sequence data. |
| Approach: | They propose to use a multimodal large language model architecture to handle biomedical graphs such as protein structure and chemical molecules to improve its performance. |
| Outcome: | The proposed architecture can handle multiple data types for biomedical graphs such as protein structure and chemical molecules. |
Copied to clipboard
| Challenge: | Existing multihop attentions for machine comprehension are recurrent and hierarchical . a proposed multi-hop attention for the Transformer refines the attention for an output symbol many times . |
| Approach: | They propose a multi-hop attention for the Transformer which integrates attentions from each head. |
| Outcome: | The proposed model outperforms the baseline Transformer in terms of translation accuracy and speed. |
Copied to clipboard
| Challenge: | Existing approaches to extract summary from document with a disproportionate ratio of selected and unselected sentences are far from human performance. |
| Approach: | They propose a model that rebalances sentence-level extractive summarization by amplifying the semantic difference between each sentence and all other sentences and applying the residual unit as the second item of the differential amplifier to deepen the architecture. |
| Outcome: | The proposed model performs competitively against state-of-the-art methods on two benchmark datasets. |
Copied to clipboard
| Challenge: | We extend the maximum context size of a neural network called Transformer to 8k characters. |
| Approach: | They propose a self-attention mechanism that can learn its optimal attention span . this allows for models with longer context and the capability to catch longer dependencies. |
| Outcome: | The proposed model achieves state-of-the-art performance on text8 and enwiki8 using 8k characters with no loss of performance, and maintains control over memory footprint and computational time. |
Copied to clipboard
| Challenge: | Existing document-level neural machine translation systems concatenate several consecutive sentences to form a pseudo-document, and then learn inter-sentential dependencies. |
| Approach: | They propose a document flattening technique that integrates Flat-Batch Attention (FBA) and Neural Context Gate (NCG) into Transformer model to utilize information beyond the pseudo-document boundaries. |
| Outcome: | The proposed method outperforms baselines on BLEU, COMET and accuracy on the contrastive test set. |
Copied to clipboard
| Challenge: | We propose a programmatic solution to generate product advertising headlines using retail content. |
| Approach: | They propose a programmatic solution to generate product advertising headlines using retail content . they use Reinforcement Learning (RL) Policy gradient methods on Transformer . |
| Outcome: | The proposed method outperforms existing methods in overlap metrics and quality audits. |
Copied to clipboard
| Challenge: | Existing methods for Hierarchical Text Classification (HTC) are expensive and require explicit injection of the hierarchy, verbalizers, and/or prompt engineering. |
| Approach: | They propose a hierarchical text classification system that uses a single classifier to predict one or more topics using differentiable prompts and labels that are learnt through backpropagation. |
| Outcome: | The proposed model outperforms existing models on several benchmarks that span a range of topics consistently. |
Copied to clipboard
| Challenge: | Abstractive document summarization is a comprehensive task in natural language processing. |
| Approach: | They propose a topic assistant that rearranges and learns document semantics . they propose TA that is compatible with Transformer-based models and user-friendly . |
| Outcome: | The proposed model is compatible with Transformer-based models and user-friendly. |
Copied to clipboard
| Challenge: | Existing approaches to describe the syntax structure of code are lacking in retaining the semantic structure of source code. |
| Approach: | They propose to use a triplet position to model hierarchical syntax structure of code by introducing a graph neural network and Transformer to preserve the structural and sequential information of code. |
| Outcome: | The proposed model preserves the structural and sequential information of code and a pointer-generator network that pays attention to both the structure and sequential tokens of code for a better summary generation. |
Copied to clipboard
| Challenge: | Current language models generate high-quality text, but are they copying it or have they learned generalizable linguistic abstractions? |
| Approach: | They propose a suite of analyses for assessing the novelty of generated text . they focus on sequential structure (n-grams) and syntactic structure (syntactical structure). |
| Outcome: | The proposed model-generated text is as novel as the baseline human-generated model- generated text, but it is copied substantially, the authors show . |
Copied to clipboard
| Challenge: | Existing studies show that deep Transformers have difficulty in training even with residual connection and layer normalization. |
| Approach: | They propose a method that leverages the Lipschitz constraint on the initialization of Transformer parameters to ease the optimization difficulties caused by its multi-layer encoder/decoder structure. |
| Outcome: | The proposed model outperforms previous RNN/CNN models but fails to converge with the original computation order. |
Copied to clipboard
| Challenge: | Deep learning has demonstrated performance advantages in a wide range of natural language processing tasks. |
| Approach: | They propose to deepen the decoder layer in a Transformer model to reduce the difficulty of deep learning. |
| Outcome: | The proposed method can deepen the model on both the encoder and decoder at the same time, resulting in a deeper model and improved performance. |
Copied to clipboard
| Challenge: | Existing vision-language pre-training methods use a two-step training procedure to learn visual features from image-text pairs. |
| Approach: | They propose a vision-language pre-trained model for V+L understanding and generation using a unified Transformer framework. |
| Outcome: | The proposed model can learn visual representation and semantic alignments between image and text on visual-text pairs and on visual processing tasks. |
Copied to clipboard
| Challenge: | Prior work suggests that Transformer captures poor word alignments through its attention mechanism. |
| Approach: | They propose two new word alignment induction methods that use attention weights to capture accurate word alignments. |
| Outcome: | The proposed methods outperform baselines on three publicly available datasets and are significantly better than GIZA++. |
Copied to clipboard
| Challenge: | Existing multilingual models for voice assistants are limited by their prohibitive inference time and limited performance. |
| Approach: | They propose to distill and deploy multilingual Transformer models for voice assistants using a teacher-student framework that uses teacher-trained models to supervise student model training. |
| Outcome: | The proposed model outperforms a teacher model trained on unlabelled data and achieves equivalent performance. |
Copied to clipboard
| Challenge: | Existing systems for grammatical error correction in English have been limited . however, there is limited progress on error correction of other languages . |
| Approach: | They propose a dataset on grammatical error correction for Czech and an annotated learner corpus for Russian and Czech. |
| Outcome: | The proposed model can reach new state-of-the-art on Czech, German and Russian datasets. |
Copied to clipboard
| Challenge: | Existing work has increased the modeling capacity of multilingual NMT by deepening or widening the Transformer. |
| Approach: | They propose to increase the model capacity by deepening the Transformer . they propose to use a multi-input-multi-output architecture to combine multiple inputs . |
| Outcome: | The proposed model surpasses previous work and is 1.31 times faster than existing models. |
Copied to clipboard
| Challenge: | Recent studies have shown that attention heads learn simple positional patterns . |
| Approach: | They propose to replace all but one attention head of each encoder layer with simple fixed – non-learnable – attentive patterns that are solely based on position and do not require external knowledge. |
| Outcome: | The proposed model improves translation quality and improves BLEU scores by up to 3 points in low-resource scenarios. |
Copied to clipboard
| Challenge: | Documentlevel NLI is an important problem for many tasks including verification of factual correctness of documents. |
| Approach: | They propose a document-level natural language inference model that builds a hierarchical document graph enriched through inter-sentence relations and performs paragraph pruning using the novel SubGraph Pooling layer. |
| Outcome: | The proposed model performs on a legal judicial reasoning task with a dataset enriched with document graphs and a proposed evidence selection algorithm. |
Copied to clipboard
| Challenge: | Recent studies have shown that multilingual NMT models can handle more than one translation direction with a single system. |
| Approach: | They propose a multilingual neural machine translation model that can handle more than one translation direction with a single system. |
| Outcome: | The proposed model performs well in low-resource settings against bilingual systems. |
Copied to clipboard
| Challenge: | a dominant approach to solving NLP tasks is pre-training a large neural language model and fine-tuning the model for specific tasks. |
| Approach: | They propose a challenge to train and fine-tune large Transformer models for historical texts . they pre-trained a RoBERTa model from scratch from the historical texts and evaluate them on benchmarks . |
| Outcome: | The proposed ML task is based on OCR-ed clippings from the Chronicling America portal. |
Copied to clipboard
| Challenge: | Existing Transformers models are computationally expensive for long context inputs. |
| Approach: | They propose a transformer that can interchange information between memory states and context . they evaluate the efficiency of their model on three dialogue datasets and two language datasets . |
| Outcome: | The proposed model is compatible with existing transformer models and can preserve dialogue history information. |
Copied to clipboard
| Challenge: | Existing methods for encoding text in tables require additional training and require additional pretraining. |
| Approach: | They propose a novel encoding strategy that preserves the critical property of permutation invariance across rows or columns. |
| Outcome: | The proposed approach outperforms state-of-the-art methods on three table interpretation tasks: column type annotation, relation extraction, and entity linking. |
Copied to clipboard
| Challenge: | a recent study shows that current NLP models operate non-incrementally, causing unacceptable delays for the user. |
| Approach: | They propose a streaming BERT-based sequence tagging model that detects disfluencies in real-time . they train the model to decide whether to immediately output a prediction or wait for further context . |
| Outcome: | The proposed model produces accurate predictions sooner than baselines, with lower flicker . disfluencies hurt readability of ASR transcripts, erode model performance on downstream tasks . |
Copied to clipboard
| Challenge: | Spectral-normalized identity priors (SNIP) is a structured pruning approach for a Transformer model. |
| Approach: | They propose a structured pruning approach which penalizes an entire residual module toward an identity mapping. |
| Outcome: | The proposed method improves on 5 GLUE benchmark tasks while maintaining comparable performance. |
Copied to clipboard
| Challenge: | Existing pre-trained language models produce large sentence embeddings, resulting in performance gap between large and small models. |
| Approach: | They propose a method that augments a small Transformer encoder model with learnable projection layers to produce compact sentences while mimicking a large pre-trained language model to retain the sentence representation quality. |
| Outcome: | The proposed method achieves 2.7-4.5 points performance gain on STS and SR tasks while maintaining the quality of the pre-trained language models. |
Copied to clipboard
| Challenge: | Existing long-document Transformers do not learn representations of document structure during pretraining. |
| Approach: | They propose to use long-document Transformers to acquire an internal representation of document structure during pre-training and evaluate the effects of structure infusion on QASPER and Evidence Inference. |
| Outcome: | The proposed models acquire implicit understanding of document structure during pre-training, which can be enhanced by structure infusion, leading to improved end-task performance. |
Copied to clipboard
| Challenge: | Existing methods to fine-tune a model for multiple tasks require a large amount of memory and computing power. |
| Approach: | They propose to factorize the weighs of a pre-trained Transformer model to improve training efficiency across multiple tasks by using BERT-Large as an instantiation of the Transformer and the GLUE as the evaluation benchmark. |
| Outcome: | The proposed method matches or improves the original fine-tuned model’s performance for each task while effectively decreasing parameter requirements by two orders of magnitude. |
Copied to clipboard
| Challenge: | Existing methods for extracting temporal information from text are not suitable for time-sensitive questions. |
| Approach: | They propose to use existing temporal information extraction systems to construct temporal graphs of events, times, and temporal relations in questions and documents. |
| Outcome: | The proposed method outperforms graph convolution-based approaches on SituatedQA and TimeQA. |
Copied to clipboard
| Challenge: | Document dependency graphs (TDGs) are used to understand the temporal relations between events mentioned in a document and to improve downstream tasks such as timeline creation and time-aware summarization. |
| Approach: | They propose a temporal dependency graph parser that takes input from a text document and produces a graph that incorporates longer range dependencies. |
| Outcome: | The proposed framework outperforms existing models on three datasets and improves tasks such as timeline creation, time-aware summarization, and temporal information extraction. |
Copied to clipboard
| Challenge: | Recent approaches to sequence to sequence learning leverage recurrence, convolution, attention or combination of recurrent and convolutional neural networks. |
| Approach: | They propose an approach that extends the self-attention mechanism to consider representations of relative positions, or distances between sequence elements. |
| Outcome: | The proposed approach yields 1.3 BLEU and 0.3 BLUE on translation tasks . it is based on a relation-aware self-attention mechanism that can generalize to arbitrary graph-labeled inputs. |
Copied to clipboard
| Challenge: | Existing pre-trained models suffer from slow inference speed due to cross-modal attention in transformer architecture. |
| Approach: | They propose a multimodal approach that accelerates the inference time of ITR by thousands of times . they extract pre-cached feature indexes offline and employ instant dot-product matching online . |
| Outcome: | The proposed approach outperforms existing models that consume 1000 times magnitude of computational hours using the same features. |
Copied to clipboard
| Challenge: | Existing hybrids lack performance, latency, and cost-efficient scaling for production LLMs. |
| Approach: | They propose a deployment-oriented parallel hybrid architecture that enables deterministic conditional computation via FLOP-aware token circulation across attention and SSM branches. |
| Outcome: | FlowHN achieves 4 higher throughput and 15% higher MFU than current models while maintaining competitive accuracy on reasoning, coding, and long-context tasks. |
Copied to clipboard
| Challenge: | generative retrieval is a new paradigm for information retrieval, enabling a sequence-to-sequence model with a single Transformer . generative encoders have been used on small corpora, but only on large ones . |
| Approach: | They propose to encode an entire document corpus within a single Transformer . they find generative retrieval is competitive with state-of-the-art dual encoders on small corpora . |
| Outcome: | The proposed approach is competitive with state-of-the-art dual encoders on small corpora, the study finds . the proposed approach only evaluates on document corporales on the order of 100K in size . |
Copied to clipboard
| Challenge: | morphological inflection models have been successful with shared tasks . but they fail at generalizing inflation patterns when trained on a limited number of lemmata . |
| Approach: | They find that standard models fail at generalizing inflection patterns when trained on a limited number of lemmata and asked to inflect previously unseen lemma. |
| Outcome: | The proposed model can perform well on morphological inflection tasks if training data covers a diversity of lemmata or some variant of the input lemma has been witnessed during training. |
Copied to clipboard
| Challenge: | Syntactic Transformer language models aim to achieve better generalization through simultaneously modeling syntax trees and sentences. |
| Approach: | They propose a class of Transformer language models with explicit dependency-based inductive bias. |
| Outcome: | Experiments show that the proposed models outperform constituency-based models on sentences annotated with dependency trees and achieve better generalization. |
Copied to clipboard
| Challenge: | Neural machine translation models can benefit from modeling translated and untranslated source contents as recurrent states, but this less interpretable recurrence hinders their power to model dynamic updating of and contents during decoding. |
| Approach: | They propose to model the dynamic updating of and contents during decoding by explicitly separating source words into groups of translated and untranslated contents through parts-to-wholes assignment. |
| Outcome: | The proposed method achieves significant improvements over both Rnmt and Transformer by producing more adequate translations. |
Copied to clipboard
| Challenge: | Experimental results show that Transformer Encoder model can't automatically capture word order, so explicit position embeddings are required to be fed into the target model. |
| Approach: | They propose a Transformer-based language model DecBERT that uses a causal attention mask to capture word order. |
| Outcome: | The proposed model improves on the GLUE language understanding benchmark and accelerates the pre-training process. |
Copied to clipboard
| Challenge: | Existing parsing systems use local or global models of the parser state to improve performance. |
| Approach: | They propose to modify the sequence-to-sequence Transformer to model global or local parser states in transition-based parsing. |
| Outcome: | The proposed model significantly improves performance on dependency and Abstract Meaning Representation (AMR) parsing tasks. |
Copied to clipboard
| Challenge: | Recent work attempts to apply incremental processing to NLUs but this is computationally expensive and does not scale efficiently for long sequences. |
| Approach: | They propose to apply Transformers incrementally via restart-incrementality by repeatedly feeding, to an unchanged model, increasingly longer input prefixes to produce partial outputs. |
| Outcome: | The proposed model has better incremental performance and faster inference speed compared to the standard Transformer and LT with restart-incrementality, at the cost of part of the non-incremental quality. |
Copied to clipboard
| Challenge: | Neural Machine Translation models are influenced by two types of context, source and target, but none explicitly evaluates relative contribution to generation decision. |
| Approach: | They propose to adopt a variant of Layerwise Relevance Propagation which evaluates relative contributions to the generation decision by a proportion of token influence. |
| Outcome: | The proposed model can evaluate the relative contribution of source and target to the generation decision by using a variant of Layerwise Relevance Propagation (LRP) |
Copied to clipboard
| Challenge: | Code summarization (CS) is a promising area in recent language understanding . previous work using structurebased traversal or non-sequential models to learn structural program semantics has shown no performance gain . |
| Approach: | They propose to use a structure-based traversal model to learn structural program semantics to generate human language automatically for programming language in the format of source code. |
| Outcome: | Experiments show that the proposed method achieves state-of-the-art on benchmarks. |
Copied to clipboard
| Challenge: | Current studies prove that attention weights are not unique and therefore unfit for interpretation. |
| Approach: | They propose a transformer encoder layer that decouples the relationship between key and value vector and provides identifiable weights up to the desired length of the input. |
| Outcome: | The proposed model is more identifiable than previously thought but still prone to be non-unique attentions that make them unfit for interpretation. |
Copied to clipboard
| Challenge: | Chain-of-Thought prompting is a powerful technique for enhancing language model’s reasoning capabilities, but generating long and correct CoT trajectories is challenging. |
| Approach: | They propose to align the steps of Chain-of-Thought reasoning with loop iterations and apply intermediate supervision during the training of Looped Transformers. |
| Outcome: | The proposed method generates accurate reasoning chains for complex problems exceeding training length, and improves performance of the auto-regressive model. |
Copied to clipboard
| Challenge: | Existing text-based recommendation frameworks that use pretrained language models (PLMs) can improve performance on text-related tasks. |
| Approach: | They propose a unified local- and global-attention Transformer encoder to better model two-level contexts of user history. |
| Outcome: | The proposed framework improves on three text-based recommendation tasks. |
Copied to clipboard
| Challenge: | Pretrained language models use the attention mechanism to contextualize input inputs . but, we find that it is not as important as thought for pretrained models . |
| Approach: | They propose a probing method that replaces input-dependent attention matrices with constant ones. |
| Outcome: | The proposed method improves performance of pretrained language models without input-dependent attention. |
Copied to clipboard
| Challenge: | Existing models for matching dialogue responses rely on semantic and functional dependencies . a recent study only uses the last utterance in context for matching a reply . |
| Approach: | They propose a model that matches a response with its multi-turn context using attention. |
| Outcome: | The proposed model outperforms the state-of-the-art models on two large-scale multi-turn response selection tasks. |
Copied to clipboard
| Challenge: | Existing approaches to model long documents are difficult due to the quadratic complexity of text length. |
| Approach: | They propose a hierarchical interactive Transformer for efficient long document modeling. |
| Outcome: | Extensive experiments on three benchmark datasets validate the efficiency and effectiveness of Hi-Transformer in long document modeling. |
Copied to clipboard
| Challenge: | Existing solutions to reduce the cost of pretraining Transformer-based models are expensive especially for large-scale models. |
| Approach: | They propose to reduce the cost of pre-training Transformer-based models by compressing the sequence of hidden states inside Transformer architecture. |
| Outcome: | The proposed model achieves state-of-the-art on several Arabic downstream tasks despite using less computational resources compared to other BERT-based models. |
Copied to clipboard
| Challenge: | Generally, documents are truncated before being inputs to deep neural networks, resulting in missing keyphrases . evaluators use layer-wise coverage attention to cover all the critical points in a document . |
| Approach: | They propose a neural keyphrase generation model that identifies the salient sentences in a document and an extractor-generator that jointly extracts and generates keyphrases from the selected sentences. |
| Outcome: | The proposed model outperforms the state-of-the-art keyphrase generation methods on keyphrases generated from scientific and web documents. |
Copied to clipboard
| Challenge: | Medical imaging reports are time-consuming and can be error-prone for inexperienced radiologists. |
| Approach: | They propose to generate radiology reports with memory-driven Transformer using relational memory and memory-based conditional layer normalization. |
| Outcome: | The proposed method outperforms existing models on IU X-Ray and MIMIC-CXR . it generates long reports with medical terms and meaningful image-text attention mappings . |
Copied to clipboard
| Challenge: | a recent study has found that commonsense reasoning models are learning transferable generalizations . commonsensibility is a human capacity that has been a core challenge to Artificial Intelligence since its inception. |
| Approach: | They conduct an analysis of benchmarks that involve commonsense reasoning . they find that most datasets experimented with are problematic . commonsensence is a quintessential human capacity . |
| Outcome: | The proposed model is able to perform well on commonsense reasoning tasks . the model is not learning transferable generalizations or taking advantage of shortcuts . |
Copied to clipboard
| Challenge: | Transformers are impressive but inefficient and costly, which limits their applications and accessibility. |
| Approach: | They first use different ways to downsample and upsamplify activations in Transformers to make them hierarchical. |
| Outcome: | The proposed model outperforms Transformers on the ImageNet32 and enwik8 benchmarks. |
Copied to clipboard
| Challenge: | Existing vision-and-language pretraining approaches rely on external object detectors to encode images in a multi-modal transformer framework. |
| Approach: | They propose an object-aware end-to-end VLP framework which feeds image grid features from CNNs into the Transformer and learns the multi-modal representations jointly. |
| Outcome: | The proposed framework achieves competitive or superior performances on vision-language tasks. |
Copied to clipboard
| Challenge: | Existing methods to extract relation extraction from sentence are limited in focusing on leveraging dependency information. |
| Approach: | They propose dependency position encoding (DPE) that incorporates dependency connections and dependency types into the self-attention mechanism to distinguish the importance of different word dependencies. |
| Outcome: | The proposed method significantly outperforms the previous methods on SemEval 2010 Task 8, KBP37, and TACRED. |
Copied to clipboard
| Challenge: | Existing approaches to Multi-document summarization are limited due to the extremely long input length. |
| Approach: | They propose an extract-then-abstract Transformer framework to overcome the problem . they leverage pre-trained language models to construct hierarchical extractors and abstractors . |
| Outcome: | The proposed framework outperforms baseline models with comparable model sizes and achieves the best results on the Multi-News, Multi-XScience, and WikiCatSum corpora. |
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is a sequence tagging task that extracts named entities from unstructured text. |
| Approach: | They propose to integrate Chinese character features with radical-level embedding to improve Chinese NER by integrating Chinese character information. |
| Outcome: | The proposed method can improve Chinese Named Entity Recognition (NER) on well-known datasets. |
Copied to clipboard
| Challenge: | Existing studies show that the lack of recurrence modeling hinders the development of a translation model. |
| Approach: | They propose to model recurrence for Transformer with an additional recurrent encoder. |
| Outcome: | The proposed model outperforms the deep model on EnglishGerman and ChineseEnglish translation tasks. |
Copied to clipboard
| Challenge: | Standard decoders for neural machine translation generate a single token per timestep, which slows inference . a series of controlled experiments demonstrates that SynST decodes sentences 5x faster than the baseline autoregressive Transformer. |
| Approach: | They propose a syntactically supervised Transformer that generates all target tokens in one shot . synST is a variant of the Transformer architecture that autoregressively predicts a chunked parse tree . |
| Outcome: | The proposed method decodes sentences 5x faster than the baseline method on En-De and En-Fr datasets while achieving higher BLEU scores. |
Copied to clipboard
| Challenge: | Abstractive summarization models have been widely used to extract words from source into summary, but how to ensure that important words in source are copied remains a challenge. |
| Approach: | They propose a Transformer-based model to enhance copy mechanism by identifying the importance of each source word based on the degree centrality. |
| Outcome: | The proposed model outperforms baseline methods on CNN/Daily Mail and Gigaword datasets. |
Copied to clipboard
| Challenge: | Neural networks are at the center of a debate about human behavior in inflectional morphology. |
| Approach: | They measure correlation between human judgments and neural network probabilities for unknown word inflections. |
| Outcome: | The proposed model for morphological inflections correlates best with human wug ratings, but not with humans. |
Copied to clipboard
| Challenge: | Current approaches to speech-to-text translation (ST) use a pipeline of two sub-components - an automatic speech recognition (ASR) and a machine translation (MT) model. |
| Approach: | They propose an architecture that avoids initial lossy compression and aggregates information only at a higher level according to more informed linguistic criteria. |
| Outcome: | The proposed architecture achieves gains of up to 0.8 BLEU on the standard MuST-C corpus and up to 4.0 BLUE in a low resource scenario. |
Copied to clipboard
| Challenge: | Existing approaches to extract multiple relations from a paragraph require multiple passes over the paragraph. |
| Approach: | They propose a method to extract multiple relations from a paragraph by encoding the paragraph only once. |
| Outcome: | The proposed approach can perform state-of-the-art on the benchmark ACE 2005. |
Copied to clipboard
| Challenge: | Existing work on memory-efficient parallelisms to reduce time and space complexity focuses on reducing time and complexity from system perspective. |
| Approach: | They propose a memory-efficient parallelism to reduce time and space complexity . they split input sequence into multiple chunks and feed each chunk into GPU . |
| Outcome: | The proposed approach is compatible with most existing parallelisms and makes 4D parallelismal possible. |
Copied to clipboard
| Challenge: | Recent advances on neural approaches to natural language processing have triggered a resurgent interest on building intelligent open-domain chatbots. |
| Approach: | They propose a dialoguE COntradiction DEtection task and a conversational dataset . they show that their best contradiction detection model correlates well with human judgments . |
| Outcome: | The proposed model is more robust and generalizes well on analysis and out-of-distribution dialogues than standard (unstructured) Transformer models that explicitly hinge on utterance structures are more robust, the study shows . |
Copied to clipboard
| Challenge: | Existing research explores to enhance the two sublayers separately to improve the capability of Transformer for text representation. |
| Approach: | They propose to combine SAN and Feed-Forward Networks to create a dynamic mask attention network with a learnable mask matrix which can model localness adaptively. |
| Outcome: | The proposed model outperforms the original Transformer on translation and text summarization tasks. |
Copied to clipboard
| Challenge: | Experimental results show that Generative pre-trained Transformers (GPT) have great success in natural language processing. |
| Approach: | They propose a unified language model of text and molecules pre-trained on SMILES wrapped by text. |
| Outcome: | The proposed model outperforms strong baselines of molecular property prediction on MoleculeNet and performs comparably to the best model in text-molecule translation while using less than half of its parameters. |
Copied to clipboard
| Challenge: | Existing models ignore the inherent causality during related work generation, leading to spurious correlations which downgrade the models’ generation quality and generalizability. |
| Approach: | They propose a Causal Intervention Module for Related Work Generation (CaM) that captures causal relationships in related work generation and implements causal interventions to mitigate the negative impact of spurious correlations. |
| Outcome: | The proposed framework improves the quality and coherence of generated related work by capturing causalities in the generation process. |
Copied to clipboard
| Challenge: | Despite the success of sequence-to-sequence models, dialogue logics are often ignored. |
| Approach: | They propose a network architecture to explore the current dialog context and similar dialogue instances’ logical structure simultaneously. |
| Outcome: | The proposed network architecture is superior to existing state-of-the-art models. |
Copied to clipboard
| Challenge: | eschewing separate architecture and training for knowledge-intensive tasks is cumbersome . end-to-end training only based on supervision from the end task is awkward . |
| Approach: | They propose a single Transformer that performs retrieval as attention and end-to-end training solely based on supervision from the end QA task. |
| Outcome: | The proposed model outperforms state-of-the-art retrievers and readers on in-domain datasets. |
Copied to clipboard
| Challenge: | BERT is a promising technique to improve NMT, but how it outperforms standard NMT is understudied. |
| Approach: | We compare MT engines trained with pre-trained BERT and back-translation with incrementally larger amounts of data. |
| Outcome: | The proposed technique outperforms standard NMT models on morphology and syntax. |
Copied to clipboard
| Challenge: | Using mixture-of-experts (MoE) to deal with language heterogeneity is a challenge in neural machine translation (NMT). |
| Approach: | They propose a lightweight MoE-based NMT model that is trained via an elaborate stage-wise training strategy. |
| Outcome: | The proposed model achieves stable improvements in translation tasks by introducing fewer extra parameters compared to baseline models. |
Copied to clipboard
| Challenge: | Neural document rerankers require dedicated hardware for serving, which is costly and often not feasible. |
| Approach: | They propose a method that captures 86% of the gains of a Transformer cross-attention model with a lexicalized scoring function that only requires 10-6% of . the model architecture is compatible with recent encoder-decoder and decoder-only large language models, such as T5, GPT-3 and PaLM. |
| Outcome: | The proposed model captures 86% of the gains of a Transformer cross-attention model with a lexicalized scoring function. |
Copied to clipboard
| Challenge: | Existing methods to analyze malware behavior only disclose a subset of behaviors due to inherent difficulties. |
| Approach: | They propose a novel malware behavior search technique that is based on graph isomorphism at the attention layers of Transformer models. |
| Outcome: | The proposed technique outperforms state-of-the-art methods in a case study of 10 real-world malwares by 6-14%. |
Copied to clipboard
| Challenge: | Existing methods to improve performance of pre-trained language models are limited due to large-scale parameters and the universal autoregressive decoding paradigm. |
| Approach: | They propose a novel fine-tuning method which can make a single pre-trained model support Dynamic and Efficient infERence and achieve an adaptive trade-off between model performance and latency. |
| Outcome: | The proposed method achieves higher BLEU scores than the strong autoregressive Transformer model on translation tasks with 3 12 times speedup and faster inference speed compared with the BART model on four GLGE benchmark tasks. |
Copied to clipboard
| Challenge: | a limited human translation budget is required to train neural machine translation models. |
| Approach: | They propose to integrate active learning into neural machine translation techniques . they propose a word frequency based acquisition function and an uncertainty based method . |
| Outcome: | The proposed method outperforms other acquisition functions on a limited human translation budget. |
Copied to clipboard
| Challenge: | Existing work on how Transformers can solve synthetic tasks has not explored how to extend this to a conversational setting. |
| Approach: | They propose to use ELIZA as a framework for formal mechanistic analysis of Transformers . they propose to model local pattern matching and long-term dialogue state tracking . |
| Outcome: | The proposed model can be extended to model key aspects of conversation, the authors show . their model favors an induction head mechanism over a more precise copying mechanism . |
Copied to clipboard
| Challenge: | Existing models that capture token distances are not optimal for modeling the orders and relations of contexts. |
| Approach: | They propose a distance-aware Transformer that can exploit the real distances between tokens to re-scale the raw self-attention weights. |
| Outcome: | The proposed model outperforms the existing Transformer and its variants on five benchmark datasets and can improve the performance of many tasks. |
Copied to clipboard
| Challenge: | Neural Machine Translation (NMT) has been replaced by convolutional or self-attentional approaches. |
| Approach: | They propose an architecture definition language that allows for a flexible combination of common building blocks. |
| Outcome: | The proposed architectures can bring recurrent and convolutional models close to the Transformer architecture, but not using self-attention. |
Copied to clipboard
| Challenge: | Existing models with seq2seq framework lack ability to effectively manage concept transitions . lack of concept management strategies might lead to incoherent dialogue due to loosely connected concepts . |
| Approach: | They propose a concept-guided non-autoregressive model for open-domain dialogue generation that learns to identify multiple associated concepts from a conceptual graph and a customized Insertion Transformer to perform concept-directed generation to complete a response. |
| Outcome: | The proposed model outperforms state-of-the-art models in automatic and human evaluations with substantially faster inference speed. |
Copied to clipboard
| Challenge: | GINCO is a new training dataset for automatic genre identification based on 1,125 crawled Slovenian web documents that consist of 650,000 words. |
| Approach: | They propose to use 1,125 crawled Slovenian web documents to train a new genre classification system based on a GINCO training dataset . |
| Outcome: | The proposed classifiers perform better on the 1,125 crawled Slovenian web documents than the existing models and achieve higher scores on the task. |
Copied to clipboard
| Challenge: | Existing sparse attention methods use fixed patterns to select words without considering similarities between words. |
| Approach: | They propose a neural clustering method which integrates into the Self-Attention Mechanism in Transformer and integrates it into the target task. |
| Outcome: | The proposed method outperforms two typical sparse attention methods on translation, text classification, and text matching tasks while having a comparable or even better time and memory efficiency. |
Copied to clipboard
| Challenge: | Recent studies have shown that Transformers is implicitly learning syntactic information from data, albeit is highly dependent on the quality and scale of the training data. |
| Approach: | They propose a syntax-guided localized self-attention model that allows directly incorporating grammar structures from an external constituency parser. |
| Outcome: | The proposed model improves translation performance on a variety of datasets, from small to large datasets and with different source languages. |
Copied to clipboard
| Challenge: | Existing work exploits the reordering information in neural machine translation . experimental results show that the proposed methods can significantly improve the performance of the transformer translation system. |
| Approach: | They propose a reordering mechanism to learn the re ordering embedding of a word based on contextual information and stack them together with self-attention networks to learn sentence representation for machine translation. |
| Outcome: | The proposed method improves translation performance on English-to-German, NIST Chinese-to English, and WAT Japanese-toEnglish translation tasks. |
Copied to clipboard
| Challenge: | Dynamic neural networks can scale up pretrainable models with sub-linear increases in computation and time. |
| Approach: | They summarize the progress of three types of dynamic neural networks in NLP . skimming, mixtures of experts, and early exit are among the most popular . |
| Outcome: | The proposed models can scale up with sub-linear increases in computation and time . skimming, mixture of experts, and early exit are the most popular approaches . |
Copied to clipboard
| Challenge: | Aerial visionand-dialling navigation (AVDN) is a new approach to autonomous drones that can converse with humans and follow natural language commands to complete tasks. |
| Approach: | They propose to use Aerial Visionand-Dialog Navigation (AVDN) to navigate a drone via natural language conversation by collecting a dataset of over 3k recorded navigation trajectories with asynchronous human-human dialogs between commanders and followers. |
| Outcome: | The proposed system can converse with humans and follow natural language commands to fly to the expected destination. |
Copied to clipboard
| Challenge: | Existing compositional generalization benchmarks focus on lexical generalisation, the interpretation of novel lexicals in syntactic structures familiar from training. |
| Approach: | They propose a semantic parsing dataset that extends COGS with 17 structural generalization cases to evaluate how well models generalize to new complex linguistic expressions. |
| Outcome: | The proposed model generalization accuracy is far below the near-perfect accuracy of existing models on COGS, demonstrating the role of SLOG in foregrounding the large discrepancy between models’ lexical and structural generalization capacities. |
Copied to clipboard
| Challenge: | Document-level contextual information has shown benefits to text-based machine translation, but whether and how it helps end-to-end speech translation is still under-studied. |
| Approach: | They propose a concatenation-based ST model with adaptive feature selection for computational efficiency. |
| Outcome: | The proposed model improves translation quality and robustness to (artificial) audio segmentation errors. |
Copied to clipboard
| Challenge: | Singing Voice Synthesis (SVS) synthesizes pleasing vocals based on music scores and lyrics . current acoustic models ignore the significance of local modeling within the sequence and the hard-to-synthesize parts in the predicted mel-spectrogram . |
| Approach: | They propose a method to enhance local modeling in the acoustic model by focusing on phoneme tokens located before and after the phoneme. |
| Outcome: | The proposed method improves local modeling in the acoustic model by focusing on the hard-to-synthesize parts of the predicted mel-spectrogram. |
Copied to clipboard
| Challenge: | Neural machine translation models are trained to maximize the likelihood of the next token given previous golden tokens as inputs, but at the inference stage, golden token is unavailable. |
| Approach: | They propose a scheduled sampling method that randomly replaces groundtruth tokens with predicted ones during training, ignoring real-time model competence. |
| Outcome: | The proposed method outperforms the Transformer and vanilla scheduled sampling on large-scale translations. |
Copied to clipboard
| Challenge: | Intent detection and slot filling are two main tasks for building a spoken language understanding system. |
| Approach: | They propose a framework to incorporate intent information into slot filling tasks . they use a joint model with Stack-Propagation to capture intent semantic knowledge . |
| Outcome: | The proposed model outperforms existing models on two publicly available datasets and outperformed existing models by a large margin. |
Copied to clipboard
| Challenge: | a new study investigates the quality and novelty of generated paraphrases . paraphrase models can be used for information retrieval and data mining . |
| Approach: | They use state-of-the-art neural machine translation models trained on the Opusparcus corpus to generate paraphrases in six languages. |
| Outcome: | The proposed model outperforms the existing model on human evaluation in five of the six languages. |
Copied to clipboard
| Challenge: | Existing work on rule mining focuses on mining rules, but how to select appropriate rules for completion of different triplets has not been discussed. |
| Approach: | They propose to take context information into consideration when selecting suitable rules . they devise a transformer-based rule mining approach, Ruleformer . |
| Outcome: | The proposed model takes context information into consideration, which helps select suitable rules for inference tasks. |
Copied to clipboard
| Challenge: | Biomedical data and benchmarks are highly valuable but limited in low-resource languages such as English. |
| Approach: | They propose a translation model in Vietnamese that trains a pretrained Encoder-Decoder Transformer model on 20 million translated abstracts. |
| Outcome: | The proposed model can translate and produce both pretrained and supervised biomedical data in two biomedically important domains. |
Copied to clipboard
| Challenge: | Multi-head self-attention-based Transformers have shown promise in different learning tasks . but encoders of Transformers and their variants fail to preserve layer-wise contextual information . |
| Approach: | They propose an encoder model that guarantees a theoretical bound for layer-wise distance preservation between a pair of tokens. |
| Outcome: | The proposed model preserves equivalence between tokens and performs better than Transformers. |
Copied to clipboard
| Challenge: | Existing glyph-based models neglect the relationship between pictorial elements and radicals for Named Entity Recognition (NER) tasks. |
| Approach: | They propose a model that integrates multi-source visual and phonetic information of Hanzi . they propose combining pictographic features with radicals to facilitate integration . |
| Outcome: | The proposed model improves performance on benchmark datasets. |
Copied to clipboard
| Challenge: | Pre-trained seq2seq models suffer from a prediction bias due to their unidirectional decoding. |
| Approach: | They propose a bidirectional Transformer reranker that re-estimates the probability of each candidate sentence generated by pre-trained seq2seq models. |
| Outcome: | The proposed model improves on the original model and gives a 59.52 GLEU score on the JFLEG corpus. |
Copied to clipboard
| Challenge: | Practical applications of automatic image description systems include leveraging descriptions for image indexing or retrieval, and helping those with visual impairments by transforming visual signals into information that can be communicated via text-to-speech technology. |
| Approach: | They propose to extract and filter image caption annotations from billions of webpages and use them to train models. |
| Outcome: | The proposed model architectures perform better when trained on the Conceptual Captions dataset. |
Copied to clipboard
| Challenge: | Recent Transformer based models have achieved state-of-the-art performance for many natural language processing tasks including machine translation, question-answering tasks and semantic role labeling. |
| Approach: | They propose to reduce the number of parameters of BERT to obtain a much efficient light model. |
| Outcome: | The proposed model achieves 6.6% higher average accuracy on GLUE and SQuAD datasets than the previous model with three encoder layers while having the same number of parameters. |
Copied to clipboard
| Challenge: | Existing studies on pre-trained Transformers show that they learn fine-grained neuron functions. |
| Approach: | They examine the presence of modularity in pre-trained Transformers . they focus on Mixture-of-Experts, a promising candidate for modularity . |
| Outcome: | The proposed structure stabilizes at the early stage, which is faster than neuron stabilization. |
Copied to clipboard
| Challenge: | Unsupervised methods for dialogue topic segmentation are difficult to surpass due to short sentences, serious references and non-standard language. |
| Approach: | They propose a method to divide a dialogue into different topic paragraphs to better understand its structure and content. |
| Outcome: | The proposed method achieves the best results on multiple benchmark datasets across different scenarios. |
Copied to clipboard
| Challenge: | Existing models that use natural language rationales provide intuitive, higher-level explanations that are easily understandable by humans. |
| Approach: | They propose a model that generates free-text rationales by combining pretrained language models with object recognition, grounded visual semantic frames, and visual commonsense graphs. |
| Outcome: | The proposed model generates free-text rationales by combining pretrained language models with object recognition, grounded visual semantic frames, and visual commonsense graphs. |
Copied to clipboard
| Challenge: | Neural machine translation systems require a number of stacked layers for deep models, but the prediction depends on the sentence representation of the top-most layer with no access to low-level representations. |
| Approach: | They propose a multi-layer representation fusion approach to fusing stacked layers to learn a better representation from the stack. |
| Outcome: | The proposed approach yields 0.92 and 0.56 BLEU points over the strong Transformer baseline on IWSLT German-English and NIST Chinese-English MT tasks respectively. |
Copied to clipboard
| Challenge: | Recent research questions the importance of dot-product self-attention in Transformer models and shows that most attention heads learn simple positional patterns. |
| Approach: | They propose a novel mechanism to replace dot-product self-attention with a recurrent atteNtion mechanism that directly learns attention weights without token-to-token interaction. |
| Outcome: | The proposed model outperforms the Transformer model on translation tasks with fewer parameters and inference time. |
Copied to clipboard
| Challenge: | Existing grammar induction methods do not provide sufficient performance in downstream tasks. |
| Approach: | They propose an unsupervised grammar induction method for language understanding and generation using a grammar parser and a syntactic mask. |
| Outcome: | The proposed method performs better on from-scratch and pre-trained scenarios. |
Copied to clipboard
| Challenge: | Existing methods for Neural Machine Translation (NMT) have been proven effective in improving the performance of computer vision tasks without pre-training a teacher. |
| Approach: | They propose a rank-order augmented Pearson correlation loss and an iterative distillation method to prevent the discrepancy of predictions between the student and a stronger teacher from disturbing the training. |
| Outcome: | The proposed method can lead to significant improvements over the strong Transformer baseline on low/middle/high-resource tasks, obtaining comparable or better performance with fewer layers. |
Copied to clipboard
| Challenge: | Existing work on societal bias in NLP focuses on race and gender . linguistic background is a unique attribute that has been largely ignored in the field . |
| Approach: | They examine linguistic background to craft plausible adversarial examples that expose biases in popular NLP models. |
| Outcome: | The proposed model improves robustness without sacrificing performance on clean data. |
Copied to clipboard
| Challenge: | Recent work on structured prediction has produced very effective supervised clustering algorithms using linear classifiers. |
| Approach: | They propose to use latent structured prediction loss and Transformer models to approach supervised clustering. |
| Outcome: | The proposed approach outperforms the state-of-the-art in recreating intents from public question corpora. |
Copied to clipboard
| Challenge: | a recent study has attempted to decode linguistic structure from the Transformer . but, much of the work focused on English, a language with rigid word order and a lack of inflectional morphology. |
| Approach: | They propose to fine-tune a feature encoder for BERT to learn linguistic structure from its multi-head attention mechanism. |
| Outcome: | The proposed model can decode full trees above baseline accuracy from single attention heads across languages. |
Copied to clipboard
| Challenge: | Existing models for generating instructions for navigation generate references to objects or actions that are inconsistent with what a human follower would perform or encounter along the path. |
| Approach: | They propose a weakly supervised approach that detects hallucinated references by using a pre-trained vision-language model. |
| Outcome: | The proposed model outperforms baseline models and supervised models on generating navigation instructions. |
Copied to clipboard
| Challenge: | Existing work extends translation unit from single sentence to multiple sentences. |
| Approach: | They propose to introduce locality assumption as an inductive bias into Transformer and reduce the hypothesis space of attention from target to source. |
| Outcome: | The proposed model achieves state-of-the-art BLEU scores on three benchmark datasets. |
Copied to clipboard
| Challenge: | Existing benchmarks on long-range attention models have not been sufficient to develop efficient Transformers and their practical application on complex NLP tasks. |
| Approach: | They propose to benchmark 7 Transformer variants on 5 difficult NLP tasks and 7 datasets to examine their capacity for long-range attention. |
| Outcome: | The proposed models have advantages on content selection and query-guided decoding, but they come with previously unrecognized drawbacks such as insufficient attention to distant tokens and accumulated approximation error. |
Copied to clipboard
| Challenge: | Existing work has shown that Transformer embeddings are anisotropic, which is called the representation degradation problem. |
| Approach: | They identify a set of Transformer models with isotropic embedding spaces, the large Pythia models. |
| Outcome: | The proposed model sets show that isotropic models do not develop as previously theorized. |
Copied to clipboard
| Challenge: | Neural machine translation models are weak enough for document-level translation . current models only translate sentences individually, resulting in poor document coherence . |
| Approach: | They propose to use the original Transformer model to test document-level neural machine translation . they find that the original transformer models can achieve strong results for document translation if trained properly . |
| Outcome: | The proposed model outperforms sentence-level models on nine datasets and two sentence- level datasets across six languages. |
Copied to clipboard
| Challenge: | Neural machine translation (NMT) takes deterministic sequences for source representations. However, word-level or subword-level segmentation has multiple choices to split a source sequence with different word segmentors or different subword vocabulary sizes. |
| Approach: | They propose lattice-based encoders to explore effective word or subword representations in an automatic way during training. |
| Outcome: | The proposed encoders can explore effective word or subword representation in an automatic way during training. |
Copied to clipboard
| Challenge: | Recent studies show that publicly shared gradients in the training process can reveal the private training data to a third-party. |
| Approach: | They propose a gradient attack algorithm to reconstruct the local training data using GLUE benchmarks. |
| Outcome: | The proposed algorithm achieves 1.5x recover rate and 2.5x ROUGE-2 over previous methods without the need of ground truth label. |
Copied to clipboard
| Challenge: | Variational Auto-Encoders (VAEs) for learning disentangled latent representations in speech fail to learn latent clusters of speaker attributes when trained on limited or noisy datasets. |
| Approach: | They propose a Variational Auto-Encoder (VAE) that minimizes mutual information between latent variables and learns controllable latent representations in speech data. |
| Outcome: | The proposed model reduces the cluster overlap of speaker attributes by 30% over LSTM-VAE. |
Copied to clipboard
| Challenge: | Recent work in context-aware NMT considers only a few previous sentences as context . current systems fail to achieve fluent, good quality translation for a full document . |
| Approach: | They propose a top-down approach to hierarchical attention for context-aware NMT which uses sparse attention to selectively focus on relevant sentences in the document context. |
| Outcome: | The proposed approach outperforms context-agnostic baselines and context-based baselines on English-German datasets. |
Copied to clipboard
| Challenge: | Recent work on Chinese word segmentation has been concerned about the following three perspectives. |
| Approach: | They propose to use a greedy decoding algorithm to improve Chinese word segmentation model. |
| Outcome: | The proposed model achieves state-of-the-art or comparable performance against strong baselines in strict closed test setting. |
Copied to clipboard
| Challenge: | Existing work on event argument extraction (EE) is limited due to data scarcity and lack of a model encoder. |
| Approach: | They propose to capture the long-range dependency between an event trigger and a distant event argument using unlabeled data. |
| Outcome: | Experiments on the English ACE 2005 benchmark show that the proposed method achieves a new state-of-the-art. |
Copied to clipboard
| Challenge: | Neural machine translation systems fail on less decent inputs, which may harm the credibility of these systems. |
| Approach: | They propose a paradigm that generates adversarial examples using reinforcement learning to expose pitfalls for a given performance metric. |
| Outcome: | The proposed paradigm produces stable attacks with meaning-preserving adversarial examples. |
Copied to clipboard
| Challenge: | Recent studies have focused on improving dialogue generation models that include knowledge related to the posts. |
| Approach: | They propose to use a novel method to generate responses from posts and related knowledge by injecting knowledge into dialogue generation models. |
| Outcome: | The proposed method outperforms baseline models in terms of knowledge relevance and quality. |
Copied to clipboard
| Challenge: | Text style transfer is the task of transferring the style of text having certain stylistic attributes, while preserving non-stylistic or content information. |
| Approach: | They propose a new approach to rewriting sentences to a target style in the absence of parallel style corpora by exploiting the Transformer. |
| Outcome: | The proposed method outperforms state-of-the-art systems across 5 datasets on sentiment, gender and political slant transfer. |
Copied to clipboard
| Challenge: | Existing studies use pseudo cross-lingual abstractive summarization data to train neural encoder-decoders. |
| Approach: | They propose a multi-task learning framework for cross-lingual abstractive summarization that attaches a special token to the beginning of the input sentence to indicate the target task. |
| Outcome: | The proposed model achieves better performance than the model trained with only pseudo cross-lingual abstractive summarization data. |
Copied to clipboard
| Challenge: | Recent advances in NMT have shown promising results but are vulnerable to noise. |
| Approach: | They propose a data-driven technique called Target Augmented Fine-tuning to incorporate noise during training. |
| Outcome: | The proposed techniques perform with no degradation where up to 10% of entire test words are infected by noise. |
Copied to clipboard
| Challenge: | Compared to previous studies, the performance of neural models is likely to be affected by the choice of hyper-parameters. |
| Approach: | They propose to automatically and dynamically determine batch sizes by accumulating gradients of mini-batches and performing an optimization step at just the time when the direction of gradients starts to fluctuate. |
| Outcome: | The proposed approach improves the Transformer model with a fixed 25k batch size by +0.73 and +0.82 BLEU respectively. |
Copied to clipboard
| Challenge: | Existing Diffusion Language Models lack a structural constraint to stabilize attention sinks. |
| Approach: | They propose a simple but effective extra sink token that is constrained to attend to itself while remaining globally visible to all other tokens. |
| Outcome: | The proposed token is able to stabilize attention sinks and improve model performance. |
Copied to clipboard
| Challenge: | Document-level Neural Machine Translation aims to increase the quality of neural translation models by taking into account contextual information. |
| Approach: | They propose to use document-level corpus for Basque-Spanish language pairs to take into account contextual information and perform fine-grained evaluations of gender and gender. |
| Outcome: | The proposed corpus is suitable for fine-grained evaluation of document-level machine translation systems. |
Copied to clipboard
| Challenge: | Existing models based on textual data do not capture context beyond the sentence. |
| Approach: | They propose a framework that enables the model to learn multi-omnics biological information about entities (proteins) with the help of additional multi-modal cues like molecular structure. |
| Outcome: | The proposed model is generalized and optimized for protein-protein interaction task and benefited from additional domain-specific cues. |
Copied to clipboard
| Challenge: | Transformer-based pre-trained models achieve state-of-the-art results, but they can be prohibitively costly. |
| Approach: | They propose a fine- and coarse-granularity hybrid self-attention that shortens the computational sequence length in self- attention by progressively shortening the computational time. |
| Outcome: | The proposed model reduces computation cost by shortening the computational sequence length in self-attention. |
Copied to clipboard
| Challenge: | Existing automatic headline generation methods cannot include a given phrase in the generated headline. |
| Approach: | They propose a Transformer-based method that guarantees to include a given phrase in a generated headline. |
| Outcome: | The proposed method achieves ROUGE scores comparable to previous methods with Japanese news corpus. |
Copied to clipboard
| Challenge: | Current sentence encoders are word order sensitive, resulting in poor performance . Adapting word order from one language to another is key in cross-lingual structured prediction. |
| Approach: | They propose a new module to organize words following the source language order . they build structured prediction models with bag-of-words inputs and introduce a module to do this . |
| Outcome: | The proposed model significantly improves target language performance for languages that are distant from the source language. |
Copied to clipboard
| Challenge: | Existing NMT models are shallow in comparison to convolutional models used for both text and vision tasks. |
| Approach: | They propose to modify the attention mechanism to ease the optimization of deeper models by a simple modification to the seq2seq with attention paradigm. |
| Outcome: | The proposed model achieves consistent gains of 0.7-1.1 BLEU on the benchmark WMT’14 English-German and WMT'15 Czech-English tasks. |
Copied to clipboard
| Challenge: | Recent studies have found that information relevant to the next token prediction task accumulates in the hidden representations of just a few tokens. |
| Approach: | They propose a method that integrates attention preferences useful for a downstream task into the eviction process of hidden states. |
| Outcome: | The proposed method performs better on comprehension and retrieval tasks while preserving language modeling perplexity. |
Copied to clipboard
| Challenge: | Existing paradigms for semantic parsing are sequence-to-sequence and AMR parsers. |
| Approach: | They propose to formulate parsing as a sequence-to-sequence task using graph-based decoding techniques developed for syntactic parsers. |
| Outcome: | The proposed approach is competitive with sequence decoders on the standard setting and offers significant improvements in data efficiency and data availability. |
Copied to clipboard
| Challenge: | a new method for learning unsupervised sentence embeddings is proposed . unsup-SimCSE is biased because of the length information encoded into the sentence embeds . |
| Approach: | They propose a new unsupervised sentence embedding method that uses dropout to obtain positive pairs from a pre-trained Transformer encoder. |
| Outcome: | The proposed method outperforms the state-of-the-art unsup-SimCSE on a STS task. |
Copied to clipboard
| Challenge: | Recent improvements in NLP tasks can be attributed to the Transformer model. |
| Approach: | They propose to use parameter-sharing methods to reduce parameter budgets in generative models by using sandwich-style parameter sharing and self-attentive embedding factorization. |
| Outcome: | The proposed model outperforms the current RNN model even with significantly fewer parameters. |
Copied to clipboard
| Challenge: | Existing models for encoding long sequences in deep learning suffer from high latency and memory demands. |
| Approach: | They propose a clustering-based sparse Transformer framework to perform attention across chunked sequences. |
| Outcome: | The proposed framework achieves state-of-the-art on several major QA benchmarks. |
Copied to clipboard
| Challenge: | Improving Transformer efficiency has become increasingly attractive in recent years. |
| Approach: | They propose to combine pruning, quantization, new architectures and training strategies to improve Transformer efficiency. |
| Outcome: | The proposed methods improve the inference efficiency of a strong Transformer system by 3.80x on CPU and 2.52x on GPU. |
Copied to clipboard
| Challenge: | Dialogue systems using deep learning have achieved generation of fluent response sentences to user utterances, but they tend to produce responses that are not diverse and less context-dependent. |
| Approach: | They propose an Inverse N-gram loss function which incorporates contextual fluency and diversity at the same time by a simple formula. |
| Outcome: | The proposed loss function outperforms baseline models in automatic evaluations such as DIST-N and ROUGE and achieves higher scores on human evaluations of coherence and richness. |
Copied to clipboard
| Challenge: | Transformer architecture is composed of multi-head attention, which has been extensively analyzed. |
| Approach: | They extended the scope of the analysis of Transformers from solely the attention patterns to the whole attention block, i.e., multi-head attention, residual connection, and layer normalization. |
| Outcome: | The proposed method incorporates the whole attention block, i.e., multi-head attention, residual connection, and layer normalization into the analysis. |
Copied to clipboard
| Challenge: | Low-resource language translation is a challenging but socially valuable NLP task. |
| Approach: | They propose a normalization technique that modifies the attention mechanism to make the softmax function less prone to arbitrary saturation without sacrificing expressivity. |
| Outcome: | The proposed technique improves 0.928 BLEU over state-of-the-art bilingual benchmarks for 5 low-resource translation pairs from the TED Talks corpus and IWSLT’15. |
Copied to clipboard
| Challenge: | Existing methods for extracting interpersonal relationships from dialogues are limited to end-to-end learning. |
| Approach: | They propose a neural multi-label classifier that infers relationships from dialogues by external knowledge about speaker features and conversation style. |
| Outcome: | The proposed method outperforms the state-of-the-art methods on large-scale datasets with directed relationships of conversation participants. |
Copied to clipboard
| Challenge: | Abstractive summarization is the task of generating a concise summary of input documents . a middle-aged man and a young girl died after they were unable to avoid the plane . |
| Approach: | They propose a model that enriches the original Transformer with a Tensor Product Representation for abstractive summarization. |
| Outcome: | The proposed model outperforms the Transformer and the original TP-Transformer significantly on several datasets. |
Copied to clipboard
| Challenge: | Recent years have witnessed the successful application of natural language generation. |
| Approach: | They propose a model that uses user and item IDs to predict the words in the target explanation to make personalized Transformer. |
| Outcome: | The proposed model outperforms BERT on the explainable recommendation task in terms of effectiveness and efficiency. |
Copied to clipboard
| Challenge: | Neural network-based language models (LMs) have been shown to learn relevant properties of language without being explicitly trained for them. |
| Approach: | They extend their previous work to analyze whether language models capture anaphoric relations and pronoun-antecedent relations in English. |
| Outcome: | The Transformer outperforms the LSTM in all analyses. |
Copied to clipboard
| Challenge: | In the Transformer model, “self-attention” combines information from attended embeddings into the representation of the focal embeddable in the next layer. |
| Approach: | They propose two methods to quantify flow of information through self-attention using attention weights as relative relevance of input tokens. |
| Outcome: | The proposed methods give complementary views on the flow of information and yield higher correlations with importance scores of input tokens. |
Copied to clipboard
| Challenge: | Forced labour is the most common type of modern slavery, affecting at least 24.9 million people worldwide. |
| Approach: | They propose to annotate an English corpus for multi-class and multi-label forced labour detection using specialised data from specialised sources. |
| Outcome: | The proposed corpus consists of 989 news articles annotated according to risk indicators defined by the International Labour Organization (ILO). |
Copied to clipboard
| Challenge: | Recursive noun phrases have interesting semantic properties, yet it is unknown whether language models have such knowledge. |
| Approach: | They propose a dataset of three textual inference tasks targeting recursive noun phrases . they show that such knowledge is learnable with appropriate data . |
| Outcome: | The proposed model achieves strong zero-shot performance on an extrinsic Harm Detection task. |
Copied to clipboard
| Challenge: | Existing evaluation metrics do not capture meeting-specific errors, leading to ineffective assessment. |
| Approach: | They examine the relationship between established metrics and human evaluations to determine what challenges and errors are captured by correlating metric scores with human evaluation. |
| Outcome: | The proposed measures show weak correlations with human evaluations and a third of the correlations show error masking. |
Copied to clipboard
| Challenge: | Neural machine translation suffers from exposure bias and error propagation problem. |
| Approach: | They conduct a series of analyses to deeply understand the accuracy drop problem . they find that the left part of the translated sentence is often better than its right part . |
| Outcome: | The results show that the left part of the translated sentence is often better than its right part in left-to-right decoding models. |
Copied to clipboard
| Challenge: | Query translation (QT) is a critical factor in successful cross-lingual information retrieval (CLIR). |
| Approach: | They propose to extend query translation (QT) with a domain transfer procedure to revise synthetic candidates to search-aware examples. |
| Outcome: | The proposed method outperforms baselines and domain transfer methods on translation quality and retrieval accuracy. |
Copied to clipboard
| Challenge: | Existing methods to incorporate information from other modality, usually static images, are not considered relative to multimodal machine translation. |
| Approach: | They propose a multimodal self-attention method which learns the representation of images based on the text, which avoids encoding irrelevant information in images. |
| Outcome: | The proposed model outperforms previous studies and competitive baselines in terms of various metrics. |
Copied to clipboard
| Challenge: | Existing Transformer Architecture Search methods are limited to computer vision and natural language processing tasks. |
| Approach: | They propose a Transformer Architecture Search proxy that measures trainability and expressivity of Transformer networks separately and integrates it into an effective regularized evolution framework to demonstrate its efficacy. |
| Outcome: | The proposed proxy can achieve higher correlation with the true performance of Transformer networks on computer vision and natural language processing tasks. |
Copied to clipboard
| Challenge: | ELECTRA is more accurate than BERT, but it is not clear if this is due to its innovative architecture or to the long and extensive training, which highly increases the computation cost for obtaining the final language model. |
| Approach: | They propose to replace BERT’s Masked Language Modeling objective (MLM) with Token Detection (TD) by using a statistical approach to generate light tokens. |
| Outcome: | The proposed method can replace ELECTRA's computationally heavy generators without a significant drop in performance. |
Copied to clipboard
| Challenge: | Existing methods only conduct network growth in a single dimension, but compound growth operators are beneficial for multiple dimensions. |
| Approach: | They propose a method to train BERT progressively using a Transformer model and explore alternative growth operators in each dimension via controlled comparison. |
| Outcome: | The proposed method speeds up BERT pre-training by 73.6% and 82.2% for the base and large models respectively while achieving comparable performances. |
Copied to clipboard
| Challenge: | Existing methods for multimodal sentiment analysis require all modalities as input, thus are sensitive to missing modality at predicting time. |
| Approach: | They propose to model bi-direction interplay via couple learning and exploit multiple bi-directional translations to exploit multimodal fusion embeddings. |
| Outcome: | The proposed framework achieves state-of-the-art or often competitive performance on two multimodal benchmarks with extensive ablation studies. |
Copied to clipboard
| Challenge: | Existing approaches to graph representation only consider the local neighbors, sacrificing the Transformer’s ability to attend to elements at any distance. |
| Approach: | They propose a dual-encoding Transformer architecture that uses a structural encoder and a semantic encoder to seek for semantically relevant nodes. |
| Outcome: | The proposed architecture achieves superior performance compared to state-of-the-art attention-based methods on complex relational graphs like KGs and citation networks. |
Copied to clipboard
| Challenge: | Attention-based language models rely on the softmax function to convert attention logits into probability distributions, but this process can result in attention entropy collapse. |
| Approach: | They propose to use the softmax function to re-weight attention logits to create probability distributions, but this reweighting can lead to attention entropy collapse . they find that entropic-stable attention methods can prevent entrapment and enable more stable training by controlling or insensitive to variance of attention logit variance. |
| Outcome: | The proposed methods prevent attention entropy collapse and enable more stable training. |
Copied to clipboard
| Challenge: | Automated fact checking is rapidly gaining attention of the NLP and AI communities. |
| Approach: | They propose lightweight strong baselines for automated fact-checking systems . they propose to combine multiple pieces of evidence to verify a claim . |
| Outcome: | The proposed methods outperform heavier models on the leaderboard with blind TEST set. |
Copied to clipboard
| Challenge: | Existing autoregressive models suffer from the exposure bias problem due to mismatches between training and generation stages. |
| Approach: | They propose a Transformerbased Implicit Latent GAN which combines a transformer autoencoder and GAN in the latent space with a novel design and distribution matching based on the Kullback-Leibler divergence. |
| Outcome: | The proposed model improves local and global coherence and quality-diversity trade-off on three benchmark datasets. |
Copied to clipboard
| Challenge: | Using a sparse Mixture-of-Experts Transformer architecture, our model is highly efficient and efficient across languages. |
| Approach: | They propose a multilingual large language model trained on more than 160 languages using a sparse Mixture-of-Experts Transformer architecture. |
| Outcome: | The proposed model performs well on natural language generation and understanding tasks while avoiding the common pitfalls of multilinguality. |
Copied to clipboard
| Challenge: | Quantization is an effective technique to address heavy computation load and memory overhead during inference. |
| Approach: | They propose a low-bit quantization strategy to represent Transformer weights by an extremely low number of bits. |
| Outcome: | The proposed model achieves 11.8 smaller model size than baseline model, with less than -0.5 BLEU. |
Copied to clipboard
| Challenge: | Existing feed-forward neural networks have significant computational and parametric overhead. |
| Approach: | They propose a parameter-efficient Transformer architecture that utilizes multiple smaller FFNs to reduce parameters and computation while maintaining essential hidden dimensions. |
| Outcome: | The proposed architecture reduces computational and parameter overhead while maintaining essential hidden dimensions. |
Copied to clipboard
| Challenge: | Transformer is a powerful architecture that achieves superior performance on various sequence learning tasks, including neural machine translation, language understanding, and sequence prediction. |
| Approach: | They propose a new formulation of attention via the lens of the kernel which allows us to understand individual components of Transformer's attention. |
| Outcome: | The proposed model outperforms existing models on language understanding and sequence prediction tasks and is more efficient than existing models. |
Copied to clipboard
| Challenge: | Abstract Meaning Representation parsing is a sentence-to-graph prediction task . graph nodes are semantically based on one or more sentence tokens, so implicit alignments can be derived. |
| Approach: | They propose a transition-based system that decouples hard-attention over sentences with a target-side action pointer mechanism to decouple source tokens from node representations and address alignments. |
| Outcome: | The proposed system achieves the second best Smatch score on AMR 2.0 (81.8) it decouples source tokens from node representations and addresses alignments, but lacks expressiveness. |
Copied to clipboard
| Challenge: | Pre-trained Transformer models have proven their effectiveness in adapting to multiple NLP tasks and domains. |
| Approach: | They evaluated three categories of out-of-vocabulary words using three French domain-specific datasets on the legal, medical, and energetical domains to robustly analyze these categories. |
| Outcome: | The proposed models can create new representations for out-of-vocabulary words by adding external morpho-syntactic context rather than improving the semantic understanding of the words directly. |
Copied to clipboard
| Challenge: | Existing studies have shown that attention heads have a temporal induction property that allows them to learn and reproduce sequences of tokens. |
| Approach: | They analyze attention heads and transformer outputs to examine in-context temporal biases . they find that transformer output has a tendency toward in-constext serial recall . |
| Outcome: | The findings shed light on similarities and differences between LLMs and human memory and learning. |
Copied to clipboard
| Challenge: | Several phenomena where asymmetry arises have been identified as challenging problems for machine translation. |
| Approach: | They perform a fine-grained analysis of how an SMT system compares with two NMT systems when translating bare nouns into English. |
| Outcome: | The proposed model outperforms the SMT and BiLSTM models for 4 categories and the BiLST outperformed the SLT models for 3 categories. |
Copied to clipboard
| Challenge: | In neural machine translation, the general task of translating is to reduce the input sentence into smaller units (also known as statistical phrases), select an optimal translation for each unit, and place them in the correct order. |
| Approach: | They propose a novel architecture that relies on a feed-forward backbone and self-attention mechanism to encode sequential/positional information. |
| Outcome: | The proposed architecture improves on multiple datasets in French, Italian, and German and shows that it is more efficient than the current model. |
Copied to clipboard
| Challenge: | High-order numerical methods enhance performance in tasks like NLP but introduce a performance-efficiency trade-off due to increased computational overhead. |
| Approach: | They propose an iterative implicit Euler Transformer which simplifies high-order numerical methods by iterating implicit Eule. |
| Outcome: | The proposed method improves accuracy and reduces inference overhead by 55% while retaining 99.4% of the original task accuracy. |
Copied to clipboard
| Challenge: | Extensive experiments on ten WMT machine translation tasks show that the proposed model yields an average of 1.35x faster (with almost no decrease in BLEU) |
| Approach: | They propose a weighted residual network which reconstructs attention by reusing the features across layers. |
| Outcome: | The proposed model is 1.35x faster than the state-of-the-art inference model on translation tasks compared to AAN and SAN models with fewer parameter numbers . |
Copied to clipboard
| Challenge: | Neural Machine Translation (NMT) is a promising approach to machine translation . lack of parallel training data for Hindi-English is limiting . |
| Approach: | They propose to incorporate linguistic knowledge encoded by Hindi phenomena into a Transformer model to improve the translation performance. |
| Outcome: | The proposed model incorporates linguistic features to improve the translation performance. |
Copied to clipboard
| Challenge: | Existing approaches to improve online inference efficiency of the Transformer for instantaneous Grammatical Error Correction (GEC) are sequenceto-sequence (seq2sequ) and sequenceto sequence (saq2eq) |
| Approach: | They propose a novel approach to improve the online inference efficiency of the Transformer model for instantaneous Grammatical Error Correction (GEC) it aggressively decodes as many tokens as possible in parallel instead of always decoding only one token in each step to improve computational parallelism. |
| Outcome: | The proposed approach can achieve state-of-the-art results in English and Chinese benchmarks with 10x speedup over the Transformer-big model. |
Copied to clipboard
| Challenge: | a recent study of generation order for machine translation shows it does not affect output quality . Neural sequence models have been successfully applied to a broad range of tasks in recent years . |
| Approach: | They propose a soft order-reward framework that enables models to follow arbitrary oracle generation policies. |
| Outcome: | The proposed framework explores a wide variety of generation orders including uninformed orders, location-based orders, frequency-based or model-based orderings, and model-driven orders. |
Copied to clipboard
| Challenge: | Currently, the Transformer is the de facto architecture of choice for processing sequential data. |
| Approach: | They evaluate the Transformer architecture and its modifications in a shared experimental setting . they conjecture that performance improvements may strongly depend on implementation details . |
| Outcome: | The proposed improvements do not significantly improve performance, the authors find . the proposed improvements are either developed in the same codebase or are minor changes . |
Copied to clipboard
| Challenge: | Recent studies show that the attention heads in Transformer are not equal. |
| Approach: | They propose a masking method to mask attention heads in Transformer . they empirically validate the inequality and propose 'head mask' method to avoid bottleneck . |
| Outcome: | The proposed masking method improves translation performance on multiple languages . it can be used to remove a small subset of heads without affecting performance . |
Copied to clipboard
| Challenge: | Existing studies attributed zero-shot translation to domination of central language, e.g. English, but we supplement this viewpoint with the strict dependence of non-centered languages. |
| Approach: | They propose a language-specific modeling method that adapts to non-centered languages to counteract the instability of zero-shot translation. |
| Outcome: | The proposed method performs better than baselines in centered data conditions and can easily fit non-centered data. |
Copied to clipboard
| Challenge: | Variational Auto-Encoder (VAE) has been widely adopted in text generation due to its ability to learn flexible representations. |
| Approach: | They propose a Transformer-based recurrent VAE structure that imposes recurrence on segment-wise latent variables with arbitrarily separated text segments and constructs the posterior distribution with residual parameterization. |
| Outcome: | The proposed structure can deduce a non-zero lower bound of the KL term and enhance the entanglement of each segment and preceding latent variables, providing a theoretical guarantee of generation diversity. |
Copied to clipboard
| Challenge: | Adaptive Boundary-Token Fusion and a Morpheme-Aware Attention Bias are used to encode monosyllabic morphemes. |
| Approach: | They propose a morpheme-aware Transformer that augments a pretrained Vietnamese encoder with two lightweight inductive biases. |
| Outcome: | The proposed morpheme-aware Transformer outperforms strong baselines on Vietnamese POS, NER, and sentence-level classification benchmarks. |
Copied to clipboard
| Challenge: | Existing methods to compress Transformer are limited to sub-components, e.g., selfattention networks or embedding layer. |
| Approach: | They propose a Hybrid Tensor-Train decomposition which retains full rank and meanwhile reduces operations and parameters. |
| Outcome: | The proposed model outperforms light-weight SOTA methods on three translation tasks and achieves 7.1 points absolute improvement in BLEU and 1.27 X speedup on IWSLT’14 De-En task. |
Copied to clipboard
| Challenge: | Transformer-based models are stretched to enormous sizes, requiring increasingly larger training datasets and unsustainable amount of compute resources. |
| Approach: | They propose an alternative compatibility function for the Transformer-based attention mechanism that exploits an overlap in the learned representation of the traditional scaled dot-product attention mechanism. |
| Outcome: | The proposed model achieves 79.36 on the GLUE benchmark against 78.74 for the traditional implementation and reduces the number of trainable parameters by 6%. |
Copied to clipboard
| Challenge: | a large number of medical encounters need to be coded everyday due to long document sets and large label set. |
| Approach: | They propose a convolutional attention network for multi-label document classification problem . they use convolution-based encoders and convolution networks to aggregate information across documents . |
| Outcome: | The proposed model outperforms prior best model and multilingual Transformer model on a widely used dataset in the medical domain. |
Copied to clipboard
| Challenge: | Recent advances in machine translation have focused on a single pre-trained decoder . encoder-decoder architectures have received relatively little attention in NMT . |
| Approach: | They propose a method that leverages LLMs as MT encoders and pairs them with lightweight decoders to develop universal translation models. |
| Outcome: | The proposed method matches or surpasses baselines in terms of translation quality but achieves 75% reduction in memory footprint of the KV cache. |
Copied to clipboard
| Challenge: | Prior work has focused on training one network on multiple datasets to build a model that performs well on all of the training datasets and generalizes and transfers better to new datasets. |
| Approach: | They combine multiple reading comprehension datasets to build a multi-dataset question answering model with an ensemble of single-data set experts. |
| Outcome: | The proposed model outperforms baseline models in in-distribution accuracy and generalization and transfer performance. |
Copied to clipboard
| Challenge: | Existing neural models lack systematic compositionality in learning symbolic structures . existing models lack this ability in learning symbols, despite being able to understand complex structures. |
| Approach: | They propose to use auxiliary sequence prediction tasks to train a Transformer model to understand compositional symbolic structures of input data. |
| Outcome: | The proposed model improves on the SCAN compositionality challenge, with only 418 (5%) training instances, and achieves 97.8% accuracy on the MCD1 split. |
Copied to clipboard
| Challenge: | We train a 170Mparameter Backpack language model on OpenWebText, matching the loss of a 6Bparameter Transformer. |
| Approach: | They propose a neural architecture that learns multiple non-contextual sense vectors for each word in a vocabulary and represents a word as a context-dependent, non-negative linear combination of sense vector. |
| Outcome: | The proposed model outperforms a GPT-2's word embeddings on lexical similarity evaluations and can be used to perform controllable text generation and debiasing. |
Copied to clipboard
| Challenge: | a new method to learn which compressions to apply is based on syntactic rules for deleting spans . plausibility and salience are the two main criteria for determining which compression to apply . a recent study shows that the plausability model generally selects for grammatical and factual deletions compared to extractive methods . |
| Approach: | They propose to leave the decision about what to delete to two data-driven criteria . they show that plausibility and salience are the most important criteria if a span is deleted . |
| Outcome: | The proposed method achieves strong in-domain results on benchmark datasets and human evaluation shows that plausibility model generally selects for grammatical and factual deletions. |
Copied to clipboard
| Challenge: | Existing methods for entity prediction cannot predict when an event will occur . there are many facts not related to the query that can confuse the model . |
| Approach: | They propose a temporal knowledge Graph reasoning model based on Graph Hawkes Transformer . the model captures instantaneous structural and temporal evolution information . |
| Outcome: | The proposed model performs much better under long-term evolution scenarios. |
Copied to clipboard
| Challenge: | Abstractive opinion summarization framework outperforms competitors' summarizing frameworks . extractive approaches produce well-formed text, but selecting the most popular opinions is challenging . |
| Approach: | They propose an abstractive opinion summarization framework that trains a Transformer model to reconstruct reviews from extracted opinions. |
| Outcome: | The proposed framework outperforms baselines on Yelp and shows promising customization capabilities. |
Copied to clipboard
| Challenge: | Recent research has shown a strong fit between surprisal values from Transformers and reading times. |
| Approach: | They evaluate a Transformer model that uses a recency bias added to attention scores to improve the fit to human reading times. |
| Outcome: | The proposed model improves on a Transformer that includes a recency bias added to attention scores. |
Copied to clipboard
| Challenge: | Current Language Models (LMs) lack essential In-Context Learning capabilities, a domain where the Transformer excels. |
| Approach: | They propose a Linear Transformer with a kernel inspired by the Taylor expansion of exponential functions, augmented by convolutional networks. |
| Outcome: | The proposed model amplifies its In-Context Learning abilities on the Pile dataset. |
Copied to clipboard
| Challenge: | Contextualized embeddings vary by context, even for the same token . a recent study shows a trade-off between the norm and the variance of the embedded word . |
| Approach: | They show that contextualized embeddings vary by context, even for the same token . they focus on the norm of the mean embeddment and the variance of the embeddables . |
| Outcome: | The proposed method is efficient and efficient for embeddings in sentences. |
Copied to clipboard
| Challenge: | Empirical results show that certain components are more important than others . we propose a new training strategy that can improve Transformer models by distinguishing unimportant components . |
| Approach: | They propose a training strategy that distinguishes the unimportant components in training . they compare the impact of individual component (sub-layer) on model performance . |
| Outcome: | The proposed training strategy can improve translation performance by distinguishing unimportant components in training. |
Copied to clipboard
| Challenge: | Traditional document similarity measures do not consider in what aspects two documents are similar. |
| Approach: | They extend document similarity with aspect information by performing a pairwise document classification task. |
| Outcome: | The proposed approach is best performing on 172,073 research paper pairs from the ACL Anthology and CORD-19 corpus. |
Copied to clipboard
| Challenge: | Recent studies on AMR-to-text generation formalize the task as a sequence-tosequence learning problem . previous approaches only consider the relations between directly connected concepts while ignoring the rich structure in AMR graphs. |
| Approach: | They propose a structure-aware self-attention approach to model the relations between indirectly connected concepts in the seq2seq model. |
| Outcome: | The proposed approach outperforms the state-of-the-art on English AMR benchmarks . it significantly outperformed the state of the art on the benchmarks, with 29.66 and 31.82 BLEU scores . |
Copied to clipboard
| Challenge: | Neural machine translation models with tens and even more than a hundred blocks have shown effectiveness in image recognition. |
| Approach: | They propose a two-stage approach with three specially designed components to construct deeper NMT models. |
| Outcome: | The proposed approach improves on WMT14 EnglishGerman and EnglishFrench translation tasks. |
Copied to clipboard
| Challenge: | a quarter century ago, linguists assumed that language knowledge needed to be innate . but vector-space representations and machine learning algorithms are much more powerful than was thought . |
| Approach: | They trace the history of neural networks applied to natural language understanding tasks . they argue that Transformer is not a sequence model but an induced-structure model . |
| Outcome: | The proposed model is not a sequence model but an induced-structure model, the authors argue . they argue that the nature of language has had a profound impact on progress in machine learning . |
Copied to clipboard
| Challenge: | Recent work on why Transformer-based large language models make predictions has made their behavior opaque due to the complexity of the computations performed within each layer. |
| Approach: | They propose a linear decomposition of final hidden states from autoregressive language models based on each initial input token, which is exact for virtually all contemporary Transformer architectures. |
| Outcome: | The proposed method analyzes the influence of input tokens on model probabilities over a sequence of upcoming words with only one forward pass from the model. |
Copied to clipboard
| Challenge: | Uncertainty estimation (UE) of model predictions is crucial step for a variety of tasks such as active learning, misclassification detection, adversarial attack detection, etc. |
| Approach: | They propose to modify UE methods for Transformer models for misclassification detection in named entity recognition and text classification tasks to improve model expressiveness and computational performance. |
| Outcome: | The proposed methods outperform computationally intensive methods on misclassification detection tasks and are based on a large dataset of simulated datasets. |
Copied to clipboard
| Challenge: | In many NLP tasks, the input text can be seen as a sequence of related segments. |
| Approach: | They propose a layer-adjustable interactions framework that contextualizes token representations by attending to all other tokens at each layer, leading to quadratic increase in compute effort with the input length. |
| Outcome: | The proposed model reduces 30-50% of attention FLOPs while maintaining high accuracy. |
Copied to clipboard
| Challenge: | Residual networks are an Euler discretization of solutions to Ordinary Differential Equations (ODE). |
| Approach: | They propose a residual block of layers in Transformer that can be described as a higher-order solution to ODE. |
| Outcome: | The proposed architecture can gain large improvements over strong baselines at a slight cost in inference efficiency. |
Copied to clipboard
| Challenge: | Unsupervised machine translation models are limited by the run-time of autoregressive inference during back-translation and lack of synthetic data efficiency. |
| Approach: | They propose a two-for-one improvement to Transformer back-translation: Quick Back-Translation (QBT). QBT re-purposes the encoder as a generative model, and uses encoder-generated sequences to train the decoder. |
| Outcome: | Experiments on various WMT benchmarks show that QBT dramatically outperforms standard back-translation only method in terms of training efficiency for comparable translation qualities. |
Copied to clipboard
| Challenge: | Existing work on local explanation generation attempts to understand model dynamics on word-level or phraselevel by assigning importance scores on input features. |
| Approach: | They propose to interpret neural networks by linear decomposition by a Transformer model on a single input and a linear decomposing of the output to generate local explanations. |
| Outcome: | The proposed method achieves competitive performance in sentiment classification and machine translation, and fidelity of explanation. |
Copied to clipboard
| Challenge: | Attention is a key component of Transformers, which have achieved considerable success in natural language processing. |
| Approach: | They propose to integrate attention weights and the norm of transformed input vectors into a norm-based analysis that incorporates the norm. |
| Outcome: | The proposed analysis shows that attention weights alone determine the output of attention and that reasonable word alignment can be extracted from attention mechanisms of Transformers. |
Copied to clipboard
| Challenge: | et al., 2017) show that multi-head attention is important for neural machine translation. |
| Approach: | They evaluate the contribution made by individual attention heads to the overall performance of the Transformer model and analyze the roles played by them in the encoder. |
| Outcome: | The proposed pruning method removes the vast majority of heads without affecting performance. |
Copied to clipboard
| Challenge: | Existing MoE designs do not consider computational constraints (e.g., FLOPs, latency) Existing works in MoE consider homogeneous design where the same number of experts of the same size are placed uniformly throughout the network. |
| Approach: | They propose a framework for designing heterogeneous MoEs under computational constraints. |
| Outcome: | The proposed framework achieves 4x inference speedup and FLOPs reduction over manual models and within 1 BLEU point of MoE SwitchTransformer over benchmark datasets for NMT. |
Copied to clipboard
| Challenge: | Structured pruning is a feasible solution for end-side LLM deployment . however, achieving a high compression ratio for scaled-up LLMs remains a challenge . |
| Approach: | They propose a task-agnostic structured pruning approach coupled with a compact Transformer architecture to prune LLMs into an intra-module low-rank architecture. |
| Outcome: | The proposed approach reduces transitional activations inside multi-head attention (MHA) and multi-layer perceptron (MLP) modules while preserving inter-module activations sensitive to perturbations. |
Copied to clipboard
| Challenge: | Existing methods to enhance length extrapolation of large language models have been developed, but a systematic survey is lacking. |
| Approach: | They propose to examine the effects of positional encoding on length extrapolation. |
| Outcome: | The proposed methods improve the extrapolation of large language models, but they are still lacking a systematic survey. |
Copied to clipboard
| Challenge: | Existing Mixture-of-Expert (MoE) models allow us to scale up model sizes while keeping the amount of compute time fixed. |
| Approach: | They propose to use a router to route inputs to experts in a layer to scale up model sizes while keeping the amount of compute time fixed. |
| Outcome: | The proposed model scales up with the help of a router that routes input tokens to experts in a layer and shows that it is more efficient than a non-trainable router. |
Copied to clipboard
| Challenge: | Recent studies of the computational power of recurrent neural networks reveal a hierarchy of RNN architectures, given finite-precision assumptions. |
| Approach: | They propose to use auto-regressive Transformers with linearised attention to build RNNs . they show that many well-known results for the standard Transformer directly transfer to LTs - a new approach is proposed . |
| Outcome: | The proposed extensions overcome limitations of the LT and self-referential weight matrices. |
Copied to clipboard
| Challenge: | Recent work on pretrained language models has led to significant improvements in a range of NLP tasks. |
| Approach: | They propose a coherence model which interprets sentences incrementally to capture lexical relations between them. |
| Outcome: | The proposed model interprets sentences incrementally to capture lexical relations between them. |
Copied to clipboard
| Challenge: | Tables are ubiquitous on the web, and are rich in information. |
| Approach: | They propose a sparse-attention Transformer architecture for modeling documents that contain large tables. |
| Outcome: | The proposed architecture scales linearly with respect to speed and memory, and can handle documents containing more than 8000 tokens with current accelerators. |
Copied to clipboard
| Challenge: | Existing approaches to character-level language modeling have suffered from high learning complexity caused by inherently long character sequences. |
| Approach: | They propose a method that efficiently reduces the computational cost and parameter size of Transformer by splitting feature space into multiple groups, factorizing the calculation paths, and reducing computations for the group interaction. |
| Outcome: | The proposed model reduces the computational cost and parameter size of Transformer on two benchmark tasks, enwik8 and text8, and it performs well. |
Copied to clipboard
| Challenge: | a recent study has shown that inference steps are not equally challenging, with some being "harder" and others "easier." |
| Approach: | They propose a modified Transformer forward pass that selectively applies additional computation when the model encounters uncertainty during token generation. |
| Outcome: | The proposed method achieves performance gains while maintaining inference times twice faster than beam search. |
Copied to clipboard
| Challenge: | Existing approaches to exploit sentential context for machine translation are not well studied. |
| Approach: | They propose a shallow sentential context that exploits top encoder layer, and a deep sentential one that aggregates sentential representations from all internal layers. |
| Outcome: | The proposed model outperforms the strong Transformer model on the English-German and English-French benchmarks. |
Copied to clipboard
| Challenge: | Recent work on text sequence matching tasks uses task specific supervised datasets, which are always limited to the amount due to the cost of annotation. |
| Approach: | They propose an aggregation method to combine Bidirectional Encoder Representations from Transformer (BERT) with a MatchLSTM layer for Sequence Matching. |
| Outcome: | The proposed model improves on two publicly available datasets, WikiQA and SNLI. |
Copied to clipboard
| Challenge: | Recent results suggest that positional encodings are not necessary when training decoder-only Transformer language models. |
| Approach: | They propose a causal attention mechanism that allows Transformers to store positional information without positional encodings. |
| Outcome: | The proposed model can reconstruct the positions of tokens without positional encodings. |
Copied to clipboard
| Challenge: | Existing approaches to improve accuracy of neural networks are slow due to computational complexity. |
| Approach: | They propose a vector-vector-matrix architecture which greatly reduces latency at inference time for NLP applications by a factor of four. |
| Outcome: | The proposed framework reduces the latency of sequence-to-sequence and Transformer models used for NMT by a factor of four. |
Copied to clipboard
| Challenge: | Inductive biases play a critical role in NLP, especially in learning from limited data and generalizing systematically outside of the training distribution. |
| Approach: | They propose to strengthen the structural inductive bias of a Transformer by intermediate pre-training to perform syntactic transformations of dependency trees given a description of the transformation. |
| Outcome: | The proposed model can perform syntactic transformations and generalize semantic parsing with attention heads that keep track of which syntaktic transformation needs to be applied to which token. |
Copied to clipboard
| Challenge: | Current Transformer-based sequence-to-sequence architectures can suffer from overfitting during training. |
| Approach: | They propose to use Transformer-based sequence-to-sequence architectures to overcome overfitting problems when generating very long sequences. |
| Outcome: | The proposed model performs worse on very long sequences than previous approaches on string editing and translation tasks when faced with sequences of length diverging from the length distribution in training data. |
Copied to clipboard
| Challenge: | Current captioning models are trained to generate captions that only contain common object names, thus falling short on an important “informativeness” dimension. |
| Approach: | They propose a mechanism for integrating image information and fine-grained labels into a caption that describes the image in a fluent and informative manner. |
| Outcome: | The proposed model integrates image information with fine-grained labels to produce fluent captions . it can control the appearance of these labels in the output, resulting in fluent and informative captions. |
Copied to clipboard
| Challenge: | Deploying Transformer networks on resource-constrained edge devices is challenging. |
| Approach: | They propose a low-rank factorization initialized by SVD-based weight transfer and parameter sharing to compress and accelerate Transformer networks. |
| Outcome: | The proposed method achieves similar performance to the baseline Transformer with 3.8 times and 1.8 times fewer parameters and achieves 2.3 times speedup and 1.5 times speed up respectively. |
Copied to clipboard
| Challenge: | Existing question answering systems mainly focus on text data, but few Korean datasets exist . a dataset for table question answering is written in English, but it lacks Korean-specific datasets . |
| Approach: | They construct Korean-specific datasets for table question answering using crowd-sourced workers . they then fine-tune the model with these datasets and report the evaluation results . |
| Outcome: | The proposed model is based on Korean datasets and is publicly available . the model is evaluated against other datasets from Korean question answering systems . |
Copied to clipboard
| Challenge: | Existing linear attention models use a decay factor based positional encoding (PE), but the decay factor is manually designed and non-trainable, limiting further optimization. |
| Approach: | They propose a PE-based positional encoding that disentangles decay factor into two parts to achieve further optimization and stable training. |
| Outcome: | The proposed model achieves stable training of decay factor and improves inference efficiency in normal context and extrapolation scenarios. |
Copied to clipboard
| Challenge: | Existing neural nets fail to generalize systematically due to superficial differences in training data. |
| Approach: | They propose a new diagnostic dataset based on compositions of unary symbolic functions that tests systematicity of NNs. |
| Outcome: | The proposed dataset shows that recent CTL-solving Transformer variants fail on CTL++. |
Copied to clipboard
| Challenge: | Prepositional phrase attachment ambiguity is structural in nature, while garden path constructions are incremental in nature. |
| Approach: | They pretrain and evaluate an unsupervised Transformer model that induces tree representations internally and compare it to a pretrained supervised BiLSTM model. |
| Outcome: | The Tree Transformer model induces tree representations internally, but its parsing ability is inferior to the supervised BiLSTM model, and it is not as sensitive to lexical cues as other large LSTM models. |
Copied to clipboard
| Challenge: | Deep attention models have advanced the modelling of sequential data across many domains. |
| Approach: | They propose to use a Transformer augmented with a long-range memory to model sequential data across many domains. |
| Outcome: | The Transformer-XL has a long-range memory at every layer of the network, rendering its state thousands of times larger than RNN predecessors. |
Copied to clipboard
| Challenge: | Multilingual neural machine translation models can translate multiple language pairs in a single model but lacks ability to capture language-specific features. |
| Approach: | They propose a token-level feature mixing method that captures different features and dynamically determines feature sharing across languages. |
| Outcome: | The proposed method outperforms baselines and can be extended to zero-shot translation. |
Copied to clipboard
| Challenge: | Existing zero-shot singing voice synthesis models depend on phoneme and note boundary annotations, limiting their robustness and producing poor transitions between phonemes and notes. |
| Approach: | They propose a multi-task multilingual zero-shot SVS model with style transfer and style control based on various prompts. |
| Outcome: | Experimental results show that TCSinger 2 outperforms baseline models in subjective and objective metrics across multiple related tasks. |
Copied to clipboard
| Challenge: | Despite the success of speech recognition, how to encode the speech features effectively remains an open problem. |
| Approach: | They propose a Progressive Down-Sampling technique which compresses acoustic features into coarser-grained units containing more complete semantic information, like text-level representation. |
| Outcome: | The proposed method yields comparable or better results on the speech recognition task and inference speedups ranging from 1.20x to 1.47x. |
Copied to clipboard
| Challenge: | Existing vision-language-action models rely on causal attention for processing sequences composed of interleaved segments from different modalities. |
| Approach: | They propose a Transformer architecture featuring trajectory attention and learnable action queries that efficiently process segmented multimodal trajectories and predict actions for imitation learning. |
| Outcome: | The proposed architecture performs better on three large-scale robot manipulation benchmarks than previous models. |
Copied to clipboard
| Challenge: | Large-scale generative Pre-trained Language Models (PLMs) are limited in their deployment in real-world applications. |
| Approach: | They propose to prune the feed-forward networks of generative pre-trained language models to smaller widths without designing extra operators. |
| Outcome: | The proposed method achieves 1.51x/6.96x inference speedup on GPU/CPU with 67% size reduction. |
Copied to clipboard
| Challenge: | Existing frameworks for learning informative latent variables are limited by limitations . existing models rely on strong assumptions on distribution of latent code . |
| Approach: | They propose to apply a variational neural machine translation framework to a Transformer . they propose to introduce a more flexible approximate posterior based on normalizing flows . |
| Outcome: | The proposed framework outperforms baseline models under in-domain and out-of-domain conditions. |
Copied to clipboard
| Challenge: | Scholars call for more personalized mechanisms of content moderation to account for multifaceted differences. |
| Approach: | They propose to annotate a news forum's user comments with a German dialect and identify their spans as vulgar language or offensive statements. |
| Outcome: | The proposed model interpretability improves on fine-tuned Transformer models and large language models in a zero-shot fashion. |
Copied to clipboard
| Challenge: | In sentiment classification, there are some good features that are indicative of class labels, but there are also many common features that do not discriminate for classification. |
| Approach: | They propose to project existing features into the orthogonal space of the common features and make them more discriminative for classification. |
| Outcome: | The proposed method improves CNN, RNN, Transformer, and Bert based text classification and obtains markedly better results. |
Copied to clipboard
| Challenge: | Adapting a model to target speakers requires a lot of compute and may cause catastrophic forgetting to the existing speakers. |
| Approach: | They propose a unified speaker adaptation approach consisting of feature adaptation and model adaptation. |
| Outcome: | The proposed model outperforms baseline models with 20.58% relative WER reduction and surpasses finetuning method by 2.54% on target speaker adaptation. |
Copied to clipboard
| Challenge: | Large multilingual models fail to successfully transfer to low-resource languages for zero-shot cross-lingual transfer . sliced fine-tuning for named entity recognition (SLICER) forces stronger token contextualization in the Transformer. |
| Approach: | They propose a simple yet highly effective approach for improving zero-shot cross-lingual transfer for named entity recognition to low-resource languages. |
| Outcome: | The proposed approach reduces decontextualization of token representations and classifiers . it yields consistent transfer gains for low-resource languages, the authors show . |
Copied to clipboard
| Challenge: | Extensive experiments show EdgeFormer can effectively outperform previous parameter-efficient Transformer baselines and achieve competitive results under both the computation and memory constraints. |
| Approach: | They propose a parameter-efficient Transformer for on-device seq2seq generation that uses two novel principles for cost-effective parameterization. |
| Outcome: | Extensive experiments show that EdgeFormer outperforms the previous parameter-efficient Transformers and achieves competitive results under both the computation and memory constraints. |
Copied to clipboard
| Challenge: | telemedicine is a medical practice that provides patient care remotely using video conferencing tools. |
| Approach: | They build large-scale medical dialogue datasets to facilitate research . they pretrain several models on the Chinese MedDialog dataset and compare their performance . |
| Outcome: | The proposed datasets show that models trained on MedDialog can generate doctor-like medical dialogues. |
Copied to clipboard
| Challenge: | Recent studies have shown that powerful Transformer architectures produce dull high-frequency phrases, severely hurting the diversity and novelty of generated text. |
| Approach: | They propose a method to control the sharpness of the attention distribution by python code and use it to learn a Bayesian approximation of posterior attention. |
| Outcome: | The proposed method improves diversity and novelty while maintaining comparable quality on conditional and unconditional generation tasks. |
Copied to clipboard
| Challenge: | Existing variational inference models ignore their latent variables, a phenomenon called posterior collapse. |
| Approach: | They propose a new loss function for conditional variational autoencoders that counteracts posterior collapse by using a modified evidence lower bound objective and a factorized decoder. |
| Outcome: | The proposed model yields improved translation quality compared to existing models on WMT RoEn and DeEn. |
Copied to clipboard
| Challenge: | Context gates are effective to control the contributions from the source and target contexts in the recurrent neural network (RNN) based neural machine translation. |
| Approach: | They propose a method to identify source and target contexts and introduce a gate mechanism to control the contributions from source and targets in the advanced Transformer architecture. |
| Outcome: | The proposed model achieves an averaged gain of 1.0 BLEU score over a strong transformer baseline. |
Copied to clipboard
| Challenge: | BERT models are often taken off-the-shelf and fine-tuned on a downstream task. |
| Approach: | They propose an extra stage of self-supervised task-adaptive pre-training to perform a task on a number of Croatian-supporting Transformer models. |
| Outcome: | The proposed approach improves performance across multilingual models but not in Croatian-dominant models. |
Copied to clipboard
| Challenge: | Multilingual people code-mix using English phonetic typing and insertion of anglicisms in their native language. |
| Approach: | They propose to use minority positive sampling to selectively induce more sample to achieve better performance. |
| Outcome: | The proposed model performs better than other models, but switching points are the main challenge . |
Copied to clipboard
| Challenge: | Quantization has proven to be effective after pre-training and during fine-tuning, but its effects on pre-trainer performance have remained unexplored. |
| Approach: | They propose a linear quantization strategy to be applied during the pre-training of Transformers to improve model efficiency and stability. |
| Outcome: | The proposed method improves model efficiency, stability, and performance while maintaining language modeling ability. |
Copied to clipboard
| Challenge: | Pre-trained language models (LMs) have shown effectiveness in literature understanding tasks, especially when tuned via contrastive learning. |
| Approach: | They propose a multi-task contrastive learning framework that enables common knowledge sharing across different scientific literature understanding tasks while preventing task-specific skills from interfering with each other. |
| Outcome: | The proposed framework outperforms state-of-the-art pre-trained language models on a comprehensive dataset. |
Copied to clipboard
| Challenge: | Existing studies on the scaling properties of model architectures have not explored the impact of inductive biases on scaling behaviour. |
| Approach: | They conduct extensive experiments to understand scaling behaviour of ten different model architectures. |
| Outcome: | The results show that the best performing model can fluctuate at different scales. |
Copied to clipboard
| Challenge: | Current evaluation of neural machine translation systems is limited by one best hypothesis and search errors brought by heuristic decoding algorithms. |
| Approach: | They propose a new evaluation protocol which defines model errors with model’s ranking capability over hypothesis space and Monte Carlo sampling evaluation to tackle the problem of exponentially large space. |
| Outcome: | The proposed evaluation protocol is consistent with what is currently used in the field and is consistent to what is being proposed. |
Copied to clipboard
| Challenge: | Existing Transformers that scale to long sequences are not compatible with relative position encoding. |
| Approach: | They propose a Performer-based model with relative position encoding that scales linearly on long sequences. |
| Outcome: | The proposed model outperforms performer on long sequences with no computational overhead and outperformed vanilla Transformer on most of the tasks. |
Copied to clipboard
| Challenge: | High-performance methods for parameter-efficient fine-tuning (PEFT) typically work with Attention blocks and overlook dense MLP blocks, which contain about half of the model parameters. |
| Approach: | They propose a selective PEFT method that performs well on MLP blocks by converting layer gradients into a sparse structure and reducing the number of updated parameters. |
| Outcome: | The proposed method outperforms LoRA and MeProp, robust state-of-the-art PEFT approaches. |
Copied to clipboard
| Challenge: | In multivariate long-term time series forecasting, it is widely believed that the effectiveness of self-attention arises from its attention matrix. |
| Approach: | They propose a multi-branch MLP that isolates the ‘multi-brain mapping with element-wise operation’ structure from the Transformer and shows that it achieves competitive performance. |
| Outcome: | The proposed model outperforms three classic and three latest Transformer models and shows that it achieves competitive performance. |
Copied to clipboard
| Challenge: | Existing supervised neural methods are underexplored for coreference resolution, especially in incremental clustering. |
| Approach: | They propose a dual-threshold incremental clustering approach based on a lightweight Transformer. |
| Outcome: | Experiments on common benchmarks show that MEIC-DT achieves highly competitive coreference performance under stringent memory constraints. |
Copied to clipboard
| Challenge: | Large language models struggle with complex reasoning tasks, such as mathematical problem-solving. |
| Approach: | They constructed a symbolic multi-step reasoning task to investigate the information propagation mechanisms in Transformer models when solving the task through direct answering and Chain-of-Thought (CoT) reasoning. |
| Outcome: | The proposed algorithm improves on 7 multi-step reasoning datasets, while introducing only 132 trainable parameters. |
Copied to clipboard
| Challenge: | Existing methods to use table pre-training to boost tabular prediction performance remain open . a bachelor's degree earns less than 50K, and a generative LM can be used to unify tasks via one LM. |
| Approach: | They propose a method that leverages table pre-training to empower tabular prediction models. |
| Outcome: | The proposed method outperforms baseline models on 12 datasets and can be easily combined with various backbone models. |
Copied to clipboard
| Challenge: | recurrent neural networks struggle to match the performance of Transformers due to limitations in parallelization and scalability. |
| Approach: | They propose a model architecture that combines the efficient parallelizable training of transformers with the efficient inference of RNNs. |
| Outcome: | The proposed model performs on par with similarly sized RNNs, suggesting future work can leverage this architecture to create more efficient models. |
Copied to clipboard
| Challenge: | Existing approaches to feature attributions rely on attention weights and attention weightings. |
| Approach: | They propose a feature attribution method that replaces attention weights with the generalized Information Tensor to enhance the performance of Transformer-based models. |
| Outcome: | The proposed method outperforms state-of-the-art feature attribution methods on sequence classification tasks and provides a more reliable interpretation of Transformer model outputs. |
Copied to clipboard
| Challenge: | Existing models to reduce computation complexity are limited in some areas . a new structure to reduce the computation complexity is proposed to accelerate Transformers . |
| Approach: | They propose a new feed forward network structure which splits matrix space to smaller space to reduce computation complexity. |
| Outcome: | The proposed model can achieve a faster speed and better accuracy on the long-range arena benchmark. |
Copied to clipboard
| Challenge: | Merger Agreement Understanding Dataset (MAUD) is an expert-annotated reading comprehension dataset based on the American Bar Association's 2021 Public Target Deal Points Study. |
| Approach: | They propose a Merger Agreement Understanding Dataset with over 39,000 examples and over 47,000 annotations. |
| Outcome: | The Merger Agreement Understanding Dataset (MAUD) is an expert-annotated reading comprehension dataset based on the American Bar Association's 2021 Public Target Deal Points Study. |
Copied to clipboard
| Challenge: | Existing medical conversation speech corpora for Burmese are limited, despite advances in ASR. |
| Approach: | They propose to use a manually curated medical conversation speech corpus for Burmese to examine the performance of ASR models. |
| Outcome: | The proposed model outperforms the Transformer model and the Recurrent Neural Network (RNN) models. |
Copied to clipboard
| Challenge: | Modern Natural Language Processing models have a huge capacity, but this makes it difficult to employ. |
| Approach: | They propose a method to quantize at least 95% of Transformer weights without access to task-specific data so the drop in performance does not exceed 0.02%. |
| Outcome: | The proposed method quantizes 95% of Transformer weights and corresponding activations to INT8 without access to task-specific data so the drop in performance does not exceed 0.02%. |
Copied to clipboard
| Challenge: | Inflection is a process of word formation in which a base word form (lemma) is modified to express grammatical categories. |
| Approach: | They develop a retrograde model and two sequence-to-sequence models based on LSTM and Transformer. |
| Outcome: | The proposed systems outperform the existing systems on 9 out of 16 languages in the OOV evaluation. |
Copied to clipboard
| Challenge: | Empirical results show that MoMs consistently outperform vanilla transformers . |
| Approach: | They propose an architecture that allows for a mixture-of-modules computation that uses a finite set of modules defined by multi-head attention and feed-forward networks. |
| Outcome: | The proposed architecture outperforms vanilla Transformers and their variants in multiple ways. |
Copied to clipboard
| Challenge: | Existing methods for token-level KV optimization and grouping of tokens are inefficient and strain compute and storage resources. |
| Approach: | They propose a mixture-of-expert approach that dynamically optimizes token-wise computation and memory allocation by a token-based expert-choice routing mechanism guided by learned importance scores. |
| Outcome: | The proposed approach retains all tokens while adaptively routing them to specialized experts with varying KV group sizes, balancing granularity and efficiency. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) exhibit strong performance on multilingual tasks, yet the process of constructing predictions in the target language remains under-explored. |
| Approach: | They propose a novel interpretability method focusing on the Feed-Forward Network (FFN) layers of Large Language Models. |
| Outcome: | The proposed interpretability method is based on the Feed-Forward Network (FFN) layer of Large Language Models. |
Copied to clipboard
| Challenge: | Existing methods to train large language models that require a non-uniform model norm are not effective. |
| Approach: | They propose a technique that allows for uniformity of the norm of the model parameters . they propose 'weight scaling as reparameterization' to adjust the norm to the parameter . |
| Outcome: | The proposed technique outperforms existing methods and stabilizes training with the transformer decoders. |
Copied to clipboard
| Challenge: | Empirical evaluations across various prominent LLMs and benchmarks show that key-favored allocations retain up to 98.3% accuracy compared to uniform allocations (e.g., 4-bit keys, 2-bit values). |
| Approach: | They propose two theorems that anchor mixed-precision KV quantization in the intrinsic geometry of Transformer models. |
| Outcome: | Empirical evaluations show that key-favored allocations retain up to 98.3% accuracy while conserving memory. |
Copied to clipboard
| Challenge: | Inductive biases are inherent in every machine learning system, argues a new study . m-local entropy measures how well symbols disambiguate the next symbol . |
| Approach: | They propose a framework that captures local uncertainty of a language by quantifying how effectively preceding symbols disambiguate the next symbol. |
| Outcome: | The proposed framework captures the local uncertainty of a language by quantifying how effectively symbols disambiguate the next symbol. |
Copied to clipboard
| Challenge: | Existing approaches to interpretable representation learning rely on masks that weight the significance of input features, but the origin of these masks is uncertain. |
| Approach: | They propose a causal framework that directly learns identifiable representations from attention weights rather than relying on importance masks. |
| Outcome: | The proposed framework learns identifiable and explainable representations from attention weights, rather than masks, and guarantees faithfulness on real-world datasets. |
Copied to clipboard
| Challenge: | Large language models face inherent performance bottlenecks under parameter constraints . challenging tokens induce abrupt gradient spikes across layers, exposing stress points . |
| Approach: | They propose an inner thinking transformer that reimagines layer computations as implicit thinking steps. |
| Outcome: | Empirical results show that ITT outperforms Transformer/Loop variants in 11 benchmarks. |
Copied to clipboard
| Challenge: | Existing decoder-only transformers fail to preserve initial token-level information in deeper layers. |
| Approach: | They propose a new architecture that incorporates value residual connections in addition to hidden state residuals. |
| Outcome: | The proposed architecture reduces KV cache size by nearly half with only a small performance penalty and can be integrated with other KV-efficient methods. |
Copied to clipboard
| Challenge: | Existing studies show that stacking causal self-attention layers alone induces a positional bias in attention scores toward earlier tokens, but this differs from the bias toward later tokens observed in Transformer decoders, known as recency bias. |
| Approach: | They propose to stack causal self-attention layers and layer norm to induce recency bias in Transformer decoders by analyzing the interaction between causal self and other architectural components. |
| Outcome: | The proposed method provides new theoretical insights into how positional information interacts with architectural components and suggests improvements in positional encoding strategies. |
Copied to clipboard
| Challenge: | Recent work focuses on syntactic tree structures of languages, in particular constituency tree structures. |
| Approach: | They propose a Graph-Infused Layers Transformer Language Model which leverages dependency graphs to augment Transformer language models. |
| Outcome: | The proposed model achieves better syntactic generalization while maintaining competitive perplexity compared with baseline models. |
Copied to clipboard
| Challenge: | Existing methods for inference in centralized cloud pose privacy risks due to sensitive data. |
| Approach: | They propose a latency-aware framework for distributed Transformer inference in resource-constrained edge networks. |
| Outcome: | The proposed framework achieves 2.01 times inference acceleration over state-of-the-art baselines with leq1.06% accuracy loss, maintaining robustness under varying edge conditions. |