SpanBERT: Improving Pre-training by Representing and Predicting Spans (2020.tacl-1)
Copied to clipboard
| Challenge: | Pre-training methods like BERT mask individual words or subword units, but many tasks involve reasoning about relationships between two or more spans of text. |
| Approach: | They propose a pre-training method that masks contiguous random spans instead of random tokens to train the span boundary representations to predict the entire content of the masked span. |
| Outcome: | The proposed method outperforms BERT and its better-tuned baselines on span selection tasks and on coreference resolution tasks. |
Similar Papers
PairSpanBERT: An Enhanced Language Model for Bridging Resolution (2023.acl-long)
Copied to clipboard
| Challenge: | bridging resolution is crucial for machine comprehension of discourse entities for various downstream applications. |
| Approach: | They propose a SpanBERT-based pre-trained model specialized for bridging resolution. |
| Outcome: | The proposed model achieves the best results on three evaluation datasets for bridging resolution despite the noise inherent in the automatically generated data . |
Span Selection Pre-training for Question Answering (2020.acl-main)
Copied to clipboard
Michael Glass, Alfio Gliozzo, Rishav Chakravarti, Anthony Ferritto, Lin Pan, G P Shrivatsa Bhargav, Dinesh Garg, Avi Sil
| Challenge: | Pre-trained BERTs provide large gains across many language understanding tasks, achieving a new state-of-the-art (SOTA). |
| Approach: | They propose a new pre-training task inspired by reading comprehension to better align the pre- training from memorization to understanding. |
| Outcome: | The proposed model outperforms BERT-BASE and BERT LARGE on a new dataset and improves answer prediction F1 by 4 points and supporting fact prediction F1. |
Coreference Resolution without Span Representations (2021.acl-short)
Copied to clipboard
| Challenge: | Pretraining has reduced many complex task-specific NLP models to simple lightweight layers. |
| Approach: | They propose a lightweight end-to-end coreference model that removes the dependency on span representations, handcrafted features, pruning heuristics, and more. |
| Outcome: | The proposed model performs competitively with the current standard model, while being simpler and more efficient. |
LinkBERT: Pretraining Language Models with Document Links (2022.acl-long)
Copied to clipboard
| Challenge: | Existing language model pretraining methods do not capture dependencies or knowledge that span across documents. |
| Approach: | They propose a language model pretraining method that leverages links between documents . they use masked language modeling and document relation prediction to model LMs . |
| Outcome: | The proposed method outperforms existing methods on downstream tasks across two domains. |
Improving Span Representation by Efficient Span-Level Attention (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for generating high-quality span representations are limited by subset of tokens . span-span interactions should play an important role in span encoding, authors argue . |
| Approach: | They propose to introduce span-span interactions and more comprehensive span-token interactions to improve span representations. |
| Outcome: | The proposed model outperforms baseline models on span-related tasks and shows superior performance. |
Span Fine-tuning for Pre-trained Language Models (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to fine-tune pre-trained language models are time-consuming and lack flexibility. |
| Approach: | They propose a span fine-tuning method which allows for a more efficient and efficient way of incorporating span-level information into pre-training. |
| Outcome: | Experiments on GLUE benchmark show that the proposed method significantly enhances the PrLM and offers more flexibility in an efficient way. |
TaCL: Improving BERT Pre-training with Token-aware Contrastive Learning (2022.findings-naacl)
Copied to clipboard
| Challenge: | Existing pre-trained MLMs produce an anisotropic distribution of token representations . this is not ideal for tasks that require discriminative semantic meanings of distinct tokens - a problem that exists in pre-training models . |
| Approach: | They propose a continual pre-training approach that encourages BERT to learn an isotropic distribution of token representations. |
| Outcome: | The proposed approach improves on a wide range of English and Chinese benchmarks. |
A Span Selection Model for Semantic Role Labeling (D18-1)
Copied to clipboard
| Challenge: | Existing models for semantic role labeling use BIO tags to predict argument spans . but performance of these approaches is weak . |
| Approach: | They propose a span-based model that takes into account all possible argument spans and scores them for each label. |
| Outcome: | The proposed model achieves state-of-the-art results on the CoNLL-2005 and 2012 datasets. |
LuxemBERT: Simple and Practical Data Augmentation in Language Model Pre-Training for Luxembourgish (2022.lrec-1)
Copied to clipboard
Cedric Lothritz, Bertrand Lebichot, Kevin Allix, Lisa Veiber, Tegawende Bissyande, Jacques Klein, Andrey Boytsov, Clément Lefebvre, Anne Goujon
| Challenge: | Pre-trained Language Models such as BERT are ubiquitous in NLP but are scarce for low-resource languages such as Luxembourgish. |
| Approach: | They propose a BERT model for Luxembourgish language that they use to augment pre-training datasets by partially translating text data from a closely related language. |
| Outcome: | The proposed model outperforms the baseline model and the mBERT model in Luxembourgish. |
NarrowBERT: Accelerating Masked Language Model Pretraining and Inference (2023.acl-short)
Copied to clipboard
| Challenge: | Large-scale language model pretraining is expensive as the models and pretraining corpora have become larger over time. |
| Approach: | They propose a modified transformer encoder that increases throughput for masked language model pretraining by more than 2x. |
| Outcome: | The proposed model increases throughput on IMDB and Amazon reviews classification and CoNLL NER tasks by 3.5x with minimal performance degradation. |