Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing
Copied to clipboard
| Challenge: | Non-autoregressive neural machine translation models suffer from the multi-modality problem . aligNART leverages full alignment information to explicitly reduce the modality of the target distribution . |
| Approach: | They propose an alignment decomposition method which explicitly reduces the modality of the target distribution. |
| Outcome: | The proposed model outperforms previous models that focus on modality reduction on two translation tasks. |
Copied to clipboard
| Challenge: | Existing work on improving cross-lingual transferability of NMT model is under-explored. |
| Approach: | They propose a model that leverages a multilingual pretrained encoder to improve cross-lingual transferability. |
| Outcome: | The proposed model outperforms mBART and m2m-100 on a zero-shot cross-lingual transfer task. |
Copied to clipboard
| Challenge: | Existing methods for pretraining cross-lingual models are limited in their size due to the limited amount of parallel corpora. |
| Approach: | They propose a method that encourages the model to align multiple languages with monolingual corpora to overcome the constraint of the parallel corpus size. |
| Outcome: | The proposed method outperforms existing cross-lingual models and delivers new state-of-the-art results in various cross-linguistic downstream tasks. |
Copied to clipboard
| Challenge: | Existing approaches to simultaneous translation are limited by monotonic constraint . a novel architecture for simultaneous translation is proposed . |
| Approach: | They propose a cross attention-augmented transducer for simultaneous translation that optimizes both policies and translation models by expanding target sequences with blank symbols. |
| Outcome: | The proposed architecture achieves better latency-quality trade-offs than state-of-the-art approaches. |
Copied to clipboard
| Challenge: | Schema translation is not well studied in the community because of morphological difference and context difference between plain text and tabular data. |
| Approach: | They propose a schema translation model augmented with schema context . they model a target header and its context as a directed graph to represent their entities . |
| Outcome: | The proposed model outperforms state-of-the-art models on schema translation . it uses a graph to represent entity types and relations, and a relational-aware transformer . |
Copied to clipboard
| Challenge: | Neural Chat Translation (NCT) models that use dialogue characteristics of chat are often incoherent and speakerirrelevant. |
| Approach: | They propose to introduce the modeling of dialogue characteristics into the NCT model by capturing the inherent dialogue characteristics. |
| Outcome: | The proposed model can translate conversational text between speakers of different languages. |
Copied to clipboard
| Challenge: | Existing methods for low-resource dialogue summarization neglect the difference between dialogues and conventional articles. |
| Approach: | They propose a multi-source pretraining paradigm to leverage external summary data . they exploit large-scale in-domain non-summary data to separate dialogue encoder and summary decoder . |
| Outcome: | The proposed model can be used to better leverage external summary data. |
Copied to clipboard
| Challenge: | Experimental results show that our proposed framework generates fluent and factually consistent summaries under various planning controls using both objective metrics and human evaluations. |
| Approach: | They propose a controllable neural generation framework that can guide dialogue summarization with personal named entity planning. |
| Outcome: | The proposed framework generates fluent and factually consistent summaries under various planning controls using objective metrics and human evaluations. |
Copied to clipboard
| Challenge: | Recent studies have shown that around 30% of the summaries generated by abstractive summarization models contain factual errors. |
| Approach: | They propose a fine-grained two-stage Fact Consistency assessment framework for summarization models that uses fine-grain consistency reasoning to find subtle clues to identify whether a model-generated summary is consistent with the original document. |
| Outcome: | The proposed framework improves on the state-of-the-art models and distinguishes detailed differences better. |
Copied to clipboard
| Challenge: | Existing summarization methods define relevance based on textual information alone without incorporating insights about a particular decision. |
| Approach: | They propose a method that summarizes relevant information for a decision using full text . they then build a model that makes the decision based on the full text while accounting for textual non-redundancy. |
| Outcome: | The proposed method outperforms text-only summarization methods and model-based explanation methods in decision faithfulness and representativeness. |
Copied to clipboard
| Challenge: | Existing methods for extractive text summarization do not consider multiple types of inter-sentential relationships, nor model intra-sententential relationships. |
| Approach: | They propose a novel method to combine different types of relationships among sentences and words to model sentence embedding. |
| Outcome: | The proposed model is compared with existing methods on CNN/DailyMail benchmark dataset to demonstrate its effectiveness. |
Copied to clipboard
| Challenge: | Previous work has used task-agnostic pretraining methods like masked language models or corrupted span prediction to improve performance on downstream tasks. |
| Approach: | They propose to use a task-agnostic pretraining to improve on low-resource tasks. |
| Outcome: | The proposed model can predict extracted gap sentences on summarization with a low resource and zero shot setup. |
Copied to clipboard
| Challenge: | Existing methods for summarizing semantic graph structure from raw text are cumbersome and inefficient for long-text documents. |
| Approach: | They propose a Transformer-based pre-trained model with multi-granularity sparse attentions for long-text extractive summarization. |
| Outcome: | The proposed model performs state-of-the-art on single- and multi-document summarization tasks while using less memory and fewer parameters. |
Copied to clipboard
| Challenge: | Embedding based methods are widely used for unsupervised keyphrase extraction tasks. |
| Approach: | They propose a method where local and global contexts are jointly modeled. |
| Outcome: | The proposed method outperforms most models while generalizing better on input documents with different domains and length. |
Copied to clipboard
| Challenge: | Distantly supervised relation extraction is used in knowledge bases but its low quality and noisy sentences are present in sentence bags. |
| Approach: | They propose a multi-layer revision network which emphasizes inner-sentence correlations before extracting relevant information within sentences. |
| Outcome: | The proposed method improves on two New York Times datasets. |
Copied to clipboard
| Challenge: | Existing methods leverage programs that contain rich logical information to enhance the verification process. |
| Approach: | They propose a table-based fact verification task as an evidence retrieval framework . they retrieve logic-level program-like evidence from the given table and a statement as supplementary evidence for the table . |
| Outcome: | The proposed method is able to retrieve logic-level program-like evidence from a table and a statement as supplementary evidence for the table. |
Copied to clipboard
| Challenge: | Existing approaches to extract entity and relation feature are flawed because they do not consider the intimate connection between NER and RE. |
| Approach: | They propose a partition filter network to model two-way interaction between tasks . they leverage two gates: entity and relation gate, to segment neurons into two task partitions and one shared partition. |
| Outcome: | The proposed model performs significantly better than previous approaches on six public datasets. |
Copied to clipboard
| Challenge: | Existing methods to label data and identify entities require large amounts of manually annotated texts for training supervised models. |
| Approach: | They propose a dictionary extension method which extracts new entities through the type expanded model. |
| Outcome: | The proposed method outperforms state-of-the-art supervised systems on different types of datasets and surpasses supervised models. |
Copied to clipboard
| Challenge: | Existing methods for aspect category sentiment analysis do not necessarily occur in a sentence. |
| Approach: | They propose a Beta Distribution-guided aspect-aware graph construction based on external knowledge . they use aspect-related words as the pivots to derive aspect-relevant weights . |
| Outcome: | The proposed approach outperforms the state-of-the-art methods on 6 benchmark datasets. |
Copied to clipboard
| Challenge: | Existing methods for pre-training can be sub-optimal in some cases . for example, aspect extraction tasks require domain and category invariant representations . |
| Approach: | They propose a domain-invariant learning scheme for BERT to fine-tune pre-trained language models on a source domain and then apply it to a different target domain. |
| Outcome: | The proposed scheme improves performance over state-of-the-art models while using fraction of the unlabeled data. |
Copied to clipboard
| Challenge: | Multimodal sentiment analysis is a trending area of research, and multimodal fusion is one of its most active topics. |
| Approach: | They propose to use modality-based penalties to measure dependency between models to improve accuracy. |
| Outcome: | The proposed methods improve accuracy on two well-known sentiment analysis datasets by 4.3 on the proposed models and by-product includes a statistical network which can interpret the high dimensional representations learnt by the model. |
Copied to clipboard
| Challenge: | Recent studies have focused on identifying the sentiment polarity of aspects in product reviews. |
| Approach: | They propose to use supervised Contrastive Pre-Training to learn implicit sentiment . they propose to train large-scale sentiment-annotated corpora from in-domain language resources . |
| Outcome: | The proposed model achieves state-of-the-art performance on SemEval2014 benchmarks and comprehensively validates its effectiveness on learning implicit sentiment. |
Copied to clipboard
| Challenge: | Existing approaches to extract aspect terms from review sentences are limited due to lack of annotated data. |
| Approach: | They propose to refine conventional self-training to progressive self-teaching to reduce noise . they use a discriminator to filter the noisy pseudo-labels. |
| Outcome: | The proposed model outperforms baseline models and achieves state-of-the-art performance on four SemEval datasets. |
Copied to clipboard
| Challenge: | Existing approaches to improve generalization ability by augmenting training data with synonymous examples or adding random noises to word embeddings cannot address spurious association problem. |
| Approach: | They propose an end-to-end reinforcement learning framework which jointly performs counterfactual data generation and dual sentiment classification. |
| Outcome: | The proposed framework outperforms strong data augmentation baselines on several benchmark sentiment classification datasets. |
Copied to clipboard
| Challenge: | Structured social variation has been extensively studied, e.g., gender based variation, but little is known about how to characterize individual styles due to their idiosyncratic nature. |
| Approach: | They propose a method to study idiolects through a massive cross-author comparison to identify and encode stylistic features. |
| Outcome: | The proposed model achieves strong performance at authorship identification on short texts and through an analogy-based probing task, showing that the learned representations exhibit surprising regularities that encode qualitative and quantitative shifts of idiolectal styles. |
Copied to clipboard
| Challenge: | a growing body of theoretical work on narrative has been focused on the field of natural language processing . this position paper aims to provide a unifying framework for the computational study of narrative . |
| Approach: | They propose to introduce dominant theoretical frameworks to the NLP community and situate current research within distinct narratological traditions. |
| Outcome: | The proposed framework would allow for new empirical questions and applications in the field of natural language processing. |
Copied to clipboard
| Challenge: | Existing stance detection methods have been evaluated in comparison to the public opinion data they promise to replace. |
| Approach: | They propose to compare an individual's self-reported stance to the stance inferred from their social media data. |
| Outcome: | The proposed models are compared to a public opinion survey with 1,129 individuals across four salient targets. |
Copied to clipboard
| Challenge: | Recent studies have shown that models trained on CAD can learn cues in the dataset which are spuriously correlated with the construct. |
| Approach: | They focus on sentiment, sexism, and hate speech as social constructs to investigate their effects on model performance. |
| Outcome: | The proposed model generalizes better on out-of-domain datasets while relying less on spurious features. |
Copied to clipboard
| Challenge: | Existing studies on explicit or overt hate speech have failed to address a more pervasive form based on coded or indirect language. |
| Approach: | They propose a theoretically-justified taxonomy of implicit hate speech and a benchmark corpus with fine-grained labels for each message and its implication. |
| Outcome: | The proposed dataset will serve as a useful benchmark for understanding this multifaceted issue. |
Copied to clipboard
| Challenge: | Knowledge distillation is a major technique for deploying vast language models in resource-strapped environments. |
| Approach: | They propose a method that transfers contextual knowledge via Word Relation and Layer Transforming Relation. |
| Outcome: | The proposed method is able to transfer contextual knowledge without restrictions on architectural changes between teacher and student on language understanding tasks. |
Copied to clipboard
| Challenge: | Existing methods conduct knowledge distillation statically, e.g., student model aligns output distribution to teacher model on pre-defined training dataset. |
| Approach: | They propose a dynamic knowledge distillation that empowers the student to adjust the learning procedure according to its competency . they find it is promising and provide discussions on potential future directions towards more efficient methods . |
| Outcome: | The proposed method can boost student model performance while accelerating training . the proposed method reduces memory usage and accelerates model inference . |
Copied to clipboard
| Challenge: | Existing approaches to text generation combine task descriptions and examples with supervised learning. |
| Approach: | They propose a method for text generation that is based on pattern-exploiting training. |
| Outcome: | The proposed approach improves on several summarization and headline generation datasets. |
Copied to clipboard
| Challenge: | Sentence Compression (SC) is an important natural language processing task . it aims to shorten sentences while preserving the original meanings of the words . improvements on Chinese SC models are still lacking due to several difficulties . |
| Approach: | They propose a neural Chinese SC model enhanced with a Self-Organizing Map from Chinese colloquial sentences from a real-life question answering system. |
| Outcome: | The proposed model achieves a promising F1 score of 89.655 and BLEU4 score of 70.116 . it improves the performance of the whole neural Chinese SC model in a valid manner . |
Copied to clipboard
| Challenge: | Multi-task auxiliary learning uses a set of relevant auxiliary tasks to improve performance of a primary task. |
| Approach: | They propose a time-efficient sampling method to select the most beneficial sub-datasets from the auxiliary tasks to achieve efficient multi-task auxiliary learning. |
| Outcome: | The proposed method significantly outperforms random sampling and ST-DNN on three benchmark datasets. |
Copied to clipboard
| Challenge: | Prior methods for detecting out-of-scope (OOS) utterances in text are limited and require a limited amount of data to obtain. |
| Approach: | They propose an orthogonal technique that augments existing data to train better OOS detectors operating in low-data regimes. |
| Outcome: | The proposed method outperforms existing methods on key metrics across three benchmarks and achieves relative gains of 52.4%, 48.9% and 50.3%. |
Copied to clipboard
| Challenge: | Dialogue-based relation extraction (RE) aims to extract relation(s) between two arguments that appear in a dialogue. |
| Approach: | They propose a dialogue-based relation extraction model which is based on emotion recognition in conversations. |
| Outcome: | The proposed model outperforms the state-of-the-art models on most of the benchmark datasets. |
Copied to clipboard
| Challenge: | Recent work suggests crowdworkers goad dialog models into generating unsafe and inconsistent responses, but humans leverage superficial clues such as hate speech, while leaving systematic problems undercover. |
| Approach: | They propose two methods to automatically trigger a dialog model into generating problematic responses by reinforcement learning. |
| Outcome: | The proposed methods expose safety and contradiction issues with state-of-the-art dialog models. |
Copied to clipboard
| Challenge: | Annotating CDCR data is laborious and expensive, explaining why existing corpora are small and lack domain coverage. |
| Approach: | They use hyperlinks to extract event coreference data from online news articles . they find that models trained on small subsets of HyperCoref are highly competitive . |
| Outcome: | The proposed system frees up CDCR research from costly human-annotated training data and opens up possibilities beyond English. |
Copied to clipboard
| Challenge: | Stereotypical character roles are important aids to narrative understanding and are often referred to as archetypes or dramatis personae. |
| Approach: | They propose an unsupervised method for learning stereotypical roles given only structural plot information using Vladimir Propp’s structural theory of Russian folktales. |
| Outcome: | The proposed method induces six out of seven of Vladimir Propp’s dramatis personae with F1 measures of up to 0.70 (0.58 average), with an additional category for minor characters. |
Copied to clipboard
| Challenge: | Discourse learning is a complex task, and schemas evolve across annotation efforts preventing compilation of smaller datasets into larger ones. |
| Approach: | They propose a multitask learning approach that can combine discourse datasets from similar and diverse domains to improve discourse classification. |
| Outcome: | The proposed approach improves on the NewsDiscourse dataset by 4.9% over current state-of-the-art benchmarks on one of the largest discourse datasets. |
Copied to clipboard
| Challenge: | Infrequent names are less similar to initial representations, and are more self-similar, suggesting that models rely on less context-informed representations of uncommon and minority names. |
| Approach: | They use a dataset of U.S. first names with labels based on predominant gender and racial group to examine effect of training corpus frequency on tokenization, contextualization, similarity to initial representation, and bias. |
| Outcome: | The results show that infrequent names are less similar to initial representations and have a Spearman’s rho between frequency and self-similarity as low as .763 . |
Copied to clipboard
| Challenge: | Ethnic bias is one of the most prevalent social stereotypes. |
| Approach: | They propose to use a multilingual model and contextual word alignment to mitigate ethnic bias in monolingual BERT for English, German, Spanish, Korean, Turkish, and Chinese. |
| Outcome: | The proposed methods alleviate ethnic bias in English, German, Spanish, Korean, Turkish, and Chinese using a multilingual model and contextual word alignment of two monolingual models. |
Copied to clipboard
| Challenge: | Existing frameworks to debias contextual representations can encode undesirable attributes, like demographic associations of the users, while being trained for an unrelated task. |
| Approach: | They propose an adversarial learning framework to debias contextual representations by encoding undesirable attributes while being trained for an unrelated task. |
| Outcome: | The proposed framework debiases representations on 8 datasets while remaining informative on the target task. |
Copied to clipboard
| Challenge: | Currently, natural language inputs are unclear or ambiguous, causing uncertainty in dialogues. |
| Approach: | They propose a framework for building a visually grounded question-asking model capable of producing polar (yes-no) clarification questions to resolve misunderstandings in dialogue. |
| Outcome: | The proposed model can produce polar (yes-no) clarification questions to resolve misunderstandings in a goal-oriented 20 questions game with synthetic and human answerers. |
Copied to clipboard
| Challenge: | Existing text infilling objectives for pretrained language models require self-supervision by masking out tokens or spans in text. |
| Approach: | They propose to extend text infilling to a self-supervised sequence-to-sequence (Seq2Sequen) task. |
| Outcome: | The proposed task improves the model's performance on various natural language generation tasks. |
Copied to clipboard
| Challenge: | Existing text classification methods focus on a fixed label set, but many real-world applications require extending to new fine-grained classes as the number of samples per label increases. |
| Approach: | They propose a problem called coarse-to-fine grained classification that leverages label surface names as the only human guidance. |
| Outcome: | The proposed method outperforms existing methods on two real-world datasets. |
Copied to clipboard
| Challenge: | Existing databases contain tens of millions of molecules; PubChem alone has 110 million compounds. |
| Approach: | They propose a task to retrieve molecules using natural language descriptions as queries . they construct a paired dataset of molecules and their corresponding text descriptions . |
| Outcome: | The proposed approach improves results from 0.372 to 0.499 MRR. |
Copied to clipboard
| Challenge: | Fig. 1 shows a simplified CT protocol. |
| Approach: | They propose to use geometric deep learning to classify hierarchical documents into different categories by using a selective graph pooling operation that arises from the fact that some parts of the hierarchy are invariable across different documents. |
| Outcome: | The proposed model achieves f1-scores around 0.85 on a publicly available large scale CT registry of around 360K protocols. |
Copied to clipboard
| Challenge: | Recent studies show that basic configurations can improve the performance of neural networks on systematic generalization. |
| Approach: | They propose to revisit basic configurations to improve the performance of Transformers on systematic generalization by revisiting scaling of embeddings, early stopping, relative positional embeddment, and Universal Transformer variants. |
| Outcome: | The proposed models improve accuracy from 50% to 85% on the PCFG productivity split and from 35% to 81% on COGS. |
Copied to clipboard
| Challenge: | Existing methods for text detection lack interpretability and robustness towards unseen models. |
| Approach: | They propose three new types of interpretable topological features based on topological data analysis which is currently understudied in the field of NLP. |
| Outcome: | The proposed features outperform count- and neural-based baselines up to 10% on three common datasets and tend to be the most robust towards unseen GPT-style generation models. |
Copied to clipboard
| Challenge: | Using uncertainty and diversity sampling, active learning acquisition functions select difficult and diverse data points from a pool of unlabeled data. |
| Approach: | They propose an active learning acquisition function that selects contrastive examples from unlabeled data. |
| Outcome: | The proposed approach performs better or equal to the best performing baseline on all tasks, on both in-domain and out-of-domain data. |
Copied to clipboard
| Challenge: | Existing methods for beam search are based on a deterministic approach, but the results are not as accurate as those used in SBS. |
| Approach: | They propose a method that turns beam search into a stochastic process by using conditional Poisson sampling design instead of taking the maximizing set at each iteration. |
| Outcome: | The proposed method produces lower variance and more efficient estimators than SBS, even showing improvements in high entropy settings. |
Copied to clipboard
| Challenge: | Existing approaches to generate synthetic data using simple sentence transformations and/or model-based techniques may not generate realistic error samples with respect to the NLG models. |
| Approach: | They propose a framework to train models to classify acceptability of responses generated by natural language generation models using a 2-stage approach . they use existing sentence transformations to generate samples that better resemble the output of the generation models. |
| Outcome: | The proposed approach outperforms existing techniques and can be used in few-shot settings using self-training. |
Copied to clipboard
| Challenge: | aaron carroll: in social settings, human behavior is governed by unspoken rules of conduct rooted in societal norms . carroll and colleagues examine whether language generation models can serve as behavioral priors if they are not . they say we examine whether they can generate descriptions of actions that accomplish predefined goals . |
| Approach: | They propose to combine multiple expert models to improve quality of generated actions, consequences, and norms. |
| Outcome: | The proposed models significantly improve the quality of generated actions, consequences, and norms compared to baselines. |
Copied to clipboard
| Challenge: | Existing models with attention mechanisms can generate fluent descriptions of salient patterns in time series, but they often generate factually incorrect descriptions. |
| Approach: | They propose a model which first runs small learned programs on the input time series, then identifies the programs/patterns which hold true for the given input, and finally conditions on *only* the chosen valid program to generate the output text description. |
| Outcome: | The proposed model extracts high-level patterns from the data and generates high precision captions even though it is built on a small space of modules. |
Copied to clipboard
| Challenge: | Recent advances in deep generative modeling have led to significant advances in natural language generation (NLG). |
| Approach: | They propose to model the entity type carefully in the decoding phase to generate contextual words accurately. |
| Outcome: | The proposed model produces a target sequence based on a given list of entities. |
Copied to clipboard
| Challenge: | Recent work on multilingual AMR-to-text generation has focused on data augmentation strategies that utilize generated silver AMRs, but this assumes a high quality of generated AMR. |
| Approach: | They propose to combine gold AMR with silver AMRs to generate multilingual AMR annotations. |
| Outcome: | The proposed models outperform the current state of the art for German, Italian, Spanish, and Chinese by a large margin. |
Copied to clipboard
| Challenge: | Recent advances in machine translation and multilingual text generation have led researchers to adopt trained metrics such as COMET or BLEURT, which treat evaluation as a regression problem and use representations from multilingual pre-trained models such as XLM-RoBERTa or mBERT. |
| Approach: | They propose to use multilingual model capacity to improve model performance by transferring knowledge from one teacher to multiple students trained on related languages. |
| Outcome: | The proposed model yields 10.5% improvement over vanilla fine-tuning and reaches 92.6% of RemBERT’s performance using only a third of its parameters. |
Copied to clipboard
| Challenge: | Several modifications have been proposed to improve monolingual language models, but none of them result in better multilingual models. |
| Approach: | They propose to add positional encodings to token embeddings to preserve word-order information in a non-autoregressive setting. |
| Outcome: | The proposed modifications tend to improve monolingual models, but none improve multilingual models. |
Copied to clipboard
| Challenge: | Large pretrained models such as BERT encode a range of features into monolithic vectors, providing strong predictive accuracy across downstream tasks. |
| Approach: | They explore whether it is possible to learn disentangled representations by identifying existing subnetworks within pretrained models that encode distinct, complementary aspects. |
| Outcome: | The proposed method disentangles sentiment from genre in movie reviews, toxicity from dialect in Tweets, and syntax from semantics. |
Copied to clipboard
| Challenge: | Recent studies have focused on enhancing existing models with the primary objective of improving downstream performance on various NLP tasks. |
| Approach: | They propose to use BERT to encode meaningful knowledge in token representations to explain probing results. |
| Outcome: | The proposed model can detect syntactic and semantic abnormalities and distinguish between grammatical number and tense subspaces. |
Copied to clipboard
| Challenge: | Language models are generally trained on short, truncated input sequences, which limits their ability to use discourse-level information present in long-range context to improve their predictions. |
| Approach: | They analyze two long-range Transformer language models that accept 8K token inputs . they find that providing long-term context only improves their predictions on a small set of tokens - not sentence-level ones . |
| Outcome: | The proposed model improves on PG-19 with only 2K tokens and does not help at all for sentence-level prediction tasks. |
Copied to clipboard
| Challenge: | Recent work has raised concerns about the inherent limitations of text-only pretraining. |
| Approach: | They first generate a color dataset of human-perceived color distributions for 521 common objects and then use it to analyze and compare the color distribution found in text and the distribution captured by language models. |
| Outcome: | The proposed model improves on the CoDa color distribution, while the language model improve on the ground-truth distribution. |
Copied to clipboard
| Challenge: | Existing models that explain text classification predictions are opaque and overfit to spurious artifacts. |
| Approach: | They propose a novel self-explaining model that explains a text classifier’s predictions using phrase-based concepts. |
| Outcome: | The proposed model shows that it is adequate, trustworthy and understandable by human judges compared to existing baselines. |
Copied to clipboard
| Challenge: | Detecting salient events is an essential part of understanding narrative, and is used to aid storyline writing and summarisation. |
| Approach: | They propose an unsupervised method for salience detection derived from Barthes Cardinal Functions and theories of surprise and apply it to longer narrative forms. |
| Outcome: | The proposed method improves performance over a non-knowledgebase and memory augmented language model on longer works. |
Copied to clipboard
| Challenge: | Existing novelty detection algorithms are coarse-grained, working at the document or topic level. |
| Approach: | They propose to use a fine-grained semantic novelty detection problem to solve a novel novel scene problem. |
| Outcome: | The proposed model outperforms baseline models on the proposed task by large margins. |
Copied to clipboard
| Challenge: | Prior work has addressed ‘cold start’ estimation of item difficulties without piloting, but a multi-task generalized linear model with BERT features is needed to jump-start new items without pilot. |
| Approach: | They propose a multi-task generalized linear model with BERT features to jump-start new item difficulties without piloting them first. |
| Outcome: | The proposed model compares test-taker proficiency, item difficulty, and language proficiency frameworks like the Common European Framework of Reference (CEFR). |
Copied to clipboard
| Challenge: | Existing methods fail to complete voice queries from incomplete prefixes because they use orthographic prefix and substrings instead of the true phonetic prefix. |
| Approach: | They propose to condition QAC approaches on intermediate transcriptions to complete voice queries. |
| Outcome: | The proposed method obtains an 18% relative improvement over previous methods on a speech-enabled smart television with real-life voice search traffic. |
Copied to clipboard
| Challenge: | Large-Scale Multi-Label Text Classification (LMTC) tasks with hierarchical label spaces include automatic assignment of ICD-9 codes to discharge summaries. |
| Approach: | They propose a set of metrics for hierarchical evaluation using the depth of the ontology to evaluate the predictions of neural LMTC models. |
| Outcome: | The proposed metrics compare with previous evaluations on prior art models for ICD-9 coding in MIMIC-III and propose further avenues of research involving the proposed representation. |
Copied to clipboard
| Challenge: | authorship verification has traditionally relied on modeling stylometric linguistic properties . but neural methods introduce a tradeoff: they obviate the need for manual feature design . |
| Approach: | They propose to use domain-specific features to improve authorship representations . they propose to study Amazon reviews, fanfiction short stories, and Reddit comments . |
| Outcome: | The proposed methods outperform existing methods in large-scale authorship verification scenarios. |
Copied to clipboard
| Challenge: | Natural language relies on a finite lexicon to express an unbounded set of emerging ideas. |
| Approach: | They propose a framework that exploits the cognitive mechanisms of chaining and multimodal knowledge to predict emergent compositional expressions through time. |
| Outcome: | The proposed framework predicts emergent compositions through time using cognitive mechanisms . it is based on modal knowledge and categorizing models of chaining in a syntactically parsed English corpus . |
Copied to clipboard
| Challenge: | Pre-trained language models perform well on a variety of linguistic tasks that require symbolic reasoning, raising the question of whether such models implicitly represent abstract symbols and rules. |
| Approach: | They investigate the performance of BERT on English subject–verb agreement by analyzing word frequency and absolute frequency of verb forms. |
| Outcome: | The proposed model generalizes well to subject–verb pairs that never occurred in training, suggesting a degree of rule-governed behavior. |
Copied to clipboard
| Challenge: | Throughout human evolution, countless languages have evolved, each with unique features. |
| Approach: | They analysed a corpus of 600 languages to find strong evidence for a surprisal–duration trade-off between languages and languages. |
| Outcome: | The proposed model shows that phones are produced faster in languages where they are less surprising and vice versa. |
Copied to clipboard
| Challenge: | The uniform information density hypothesis posits a preference among language users for utterances structured such that information is distributed uniformly across a signal. |
| Approach: | They propose to test the hypothesis by using reading time and acceptability data to examine the effect of surprisal on language comprehension and acceptabilities. |
| Outcome: | The proposed hypothesis makes predictions about language comprehension and linguistic acceptability . |
Copied to clipboard
| Challenge: | Prior work fine-tunes deep LMs to encode text sequences into single dense vector representations, but dense encoders require a lot of data and sophisticated techniques to train and suffer in low data situations. |
| Approach: | They propose to pre-train Transformer language models (LMs) with a novel Transformer architecture, Condenser, where LM prediction CONditions on DENSE Representation. |
| Outcome: | The proposed model improves on various text retrieval and similarity tasks by large margins over standard LMs. |
Copied to clipboard
| Challenge: | a recent study shows that slow emerging topics are often detected too late . a positive correlation is linked with event-like topics while a negative correlation is a sign of emergence. |
| Approach: | They propose to monitor words representation in embedding space and use one of its geometrical properties to characterize the emergence of topics. |
| Outcome: | The proposed method outperforms state-of-the-art methods on two public datasets of press and scientific articles. |
Copied to clipboard
| Challenge: | Existing approaches to conversational search use multiple inference pipelines that require long inference times . despite their effectiveness, such a pipeline often includes multiple neural models that require longer inference time. |
| Approach: | They propose to integrate conversational query reformulation directly into a dense retrieval model . they use a dataset with pseudo-relevance labels to overcome the lack of training data . |
| Outcome: | The proposed model rewrites conversational queries as dense representations in conversational search and open-domain question answering datasets. |
Copied to clipboard
| Challenge: | Recent approaches to information retrieval (IR) and natural language processing (NLP) use contextual language models, which can improve both synonymy and polysemy problems associated with words. |
| Approach: | They propose an ultra-high dimensional representation scheme equipped with directly controllable sparsity and a bucketing method where embeddings from multiple layers of BERT are selected/merged to represent diverse linguistic aspects. |
| Outcome: | The proposed representation scheme outperforms sparse models with MS MARCO and TREC CAR, and shows that it is highly efficient for storage and search. |
Copied to clipboard
| Challenge: | Recent advances in contextualized embeddings have made ranking on non-English documents cumbersome . a novel multilingual query expansion mechanism provides sense definitions as additional semantic information for the query. |
| Approach: | They propose a multilingual query expansion mechanism that leverages word sense information to enhance the model's performance. |
| Outcome: | The proposed model performs better than its supervised and unsupervised alternatives across languages while being trained on English Robust04 data. |
Copied to clipboard
| Challenge: | Neural topic models (NTMs) use deep neural networks to learn topic information. |
| Approach: | They propose a variational autoencoder model that reconstructs sentence and document word counts using bag-of-words embeddings and pre-trained semantic embedders. |
| Outcome: | The proposed model lowers reconstruction errors at sentence and document levels and finds more coherent topics from real-world datasets. |
Copied to clipboard
| Challenge: | Existing knowledge bases are organized according to manual schemas that limit their expressiveness and require significant human engineering and maintenance. |
| Approach: | They propose to organize knowledge representation strategies in LMs by the level of KB supervision provided . they propose to highlight notable models, evaluation tasks, and findings . |
| Outcome: | The proposed model can internalize and express relational knowledge in more flexible forms. |
Copied to clipboard
| Challenge: | Existing techniques for certifying robustness of LSTMs and extensions of lsts are prone to adversarial examples. |
| Approach: | They propose an approach to certify robustness of LSTMs and extensions of lstms . they show that their approach can train models more robust to combinations of string transformations - a key advantage of existing certification approaches . |
| Outcome: | The proposed approach can show high certification accuracy of the resulting models. |
Copied to clipboard
| Challenge: | Existing approaches to generate relevant Knowledge Bases from text and graph data are gaining popularity. |
| Approach: | They propose a bidirectional generation of text and graph leveraging Reinforcement Learning. |
| Outcome: | The proposed system improves on WebNLG+ 2020 and TekGen datasets. |
Copied to clipboard
| Challenge: | Pretrained Transformers achieve remarkable performance when training and test data are from the same distribution, but in real-world scenarios, out-of-distribution instances can cause semantic shift problems. |
| Approach: | They propose to fine-tune the Transformers with a contrastive loss, which improves the compactness of representations, and to use the Mahalanobis distance in the model's penultimate layer to detect OOD instances. |
| Outcome: | The proposed method outperforms baselines in the real-world and achieves near-perfect OOD detection performance. |
Copied to clipboard
| Challenge: | Creating embodied, situated agents able to move in, communicate naturally about, and collaborate on human terms in the physical world has been a persisting goal in artificial intelligence (Winograd, 1972). |
| Approach: | They propose to use a 3D Minecraft dataset to model the beliefs of human partners in situ to enable theory of mind modeling in situated interactions. |
| Outcome: | The proposed model can be used to model human collaborative behaviors in the 3D virtual blocks world of Minecraft. |
Copied to clipboard
| Challenge: | Existing studies on personas are pre-defined and hard to obtain before a conversation . a new task aims to detect speaker persona based on conversational text . |
| Approach: | They propose a task to detect speaker personas based on conversational text . they build a dataset for SPD and propose utterance-to-profile matching networks . |
| Outcome: | The proposed task outperforms baseline models and utterance-to-profile (U2P) matching networks. |
Copied to clipboard
| Challenge: | Existing methods to make multilingual systems expensive and tedious introduce pipeline of errors. |
| Approach: | They propose to use pre-trained multilingual models to enhance the transfer learning process by intermediate fine-tuning of pretrained multi-lingual models. |
| Outcome: | The proposed approach improves on the cross-lingual dialogue state tracking task with only 10% of the target language task data and zero-shot setup respectively. |
Copied to clipboard
| Challenge: | Existing Transformer-based language models (LMs) are not effective as sentence encoders when used off-the-shelf. |
| Approach: | They propose a method which turns a pretrained LM into a universal conversational encoder and task-specialised sentence encoder. |
| Outcome: | The proposed framework achieves state-of-the-art ID performance across the board with particular gains in the most challenging, few-shot setups. |
Copied to clipboard
| Challenge: | Dialogs are a building block of human natural language interactions. |
| Approach: | They propose a new edit distance metric for dialog similarity analysis using conversation semantics, conversation flow, and the participants. |
| Outcome: | The proposed method outperforms existing methods on two publicly available datasets and is better aligned with human perception of conversation similarity. |
Copied to clipboard
| Challenge: | Recent work attempts to apply incremental processing to NLUs but this is computationally expensive and does not scale efficiently for long sequences. |
| Approach: | They propose to apply Transformers incrementally via restart-incrementality by repeatedly feeding, to an unchanged model, increasingly longer input prefixes to produce partial outputs. |
| Outcome: | The proposed model has better incremental performance and faster inference speed compared to the standard Transformer and LT with restart-incrementality, at the cost of part of the non-incremental quality. |
Copied to clipboard
| Challenge: | a large amount of labeled data is needed for fine-tuning. |
| Approach: | They propose attribution methods inspired by multi-agent reinforcement learning for a feedback attribution problem in spoken language understanding. |
| Outcome: | The proposed methods can train competitive models from user feedback. |
Copied to clipboard
| Challenge: | Relation extraction systems require large amounts of labeled examples which are costly to annotate. |
| Approach: | They propose to use hand-made relation extraction tasks to refine a pretrained textual entailment engine which is run as-is or further fine-tuned on labeled examples. |
| Outcome: | The proposed system achieves 63% F1 zero-shot, 69% with 16 examples per relation and 4 points short of the state-of-the-art system on the same conditions. |
Copied to clipboard
| Challenge: | Generating or modifying graphs from natural language text has applications in many subfields, such as dependency parsing or knowledge graph construction. |
| Approach: | They propose a method that first embeds the graph and the instructions with a joint encoder and then rebuilds it using a separate generative model for graphs conditioned on h. |
| Outcome: | The proposed method improves accuracy on three scene graph modification data sets while the state-of-the-art fails to generalize. |
Copied to clipboard
| Challenge: | a number of information extraction tasks require task-specific training. |
| Approach: | They propose a text-to-triple translation framework for information extraction tasks . they propose enabling task-agnostic translation by leveraging latent knowledge of a pre-trained language model . |
| Outcome: | The proposed framework outperforms the existing methods on open information extraction tasks. |
Copied to clipboard
| Challenge: | Existing models for document-level relation extraction relied on implicitly powerful representations, which makes the model less transparent. |
| Approach: | They propose a probabilistic model for document-level relation extraction by learning logic rules. |
| Outcome: | The proposed model outperforms baseline models in relation performance and logical consistency. |
Copied to clipboard
| Challenge: | Existing empathetic datasets are limited in size and cost due to the cost of manual labor. |
| Approach: | They propose to annotate 1M dialogues with 32 fine-grained emotions and eight empathetic response intents and the Neutral category using a silver dataset. |
| Outcome: | The proposed pipeline compares the quality of the proposed dataset with a state-of-the-art gold dataset using offline experiments and visual validation methods. |
Copied to clipboard
| Challenge: | Recent research has focused on open-ended text generation tasks because they are difficult to evaluate automatically. |
| Approach: | They conduct a survey of 45 open-ended text generation papers to determine whether models are reproducible . they then run story evaluation experiments with AMT workers and English teachers . |
| Outcome: | The results show that AMT workers and English teachers perform better when shown model-generated output alongside human-generated references. |
Copied to clipboard
| Challenge: | Large text corpora are often introduced with minimal documentation . documenting collection process, composition, intended uses, and other are key for structured, task-specific datasets. |
| Approach: | They propose to document a dataset created by applying filters to a single snapshot of Common Crawl. |
| Outcome: | The proposed dataset shows that blocklist filtering removes text from minority individuals and patents. |
Copied to clipboard
| Challenge: | Existing reproducible benchmarks for machine translation are limited to high-resource or well-represented languages. |
| Approach: | They propose to use AfroMT to develop a reproducible machine translation benchmark for eight widely spoken African languages and a suite of analysis tools to take into account their unique properties. |
| Outcome: | The proposed benchmarks show significant improvements when pretraining on 11 languages, with gains of up to 2 BLEU points over strong baselines. |
Copied to clipboard
| Challenge: | a lack of standardized and reliable methods for automatic evaluation hinders ST . prior work has employed as many as nine different automatic systems to rate formality alone . |
| Approach: | They evaluate automatic metrics on the oft-researched task of formality style transfer . they outline best practices for automatic evaluation in (formality) style transfer and identify models that correlate well with human judgments. |
| Outcome: | The proposed models correlate well with human judgments and are robust across languages. |
Copied to clipboard
| Challenge: | Material science synthesis procedures require high-quality annotations, which are limited by the size and quality of the annotations. |
| Approach: | They propose a corpus of entity mention annotations over 595 Material Science synthesis procedures. |
| Outcome: | The proposed approach greatly expands the training data available for the Named Entity Recognition task. |
Copied to clipboard
| Challenge: | Recent advances in pretrained language models do not capture nuanced biases in political discourse . a new approach to represent political content is to use contextualized embeddings to create effective representations . |
| Approach: | They propose a model that captures and leverages political content to generate more effective representations . they use tweets, press releases, issues, news articles and participating entities to generate composed representations. |
| Outcome: | The proposed model generates representations for political entities over multiple issues or events . qualitative and quantitative analysis shows that the model is meaningful and effective . |
Copied to clipboard
| Challenge: | Recent years have seen the successful application of span-based neural models to entity-based information extraction tasks such as entity coreference resolution (CR) Existing event coreference resolvers focused on feature engineering are few and far between, let alone event corefers. |
| Approach: | They propose to adapt existing span-based event reference systems to event coreference by adapting the models originally developed for entity coreference to event CR. |
| Outcome: | The proposed model improves the representations of entity mentions in entity-based IE tasks compared to non-span models . |
Copied to clipboard
| Challenge: | Discourse segmentation is the first step of discourse analysis. |
| Approach: | They propose a weak supervision approach to adapt a latent model to French conversation transcripts with a linguistic and acoustic input. |
| Outcome: | The proposed model improves in situations where speaker turns are lacking or noisy, gaining up to 13% in F-score. |
Copied to clipboard
| Challenge: | a novel approach to narrative event representation uses attention to re-contextualize events across the whole story . a recent study shows that attention is used to attach event semantics to tokens . |
| Approach: | They propose an unsupervised approach to narrative event representation using attention to re-contextualize events across the whole story. |
| Outcome: | The proposed approach achieves state of the art performance on multiple choice and story cloze tasks. |
Copied to clipboard
| Challenge: | Existing approaches simplify by considering coreference only within document clusters, but this fails to handle inter-cluster coreference, common in many applications. |
| Approach: | They propose to model entities/events in a reader’s focus as a neighborhood within a learned latent embedding space which minimizes the distance between mentions and the centroids of their gold coreference clusters. |
| Outcome: | The proposed model achieves state-of-the-art for events and entities on the ECB+, Gun Violence, Football Coreference, and Cross-Domain Cross-DDocument Coreference corpora. |
Copied to clipboard
| Challenge: | Storytelling is the communication of interesting and related events that form a concrete process. |
| Approach: | They propose methods for extracting the principal chain from natural language text . they filter away non-salient events and supportive sentences to isolate them . authors propose novel methods for predicting and answering events from text based on event-based temporal question answering . |
| Outcome: | The proposed method improves narrative prediction and event-based temporal question answering tasks. |
Copied to clipboard
| Challenge: | Existing approaches to question generation require conditioning on existing answers in text . previous work required human-curated templates, limiting coverage and question fluency . |
| Approach: | They propose a task of role question generation that produces a prototype and revises it to be contextually appropriate for the passage. |
| Outcome: | The proposed model generates diverse and well-formed questions for a large, broad-coverage ontology of predicates and roles. |
Copied to clipboard
| Challenge: | Existing studies have shown that pretrained Masked Language Models are not effective as universal lexical and sentence encoders off-the-shelf, i.e., without further task-specific fine-tuning on NLI, sentence similarity, or paraphrasing tasks using annotated task data. |
| Approach: | They propose a contrastive learning technique which turns pretrained MLMs into effective universal lexical and sentence encoders without additional data. |
| Outcome: | The proposed technique can turn MLMs into effective universal lexical and sentence encoders even without additional data. |
Copied to clipboard
| Challenge: | Pre-trained language models (PLMs) are limited in their ability to capture and use common-sense knowledge. |
| Approach: | They propose to teach PLMs how to reason with soft Horn rules by leveraging logical rules to learn how to predict precise probabilities. |
| Outcome: | The proposed model performs well on logical rules that were unseen at training. |
Copied to clipboard
| Challenge: | Existing studies on "gender bias" and "racial bias" focus on stereotypical attributes of word representations . a new method to elicit stereotypical information is proposed to capture stereotypical traits in language models . |
| Approach: | They propose a method to elicit stereotypical information from pretrained language models . they use fine-tuning on news sources to study their emotional effects . |
| Outcome: | The proposed method can be used to analyze emotion and stereotype shifts due to linguistic experience using fine-tuning on news sources. |
Copied to clipboard
| Challenge: | Existing systems for word Sense Disambiguation assume that each word can be disambiguated individually . a novel approach to WSD is proposed to address this limitation . |
| Approach: | They propose a supervised semantics-based approach to Word Sense Disambiguation that takes into account the senses assigned to nearby words. |
| Outcome: | The proposed approach surpasses all its competitors and sets a new state of the art on English WSD. |
Copied to clipboard
| Challenge: | a recent study has found that commonsense reasoning models are learning transferable generalizations . commonsensibility is a human capacity that has been a core challenge to Artificial Intelligence since its inception. |
| Approach: | They conduct an analysis of benchmarks that involve commonsense reasoning . they find that most datasets experimented with are problematic . commonsensence is a quintessential human capacity . |
| Outcome: | The proposed model is able to perform well on commonsense reasoning tasks . the model is not learning transferable generalizations or taking advantage of shortcuts . |
Copied to clipboard
| Challenge: | Differential privacy provides a formal approach to privacy of individuals. |
| Approach: | They propose to use ADePT to provide differentially private auto-encoders for text rewriting to provide tight privacy guarantees for users' original utterances. |
| Outcome: | The proposed algorithm is not differentially private, thus rendering the experimental results unsubstantiated. |
Copied to clipboard
| Challenge: | Discrete adversarial attacks are symbolic perturbations to a language input that preserve the output label but lead to predicting error. |
| Approach: | They propose a discrete adversarial attack based on best-first search and random sampling attacks that are not based upon expensive search procedures. |
| Outcome: | The proposed attack outperforms offline augmentation and speedups on three datasets. |
Copied to clipboard
| Challenge: | Recent debiasing methods in natural language understanding improve performance on out-of-distribution datasets by pressuring models into making unbiased predictions. |
| Approach: | They propose a general probing-based framework that allows for post-hoc interpretation of biases in language models and use an information-theoretic approach to measure the extractability of certain biase . |
| Outcome: | The proposed framework allows for post-hoc interpretation of biases in language models and measures the extractability of certain biase . |
Copied to clipboard
| Challenge: | High-performance neural language models have achieved state-of-the-art results on a wide range of NLP tasks, but results for common benchmark datasets often do not reflect model reliability and robustness when applied to noisy, real-world data. |
| Approach: | They propose to implement character-level and word-level perturbation methods to simulate scenarios in which input texts may be slightly noisy or different from the data distribution on which NLP systems were trained. |
| Outcome: | The proposed methods simulate scenarios in which input texts may be slightly noisy or different from the data distribution on which NLP systems were trained. |
Copied to clipboard
| Challenge: | Pretraining methods are convenient, but expensive in terms of time and resources. |
| Approach: | They investigate the impact of pretraining data size on the syntactic capabilities of RoBERTa by using syntaktic structural probes to determine whether models pretrained on more data encode a higher amount of syntastic information. |
| Outcome: | The proposed models perform better on part-of-speech tagging, dependency parsing and paraphrase identification. |
Copied to clipboard
| Challenge: | Pre-trained language models have shown impressive performance on downstream NLP tasks, but we have yet to establish a clear understanding of their sophistication when it comes to processing, retaining, and applying information presented in their input. |
| Approach: | They examine how robustly pre-trained LMs retain and apply relevant context information in the face of distracting content. |
| Outcome: | The proposed models retain and use critical context information in the face of distracting content, while models are susceptible to factors of semantic similarity and word position. |
Copied to clipboard
| Challenge: | Existing methods for producing model explanations seek all causal factors at once, making them difficult to comprehend. |
| Approach: | They propose a method to produce contrastive explanations in the latent space . they use attribution and token/span attribution to produce models that consider only contrastive reasoning . |
| Outcome: | The proposed method allows model behavior to consider only contrastive reasoning . it also uncovers which aspects of the input are useful for and against particular decisions . |
Copied to clipboard
| Challenge: | Existing studies show that deep neural networks are vulnerable to adversarial examples . a small perturbation to an input alters the model prediction . |
| Approach: | They propose a genetic algorithm to find models that can induce adversarial examples to fool models . they propose word replacement rules that can be used for model diagnostics from these examples . |
| Outcome: | The proposed model can fool almost all existing models, while ignoring the data bias in the training set. |
Copied to clipboard
| Challenge: | Existing methods for probing representations are limited to predicting part-of-speech . current methods cannot detect when a representation is predictive of just aspects of part- of-seech not explainable by the word identity. |
| Approach: | They propose to condition on the information in a baseline representation to test whether it is predictive of part-of-speech. |
| Outcome: | The proposed method is based on a theory of usable information called V-information and conditions on the information in the baseline. |
Copied to clipboard
| Challenge: | Recent studies have focused on gender bias in neural machine translation (NMT) incorrectly gendered translations can reflect or amplify social biases. |
| Approach: | They propose to use a monolingual corpus to generate gender-specific pseudo-parallel corpora and filter them to improve gender translation accuracy. |
| Outcome: | The proposed approach improves gender accuracy without damaging generic quality on translations from English into five languages. |
Copied to clipboard
| Challenge: | Unsupervised neural machine translation models perform well in low-resource or distant languages. |
| Approach: | They propose a model that leverages Wikipedia for machine translation and cross-lingual tasks without supervision from external parallel data or supervised models in target language. |
| Outcome: | The proposed model outperforms supervised models in Arabic and English translation tasks. |
Copied to clipboard
| Challenge: | Multilingual T5 pretrains a sequence-to-sequence model on monolingual texts, but it has shown promising results on many cross-lingual tasks. |
| Approach: | They propose a partially non-autoregressive objective for text-to-text pre-training and propose mT6 to improve cross-lingual transferability over multilingual T5. |
| Outcome: | The proposed model improves cross-lingual transferability over existing models. |
Copied to clipboard
| Challenge: | Pre-trained multilingual language encoders do not precisely align words and phrases across languages. |
| Approach: | They propose a learning strategy for training robust models by drawing connections between adversarial examples and failure cases of zero-shot cross-lingual transfer. |
| Outcome: | The proposed model can achieve good performance even if representations of different languages are not aligned well. |
Copied to clipboard
| Challenge: | Current approaches to speech-to-text translation (ST) use a pipeline of two sub-components - an automatic speech recognition (ASR) and a machine translation (MT) model. |
| Approach: | They propose an architecture that avoids initial lossy compression and aggregates information only at a higher level according to more informed linguistic criteria. |
| Outcome: | The proposed architecture achieves gains of up to 0.8 BLEU on the standard MuST-C corpus and up to 4.0 BLUE in a low resource scenario. |
Copied to clipboard
| Challenge: | Among rare words, named entities and domain-specific terms are crucial . previous studies have neglected these important words due to limited options . |
| Approach: | They propose a benchmark to evaluate automatic translation systems for rare words . named entities and domain-specific terms are crucial for their translation . |
| Outcome: | The proposed benchmark is based on European Parliament speeches annotated with NEs and terminology. |
Copied to clipboard
| Challenge: | HintedBT provides hints (as source tags on the encoder) about the quality of each source-target pair. |
| Approach: | They propose a method which provides hints to the encoder and decoder to improve the quality of BT data by providing hints about the quality. |
| Outcome: | The proposed method improves translation quality and performance in three low/medium-resource language pairs. |
Copied to clipboard
| Challenge: | Existing approaches to train simultaneous machine translation agents have been used to find the optimal action sequences for translation quality and lag. |
| Approach: | They propose a supervised learning approach that detects minimum reads required for generating target tokens by comparing simultaneous translations against full-sentence translations. |
| Outcome: | The proposed method produces much higher quality translations while minimizing the average lag in simultaneous translation. |
Copied to clipboard
| Challenge: | Existing pre-trained models can cause over-fitting when limited data are available. |
| Approach: | They propose to use a nearest-neighbor few-shot technique to improve cross-lingual adaptation using 16 distinct languages across two NLP tasks. |
| Outcome: | The proposed approach improves fine-tuning using only a handful of labeled samples in target locales and also generalizes across tasks. |
Copied to clipboard
| Challenge: | a series of experiments show that fine-tuning only the cross-attention parameters is nearly as effective as fine-timing all parameters. |
| Approach: | They conduct experiments to fine-tune a translation model on data where either the source or target language has changed. |
| Outcome: | The proposed model can be trained to several new languages with reduced parameter storage overhead. |
Copied to clipboard
| Challenge: | Evidence is emerging that neural networks learn due to inductive bias in the training routine, typically a variant of gradient descent (GD). |
| Approach: | They propose to characterize GD as an inductive bias in transformer training . they document norm growth in transformer language models and show they are saturated . |
| Outcome: | Empirically, we document norm growth in the training of transformer language models . the results suggest saturation is a new characterization of an inductive bias implicit in GD . |
Copied to clipboard
| Challenge: | Real-world applications often require improved models by leveraging a range of cheap incidental supervision signals. |
| Approach: | They propose a unified PAC-Bayesian motivated informativeness measure that characterizes the uncertainty reduction provided by incidental supervision signals. |
| Outcome: | The proposed measure quantifies the value added by incidental supervision signals to sequence tagging tasks. |
Copied to clipboard
| Challenge: | Recent work in NLP has documented dataset artifacts, bias, and spurious correlations . how to tell which features have spurious instead of legitimate correlations is typically left unspecified . |
| Approach: | They propose a class of competency problems to formalize this notion into a classification . they show that realistic datasets will increasingly deviate from competency problems . |
| Outcome: | The proposed model can be used to show that models are inappropriately affected by these less extreme biases. |
Copied to clipboard
| Challenge: | Existing meta-learning techniques may not be well-suited to testing tasks when they are not well-supported by training tasks. |
| Approach: | They propose to use a meta-learning algorithm to add representations for each sentence learned from the extracted sentence-specific knowledge graph. |
| Outcome: | The proposed model is able to represent each sentence learned from the extracted knowledge graph under supervised adaptation and unsupervised adaptation settings. |
Copied to clipboard
| Challenge: | Existing methods for pretraining a language model on text have been used for building models in NLP, but they do not work for sentence representations derived from pretrainer models based on tokens or basic pooling operations. |
| Approach: | They propose to build a sentence-level autoencoder from a pretrained transformer language model. |
| Outcome: | The proposed model achieves better quality than previous methods on text similarity and style transfer tasks while using fewer parameters than large pretrained models. |
Copied to clipboard
| Challenge: | Recent studies describe how to apply contrastive learning to the language domain but it is difficult to apply data augmentation methods directly to language modeling. |
| Approach: | They propose a memory-efficient continual pretraining method that applies contrastive learning with novel data augmentation and curriculum learning. |
| Outcome: | The proposed method outperforms baseline models on sentence-level tasks with only 70% of memory compared to the baseline model. |
Copied to clipboard
| Challenge: | Existing systems that explore user preference through conversational interactions do not exploit the context and knowledge to make accurate recommendations. |
| Approach: | They propose a model that performs tree-structured reasoning on a knowledge graph and generates informative dialog acts to guide language generation. |
| Outcome: | The proposed model can arrive at more accurate recommendation and generate more informative and engaging responses. |
Copied to clipboard
| Challenge: | Existing knowledge grounding models focus on locating knowledge in document contexts that are relevant to the conversation. |
| Approach: | They propose a knowledge identification model that leverages document structure to provide dialogue-contextualized passage encodings and better locate knowledge relevant to the conversation. |
| Outcome: | The proposed model can be applied to document-grounded conversational datasets and shows generalization to unseen documents and long dialogue contexts. |
Copied to clipboard
| Challenge: | Communicating with humans is challenging for AIs because of its complexity and multimodality. |
| Approach: | They propose to use a game of drawing and guessing based on Pictionary to test AIs' understanding of the world and multi-modal gestures. |
| Outcome: | The proposed game is a test for mixing language and visual/symbolic communication in AI. |
Copied to clipboard
| Challenge: | Large-scale pre-trained language models have shown promising results for few-shot learning in task-oriented dialog (ToD) systems. |
| Approach: | They propose a self-training approach that iteratively labels the most confident unlabeled data to train a stronger Student model. |
| Outcome: | The proposed approach improves state-of-the-art pre-trained models in few-shot learning scenarios for task-oriented dialog (ToD) systems when only a small number of labeled data are available. |
Copied to clipboard
| Challenge: | Large-scale conversational AI based dialogue systems like Alexa, Siri, and Google Assistant, are getting more and more prevalent in real-world applications to help users across the globe. |
| Approach: | They propose a contextual rephrase detection model ContReph to automatically identify rephrasings from multi-turn dialogues using contextual information and user-agent interaction signals. |
| Outcome: | The proposed model outperforms the pairwise rephrase detection models by leveraging the context and user-agent interaction signals. |
Copied to clipboard
| Challenge: | Existing methods address few-shot intent detection tasks from two perspectives: data augmentation and task-adaptive training with pre-trained models. |
| Approach: | They propose a few-shot intent detection schema using contrastive pre-training and fine-tuning. |
| Outcome: | The proposed method achieves state-of-the-art performance on three challenging intent detection datasets under 5-shot and 10-shot settings. |
Copied to clipboard
| Challenge: | Conversational recommendation systems (CRSs) aim to refine options over multiple turns of a conversation, but they are not as flexible as real conversations. |
| Approach: | They propose a method for transforming a user critique into a positive preference . they use a large neural language model to perform critique-to-preference transformation . |
| Outcome: | The proposed method improves recommendations in restaurant domain using a new dataset of restaurant critiques. |
Copied to clipboard
| Challenge: | Keyword or keyphrase extraction is to identify words or phrases presenting the main topics of a document. |
| Approach: | They propose a hybrid attention model to identify keyphrases from a document in an unsupervised manner. |
| Outcome: | The proposed model is effective and robust on long and short documents. |
Copied to clipboard
| Challenge: | Existing methods for relation extraction use latent variables and supervised training which requires large datasets. |
| Approach: | They propose a VAE-based unsupervised relation extraction technique that uses latent variables as an intermediate variable instead of a latent variable. |
| Outcome: | The proposed method outperforms state-of-the-art methods on the NYT dataset and outperformed existing methods. |
Copied to clipboard
| Challenge: | Automating high quality knowledge graphs from a given collection of documents remains a challenging problem in AI. |
| Approach: | They propose a novel approach to slot filling that extends dense passage retrieval with hard negatives and robust training procedures for retrieval augmented generation models. |
| Outcome: | The proposed model improves on both T-REx and zsRE slot filling datasets and ranks at the top-1 position in the KILT leaderboard. |
Copied to clipboard
| Challenge: | Zero-shot cross-lingual information extraction (IE) is a technique for training data in a source language but not in . |
| Approach: | They explore techniques including data projection and self-training to improve zero-shot cross-lingual information extraction (IE) IE is a construction of an IE model for some target language given existing annotations exclusively in English. |
| Outcome: | The proposed techniques show that they perform better than any single strategy. |
Copied to clipboard
| Challenge: | Recent work analyzes, quantifies, and mitigates language model biases such as gender, race or religion-related stereotypes in static word embeddings and contextual representations. |
| Approach: | They explain the complexity of gender and language around it and examine how current representations perpetuate harms associated with binary gender. |
| Outcome: | The proposed model and dataset biases perpetuate harms associated with the treatment of gender as binary in English language technologies. |
Copied to clipboard
| Challenge: | Extensive experiments on MS-COCO and Flickr30K benchmarks show that our methods significantly reduce the gender bias in image search models. |
| Approach: | They propose a fair sampling method and a feature clipping method to debias image search models. |
| Outcome: | The proposed methods significantly reduce gender bias in image search models on MS-COCO and Flickr30K benchmarks. |
Copied to clipboard
| Challenge: | Text style can reveal sensitive attributes of the author (e.g. age and race) to the reader, which can lead to privacy violations and bias in both human and algorithmic decisions based on text. |
| Approach: | They propose a framework that obfuscates stylistic features of human-generated text through style transfer by automatically re-writing the text itself. |
| Outcome: | The proposed framework obfuscates stylistic features of human-generated text through style transfer, by automatically re-writing the text itself. |
Copied to clipboard
| Challenge: | Broader disclosive transparency is difficult to define and quantify, authors say . previous work has demonstrated trade-offs and negative consequences to disclosing transparency . |
| Approach: | They propose to use neural language model-based probabilistic metrics to model disclosive transparency . they demonstrate that they correlate with user and expert opinions of system transparency a valid objective proxy . |
| Outcome: | The proposed metrics correlate with user and expert opinions of system transparency, making them a valid objective proxy. |
Copied to clipboard
| Challenge: | Existing private learning schemes which protect data privacy can be used to train models using instance encoding. |
| Approach: | They propose to recover the private training data and use it to break a private learning scheme TextHide. |
| Outcome: | The proposed attack would advance privacy-preserving machine learning in the context of natural language processing. |
Copied to clipboard
| Challenge: | Existing studies on class imbalance and mitigating bias have focused on the latter . a skewed class distribution hurts the performance of deep learning models, and is often referred to as "stereotyping" |
| Approach: | They propose to extend a margin-loss based approach to enforce fairness by using tweet sentiment and occupation classification to mitigate class imbalance and demographic bias. |
| Outcome: | The proposed methods help mitigate class imbalance and demographic biases through controlled experiments. |
Copied to clipboard
| Challenge: | Existing approaches to combine HE and GC in RNNs suffer from long inference latency due to the slow activation functions. |
| Approach: | They propose a hybrid structure of HE and GC gated recurrent unit network, for low-latency secure inferences. |
| Outcome: | The proposed structure improves the secure inference latency by up to 138 over one of the state-of-the-art secure networks on the Penn Treebank dataset. |
Copied to clipboard
| Challenge: | a new computational task supports the construction of high quality texts and lexicons for low resource languages. |
| Approach: | They propose a computational task which is tuned to the available knowledge and interests in an Indigenous community. |
| Outcome: | The proposed method achieves a transcription density gain of 17% in a morphologically complex language . the proposed grammar includes a description of the phonology and morphosyntax . |
Copied to clipboard
| Challenge: | Recent state-of-the-art (SOTA) effective neural network methods have been used in Chinese word segmentation (CWS) However, the robustness of the previous neural methods is limited by the large-scale annotated corpus. |
| Approach: | They propose a self-supervised Chinese word segmentation approach with a straightforward and effective architecture. |
| Outcome: | The proposed approach outperforms previous methods on 9 different CWS datasets with single criterion training and multiple criteria training and achieves better robustness. |
Copied to clipboard
| Challenge: | Neural models for morphological reinflection tasks have proved to be extremely accurate given ample labeled data, yet labele d data may be slow and costly to obtain. |
| Approach: | They exploit orthographic and semantic regularities in morphological systems to exploit the orthographic regularities on their own to achieve respectable accuracy. |
| Outcome: | The bootstrapping method outperforms hallucination-based methods for morphological reinflection tasks. |
Copied to clipboard
| Challenge: | Existing methods for tokenization of text are not efficient, but they are based on Aho-Corasick's algorithm. |
| Approach: | They propose an efficient algorithm for WordPiece tokenization using a longest-match-first strategy . they propose an algorithm whose tokenization complexity is strictly O(n) |
| Outcome: | The proposed method is 8.2x faster than HuggingFace Tokenizers and 5.1x faster on average for general text tokenization. |
Copied to clipboard
| Challenge: | Neural language models typically tokenise input text into sub-word units to achieve an open vocabulary. |
| Approach: | They propose that language models should be evaluated on their marginal likelihood over tokenisations instead. |
| Outcome: | The proposed approach is unsatisfactory and may bottleneck model out-of-domain performance. |
Copied to clipboard
| Challenge: | Generally, commonsense knowledge is correlated with culture and geographic locations and is only shared locally. |
| Approach: | They construct a Geo-Diverse Visual Commonsense Reasoning dataset to test vision-and-language models’ ability to understand cultural and geo-location-specific commonsense. |
| Outcome: | The proposed models perform better in non-Western regions including East Asia, South Asia, and Africa than in the Western regions. |
Copied to clipboard
| Challenge: | Using a structured referent grounding module, we can effectively ground and inform a partner's utterances to their own context. |
| Approach: | They propose a grounded neural dialogue model that works with people in a partially-observable reference game. |
| Outcome: | The proposed model outperforms state-of-the-art models on a spatial grounding dialogue task and achieves a 20% relative improvement in human evaluations. |
Copied to clipboard
| Challenge: | Existing visual question answering models leverage spurious biases and take shortcuts to improve performance. |
| Approach: | They propose a semi-automatic framework for generating disentangled shifts by introducing a controllable visual question-answer generation module that generates highly-relevant question-announcer pairs with the desired dataset style. |
| Outcome: | The proposed framework generates highly-relevant and diverse question-answer pairs with the desired dataset style. |
Copied to clipboard
| Challenge: | Past work in NLP examined the task of goal-step inference for textual goals . wikiHow dataset shows that goal-step inference is challenging for state-of-the-art models . |
| Approach: | They propose a task where a model is given a textual goal and must choose which of four images represents a plausible step towards that goal. |
| Outcome: | The proposed task is challenging for state-of-the-art multimodal models and can be transferred to other datasets. |
Copied to clipboard
| Challenge: | a general-purpose Transformer-based model with crossmodal attention solves most of the systematic generalization problems . current models are data inefficient given the narrow scope of commands in gSCAN . |
| Approach: | They propose to use a Transformer-based model with cross-modal attention to solve gSCAN . they propose to generate data to incorporate relations between objects in the visual environment . |
| Outcome: | The proposed model outperforms specialized approaches on most splits, and is data inefficient given the narrow scope of commands. |
Copied to clipboard
| Challenge: | Existing methods for creating vision-and-language models involve structural modifications and V&L pre-training. |
| Approach: | They propose to extend a language model through structural modifications and V&L pre-training to make it inherit the capability of natural language understanding from the original language model. |
| Outcome: | The proposed method improves performance of vision-and-language models by extending pre-trained models with the same pre-training. |
Copied to clipboard
| Challenge: | Dialogue systems that generate factually incorrect responses are often unfitful and hallucinate factuality invalid. |
| Approach: | They propose a method to improve faithfulness and reduce hallucination of neural dialogue systems to known facts supplied by a Knowledge Graph. |
| Outcome: | The proposed approach improves faithfulness and reduces hallucination of dialogue systems to known facts . it leverages a token-level fact critic to identify plausible sources of hallucinism . |
Copied to clipboard
| Challenge: | Existing models with seq2seq framework lack ability to effectively manage concept transitions . lack of concept management strategies might lead to incoherent dialogue due to loosely connected concepts . |
| Approach: | They propose a concept-guided non-autoregressive model for open-domain dialogue generation that learns to identify multiple associated concepts from a conceptual graph and a customized Insertion Transformer to perform concept-directed generation to complete a response. |
| Outcome: | The proposed model outperforms state-of-the-art models in automatic and human evaluations with substantially faster inference speed. |
Copied to clipboard
| Challenge: | Empathy is a complex cognitive ability based on the reasoning of others’ affective states. |
| Approach: | They propose a method to infer emotion cause words from utterances without a word-level label and a novel method to make dialogue models focus on targeted words in the input during generation. |
| Outcome: | The proposed method improves multiple best-performing dialogue agents on generating more focused empathetic responses in terms of automatic and human evaluation. |
Copied to clipboard
| Challenge: | Current models are not satisfactory for solving out-of-vocabulary problems . current models assume that the task ontology is well defined in advance . |
| Approach: | They propose to enhance the interrelation between slots with masked hierarchical attention. |
| Outcome: | The proposed model yields a significant performance gain over current state-of-the-art model and is more robust to out-ofvocabulary problem compared with other methods. |
Copied to clipboard
| Challenge: | Existing approaches to knowledge-grounded dialogue generation perform relatively independent sub-tasks . Typical approaches tend to decompose this task into two streamlined sub- tasks . |
| Approach: | They propose a collaborative latent variable model to integrate knowledge selection and knowledge-aware response generation simultaneously in separate but collaborative latences. |
| Outcome: | The proposed model outperforms previous methods on knowledge selection and response generation. |
Copied to clipboard
| Challenge: | Existing knowledge-grounded dialogues perform poorly when transfer into new domains with limited training samples. |
| Approach: | They propose a weakly supervised three-stage learning framework based on weakly-supervised learning based upon large scale ungrounded dialogues and unstructured knowledge base. |
| Outcome: | The proposed framework outperforms state-of-the-art methods even in zero-resource setting. |
Copied to clipboard
| Challenge: | Recent years has witnessed the remarkable success in end-to-end task-oriented dialog system, especially when incorporating external knowledge information. |
| Approach: | They propose a mechanism to model deterministic entity knowledge by using an intention reasoning network to obtain intention-aware representations of conceptual tokens. |
| Outcome: | The proposed mechanism captures concept shifts and generates accurate responses on two representative multi-domain dialog datasets. |
Copied to clipboard
| Challenge: | Existing knowledge-enhanced methods use a single-source homogeneous knowledge base with limited knowledge coverage. |
| Approach: | They propose a multi-source heterogeneous knowledge-enhanced dialogue generation model that leverages multiple knowledge sources to improve knowledge coverage. |
| Outcome: | The proposed model outperforms existing knowledge-enhanced models on a Chinese dataset and shows that it can leverage multiple heterogeneous knowledge sources to improve knowledge coverage. |
Copied to clipboard
| Challenge: | Existing offline DST models require a fixed dataset to train . Existing domain-lifelong learning methods are impractical in real-world applications . |
| Approach: | They propose a domain-lifelong learning method to continuously train a DST model on new data to learn incessantly emerging new domains while avoiding catastrophically forgetting old learned domains. |
| Outcome: | The proposed method outperforms state-of-the-art lifelong learning methods by 4.25% and 8.27% on the MultiWOZ and the SGD benchmarks. |
Copied to clipboard
| Challenge: | Existing CSRL parsers struggle to handle conversational structural information. |
| Approach: | They propose a conversational semantic role labeling task which explicitly encodes speaker dependent information and proposes a multi-task learning method to further improve the model. |
| Outcome: | The proposed model outperforms baselines on benchmark datasets on conversation-based tasks. |
Copied to clipboard
| Challenge: | Pre-trained models can be fine-tuned on domain-specific unlabeled data . however, most further pre-training works just keep running the conventional pre- training task . |
| Approach: | They propose to add a further pre-training phase to the model to improve downstream tasks . they propose to use a domain-adaptive pre-tuning phase to fine-tune the models on unlabeled data . |
| Outcome: | The proposed method improves multiple task-oriented dialogue downstream tasks. |
Copied to clipboard
| Challenge: | Existing methods for dialogue generation use an external knowledge base to generate appropriate responses. |
| Approach: | They propose to use an external knowledge base to generate appropriate responses for unseen entities. |
| Outcome: | Experiments on two dialogue corpus show that pre-trained models perform poorly with unseen entities. |
Copied to clipboard
| Challenge: | Multi-turn response selection models have shown comparable performance to humans in several benchmark datasets, but in the real environment, they often have weaknesses, such as giving the highest score to the wrong response candidate containing several keywords related to the context. |
| Approach: | They propose to build a robust multi-turn response selection model in an adversarial environment and to use it to evaluate weaknesses. |
| Outcome: | The proposed model makes incorrect predictions based heavily on superficial patterns without a comprehensive understanding of the context. |
Copied to clipboard
| Challenge: | Existing work on conversation disentanglement relies heavily on human annotations, which is expensive to obtain in practice. |
| Approach: | They propose to train a conversation disentanglement model without referencing human annotations . they use a message-pair classifier and a session classifier to retrieve local relations . |
| Outcome: | The proposed method achieves competitive performance compared to previous methods on a large movie dialogue dataset. |
Copied to clipboard
| Challenge: | Consistency Identification has been used for preventing inconsistent response generation, but few efforts have been made to task-oriented dialogue. |
| Approach: | They propose a dataset for Consistency Identification in task-oriented dialog system. |
| Outcome: | The proposed dataset is based on a single label and provides fine-grained labels to encourage model to know what inconsistent sources lead to it. |
Copied to clipboard
| Challenge: | Existing grounded dialogue models are limited by the distribution of data and the type of grounded concepts. |
| Approach: | They propose a framework which edits existing responses to be grounded on a given concept by disentangling and recombining persona-related and persona agnostic parts of the response. |
| Outcome: | The proposed framework outperforms baselines on the personaMi-nEdit dataset and shows that it can improve persona consistency while preserving the use of knowledge and empathy. |
Copied to clipboard
| Challenge: | Existing knowledge-grounded conversation models lack knowledge that occurs in training data, resulting in incomplete knowledge generation. |
| Approach: | They propose an Entity-Agnostic Representation Learning method to introduce knowledge graphs to informative conversation generation using context of conversations and relational structure of knowledge graph. |
| Outcome: | The proposed model generates more informative, coherent, and natural responses than baseline models. |
Copied to clipboard
| Challenge: | Conventional approaches to learning sentence embeddings from dialogues employ the siamese-network for this task, but such architecture yields a large gap between training and evaluating. |
| Approach: | They propose a dialogue-based contrastive learning approach to learn sentence embeddings from dialogues using a siamese-network. |
| Outcome: | The proposed model outperforms baseline methods on three multi-turn dialogue datasets in terms of MAP and Spearman’s correlation measures. |
Copied to clipboard
| Challenge: | Existing sentence ordering models can be classified into pairwise ordering models and set-to-sequence models. |
| Approach: | They propose a novel sentence ordering framework which introduces two classifiers to make better use of pairwise orderings for graph-based sentence ordering. |
| Outcome: | The proposed model achieves state-of-the-art performance on five commonly-used datasets. |
Copied to clipboard
| Challenge: | Existing methods of implicit discourse relation recognition (IDRR) focus on three aspects: enhancing discourse units representation, enhancing semantic interaction, and joint learning with other tasks. |
| Approach: | They propose a joint model to recognize the relation label and generate the target sentence containing the meaning of relations simultaneously. |
| Outcome: | The proposed model achieves the best performance against several state-of-the-art systems on Chinese and English datasets. |
Copied to clipboard
| Challenge: | Existing methods to consider textual coherence are limited in labeled data. |
| Approach: | They propose a language model-based generative classifier that uses labels as input and embeds labels into their representations. |
| Outcome: | The proposed classifier achieves state-of-the-art in discourse segmentation and relation F1 scores with gold boundaries and automatically segmented boundaries. |
Copied to clipboard
| Challenge: | Existing methods to model multimodal sentiment analysis are limited due to their complexity and memory footprint. |
| Approach: | They propose a multimodal Sparse Phased Transformer to reduce self-attention complexity and memory footprint. |
| Outcome: | The proposed method achieves comparable or superior performance with a 90% reduction in the number of parameters. |
Copied to clipboard
| Challenge: | Existing approaches to hierarchical multi-label text classification ignore vertical category correlations or exploit dependencies across levels without considering horizontal correlations . |
| Approach: | They propose a hierarchical multi-label text classification framework that considers both vertical and horizontal category correlations. |
| Outcome: | The proposed framework improves on real-world HMTC datasets with significant improvements over baselines. |
Copied to clipboard
| Challenge: | Existing methods require training millions of architectures to estimate the accuracy of the search results. |
| Approach: | They propose a performance ranking method (RankNAS) that uses pairwise ranking and search space pruning to enlarge the search space. |
| Outcome: | The proposed method significantly accelerates NAS through pairwise ranking and search space pruning. |
Copied to clipboard
| Challenge: | obtaining large amounts of labeled data is expensive. |
| Approach: | They develop a semi-supervised learning framework called FLiText which improves text classification accuracy. |
| Outcome: | The proposed framework improves accuracy of lightweight models on IMDb, Yelp-5, and Yahoo! Answer . the framework improve accuracy by 6.59%, 3.94%, and 3.22% on the datasets of IMDa, Yep-5 and Yahoo. Answer compared with the fully supervised method on the full dataset . |
Copied to clipboard
| Challenge: | Existing methods for debiasing protected attributes have been limited to binary attributes in isolation, however many corpora involve multiple such attributes, possibly with higher cardinality. |
| Approach: | They propose to evaluate a bias-constrained model which is new to NLP and an extension of the iterative nullspace projection technique which can handle multiple identities. |
| Outcome: | The proposed model is based on a new iterative nullspace projection technique which can handle multiple identities. |
Copied to clipboard
| Challenge: | Existing definition generation techniques have faced various problems such as the out-of-vocabulary problem and over/under-specificity problems. |
| Approach: | They propose to leverage a pre-trained encoder-decoder model and introduce a re-ranking mechanism to model specificity in definitions. |
| Outcome: | The proposed method significantly outperforms the state-of-the-art method on standard evaluation datasets and shows that it addresses the over/under-specificity problems. |
Copied to clipboard
| Challenge: | Existing methods for style transfer are based on an inductive learning approach, which represents the style as embeddings, decoder parameters, or discriminator parameters and directly applies these general rules to the test cases. |
| Approach: | They propose a retrieval-based context-aware style representation that involves top-K relevant sentences in the target style in the transfer process. |
| Outcome: | The proposed method outperforms several strong baselines and is general and effective to the task of unsupervised style transfer. |
Copied to clipboard
| Challenge: | Existing graph-based methods only consider word relations or structure information, which neglect the correlation between them. |
| Approach: | They propose a Dual Graph network for Abstractive Sentence Summarization that captures word relations and structure information from sentences. |
| Outcome: | The proposed model outperforms state-of-the-art methods on two popular benchmark datasets. |
Copied to clipboard
| Challenge: | ZP-annotated natural language generation (NLG) corpora are scarce in pro-drop languages . despite efforts to bridge the discrepancy between human and machine, zero pronouns still persist in pro -drop tasks. |
| Approach: | They propose a highly adaptive two-stage approach to couple context modeling with ZP recovering to mitigate the ZP problem in NLG tasks. |
| Outcome: | The proposed approach can improve translation, question answering, and summarization tasks. |
Copied to clipboard
| Challenge: | Experimental results show that our model can achieve a significant improvement in terms of metric-based evaluation and human evaluation compared with the state-of-the-art exposure bias approaches. |
| Approach: | They propose a novel adaptive switching mechanism which automatically transits between ground-truth learning and generated learning regarding the word-level matching score. |
| Outcome: | The proposed model improves on Chinese and English reddit datasets compared with state-of-the-art models on the word-level matching score. |
Copied to clipboard
| Challenge: | Existing methods for paraphrase generation lack reliable supervision signals. |
| Approach: | They propose an unsupervised paradigm for paraphrase generation based on contextual language models, candidate filtering and paraphrase model training based upon the selected candidates. |
| Outcome: | The proposed paradigm outperforms existing paraphrase generation methods in supervised and unsupervised setups. |
Copied to clipboard
| Challenge: | Existing methods for conditional long text generation ignore the coherence issue of the generated texts. |
| Approach: | They propose a two-stage approach to generate coherent long text based on short input text . they first build a document-level path for each output text with each sentence embedding as its node . |
| Outcome: | The proposed approach is superior to state-of-the-art approaches on three real-world datasets. |
Copied to clipboard
| Challenge: | Existing models ignore the rich structure information that is hidden in the previously generated text. |
| Approach: | They propose to model the previous generation using a Graph Neural Network at each decoding step. |
| Outcome: | The proposed model outperforms the state-of-the-art models with sentence-level QG tasks on SQUAD and MARCO datasets. |
Copied to clipboard
| Challenge: | Existing approaches to generate high quality question-answer pairs are limited . a new framework is proposed for the question-answer generation task on real-world examination data. |
| Approach: | They propose a multi-agent communication model to generate and optimize the question and keyphrases iteratively and then apply the generated question and keys to guide the generation of answers. |
| Outcome: | The proposed framework makes great breakthroughs in the question-answer pair generation task. |
Copied to clipboard
| Challenge: | Existing studies on syntactically controlled paraphrase generation rely on large-scale parallel data. |
| Approach: | They propose a syntactically-informed unsupervised paraphrasing model based on conditional variational auto-encoder which can generate texts in a specified syntastic structure. |
| Outcome: | The proposed model can generate diverse paraphrases with specified syntactic structure using non-parallel data. |
Copied to clipboard
| Challenge: | Existing models do not distinguish hard tasks from easy ones in the learning process. |
| Approach: | They propose a novel approach that exploits relation label information to learn better representations by focusing on hard tasks. |
| Outcome: | Experiments on two standard datasets show the proposed approach performs better than previous methods. |
Copied to clipboard
| Challenge: | Recent advances in entity retrieval ignore the property that meanings of entity mentions diverge in different contexts and are related to various portions of descriptions. |
| Approach: | They propose a novel approach that constructs multi-view representations for entity descriptions and approximates the optimal view for mentions via a heuristic searching method. |
| Outcome: | The proposed approach achieves state-of-the-art performance on ZESHEL and improves quality of candidates on three standard Entity Linking datasets. |
Copied to clipboard
| Challenge: | Existing neural-based ED models are confused by changeable contexts during testing . we propose a system that extracts statistical event features from word-event cooccurrence frequencies . |
| Approach: | They propose to integrate a set of statistical event features from word-event co-occurrence frequencies into the training set to cooperate with contextual features. |
| Outcome: | The proposed model outperforms ten strong baselines on ACE2005 and KBP2015 datasets. |
Copied to clipboard
| Challenge: | Existing studies focus on identifying event factuality at sentence level, which leads to conflicts between different mentions of the same event. |
| Approach: | They propose a document-level event factuality identification model that uses local uncertainty and global structure to model event factuality. |
| Outcome: | The proposed method outperforms existing models on two widely used datasets. |
Copied to clipboard
| Challenge: | Table filling based relational triple extraction methods focus on using local features but ignore the global associations of relations and token pairs, which increases the possibility of overlooking some important information during triple extraction. |
| Approach: | They propose a global feature-oriented triple extraction model that makes full use of the two kinds of global associations of relations and token pairs. |
| Outcome: | The proposed model achieves state-of-the-art on three benchmark datasets. |
Copied to clipboard
| Challenge: | Creating keyphrases that are likely to be words absent from the given document is challenging . |
| Approach: | They propose novel keyphrase generation tasks that augment missing context by adding keyphrases to documents. |
| Outcome: | The proposed keyphrase generation task outperforms the state-of-the-art in two keyphrase tasks. |
Copied to clipboard
| Challenge: | Auxiliary information from multiple sources has been demonstrated to be effective in zero-shot fine-grained entity typing (ZFET) however, there is no comprehensive understanding of how to make better use of the existing information sources and how they affect the performance of ZFET. |
| Approach: | They propose a multi-source fusion model targeting auxiliary information from multiple sources to improve zero-shot fine-grained entity typing (ZFET) |
| Outcome: | The proposed model achieves 11.42% and 22.84% gains over state-of-the-art baselines on BBN and Wiki respectively with regard to macro F1 scores. |
Copied to clipboard
| Challenge: | Existing approaches to integrate lexical knowledge into deep learning models are limited by large-scale dynamic lexicons. |
| Approach: | They propose a plug-in lexicon incorporation approach for BERT based sequence labeling tasks . they adopt word-agnostic tag embeddings to avoid re-training the representation . |
| Outcome: | The proposed framework achieves new SOTA even with large scale lexicons, the authors show . they adopt word-agnostic tag embeddings to avoid re-training the representation . |
Copied to clipboard
| Challenge: | Neural relation extraction models have shown promising results on long-tail tasks, but performance drops dramatically as the number of instances for a relation decreases. |
| Approach: | They propose a framework considering both label-agnostic and label-aligned mapping information for low resource relation extraction. |
| Outcome: | The proposed framework improves on low-resource relation extraction tasks by incorporating label-agnostic and label-based mapping information in pretraining and fine-tuning. |
Copied to clipboard
| Challenge: | Existing approaches for keyphrase generation generate uncontrollable and inaccurate absent keyphrases. |
| Approach: | They propose a graph-based method that captures explicit knowledge from related references. |
| Outcome: | The proposed model improves on baseline keyphrase generation models on multiple benchmarks. |
Copied to clipboard
| Challenge: | Existing datasets are too small to train a model for capturing regularities underlying how event arguments are extracted. |
| Approach: | They propose to bridge implicit EAE with machine reading comprehension (MRC) by building a unified training framework and explicit data augmentation regimes via MRC. |
| Outcome: | The proposed method obtains state-of-the-art performance on two benchmarks and demonstrates superior results in a data-low scenario. |
Copied to clipboard
| Challenge: | Existing keyphrase extraction methods focus on the part of phrase that is important . experimental results show that KIEMP outperforms existing keyphrase extracting methods . |
| Approach: | They propose to estimate the importance of keyphrase from multiple perspectives using a chunking module, ranking module and matching module. |
| Outcome: | The proposed method outperforms the state-of-the-art keyphrase extraction methods on six benchmark datasets. |
Copied to clipboard
| Challenge: | Existing methods to extract relation facts from limited labeled corpora are laborintensive to obtain . Existing approaches use self-training to generate pseudo labels that will cause gradual drift problem or leverage meta-learning scheme which does not solicit feedback explicitly. |
| Approach: | They propose a Gradient Imitation Reinforcement Learning method to encourage pseudo label data to imitate gradient descent direction on labeled data and bootstrap its optimization capability through trial and error. |
| Outcome: | The proposed method handles two major scenarios in low-resource relation extraction when no unlabeled data is available. |
Copied to clipboard
| Challenge: | Taxonomies represent hierarchical relationships between terms or entities. |
| Approach: | They propose a framework for taxonomy enrichment in low-resource settings with pretrained language models as knowledge bases to compensate for the shortage of information. |
| Outcome: | The proposed framework predicts whether inputted term pairs have hierarchical relationships and leverages implicit knowledge from the LM to generate queries efficiently. |
Copied to clipboard
| Challenge: | Existing studies on key information extraction from visually rich documents focus on labeling the text within bounding boxes, while relations between words are unexplored. |
| Approach: | They propose to use a dependency parsing model to extract semantic entities from visually rich documents by combining entity labeling and relation extraction tasks. |
| Outcome: | The proposed model achieves 65.96% F1 score on the FUNSD dataset. |
Copied to clipboard
| Challenge: | Existing studies on joint entity and relation extraction fail to fully utilize the interdependence between entity types and relation types. |
| Approach: | They propose a synchronous dual network with cross-type attention via separately and interactively considering the entity types and relation types. |
| Outcome: | The proposed model achieves state-of-the-art on NYT and WebNLG datasets. |
Copied to clipboard
| Challenge: | Dense retrieval requires high-quality text sequence embeddings to support effective search in the representation space. |
| Approach: | They propose a self-learning method that pre-trains the autoencoder using a weak decoder to push the encoder to provide better sequence representations. |
| Outcome: | The proposed model significantly boosts the effectiveness and few-shot ability of dense retrieval models on web search, news recommendation, and open domain question answering. |
Copied to clipboard
| Challenge: | Recent studies show that prompts improve performance of large pre-trained language models for few-shot text classification. |
| Approach: | They propose a prompt-based framework for few-shot learning that captures cross-task transferable knowledge and uses two de-biasing techniques to make it more task-agnostic and unbiased . |
| Outcome: | The proposed framework outperforms strong baselines over multiple NLP tasks and datasets. |
Copied to clipboard
| Challenge: | Existing methods for text classification ignore keyword correlation, thus ignoring it . existing methods treat keywords independently, thus not exploiting correlation between them . |
| Approach: | They propose a framework to explore keyword-keyword correlation on keyword graph by GNN . they use a self-supervised task to pretrain annotators and fine-tune them . |
| Outcome: | The proposed method outperforms existing methods on long- and short-text datasets. |
Copied to clipboard
| Challenge: | Existing news recommendation methods rely on centralized storage of user click behavior data, which may lead to privacy concerns and hazards. |
| Approach: | They propose a federated learning framework for privacy-preserving news recommendation . they propose aggregation of news representations and user model by a client . |
| Outcome: | The proposed framework reduces computation and communication cost on clients while keeping promising model performance. |
Copied to clipboard
| Challenge: | Recent studies show that passage retrieval and passage reranking are important for achieving mutual improvement. |
| Approach: | They propose a unified listwise training approach for passage retrieval and passage reranking that incorporates a retrieval procedure and a hybrid data augmentation strategy. |
| Outcome: | The proposed approach improves on both MSMARCO and Natural Questions datasets. |
Copied to clipboard
| Challenge: | Current approaches to passage retrieval and ranking rely on pre-trained deep language models that model the semantic matching between queries and passages. |
| Approach: | They propose a typos-aware training framework for DR and BERT to address this issue. |
| Outcome: | The proposed models respond and adapt to keyword typos occurring in queries, and significantly improve their retrieval and ranking effectiveness. |
Copied to clipboard
| Challenge: | Existing methods for cross-lingual entity alignment rely on lexical matching and probability reasoning, but they inherit poor interpretability and low efficiency from neural networks. |
| Approach: | They propose a simple but effective unsupervised entity alignment method without neural networks that can be used to find the equivalent entities between crosslingual KGs. |
| Outcome: | Extensive experiments show that the proposed method beats advanced supervised methods across all datasets while having high efficiency, interpretability, and stability. |
Copied to clipboard
| Challenge: | Dense passage retrieval improves ranking accuracy in open-domain question answering but at the cost of large space and memory requirements. |
| Approach: | They propose a simple unsupervised pipeline that includes principal component analysis (PCA), product quantization, and hybrid search to improve space efficiency. |
| Outcome: | The proposed pipeline achieves good accuracy–space trade-offs, for example, 48 compression with less than 3% drop in top-100 retrieval accuracy on average or 96 compression without drop in space requirements. |
Copied to clipboard
| Challenge: | Recent studies for relation extraction (RE) leverage the dependency tree of the input sentence to improve performance. |
| Approach: | They propose to use a graph convolutional network to build a context graph without dependency parsers. |
| Outcome: | The proposed approach improves neural RE methods without dependency parsers on English benchmark datasets. |
Copied to clipboard
| Challenge: | a recent paper suggests that probing should be seen as approximating a mutual information. |
| Approach: | They propose a Bayesian mutual information framework that probes probing representations from the perspective of Bayes' agents. |
| Outcome: | The proposed framework allows for more intuitive results in scenarios with finite data. |
Copied to clipboard
| Challenge: | masked language models (MLMs) pre-train to model higher-order word co-occurrence statistics . authors suggest that such models have learned to represent syntactic structures prevalent in classical NLP pipelines . purely distributional information largely explains the success of pre-training, authors say . |
| Approach: | They propose to pre-train masked language models on sentences with random shuffled word order and show they still achieve high accuracy after fine-tuning on many downstream tasks. |
| Outcome: | The proposed model performs well according to parametric syntactic probes . the authors argue that the model is not all that different from earlier distributional models . |
Copied to clipboard
| Challenge: | Existing subnetworks of one-layer randomly weighted neural networks can achieve impressive performance without changing initializations. |
| Approach: | They find subnetworks within one-layer randomly weighted neural networks that can achieve impressive performance without ever modifying the initializations. |
| Outcome: | The proposed subnetworks match 98%/92% of the performance of a trained Transformersmall/base on IWSLT14/WMT14. |
Copied to clipboard
| Challenge: | Pre-trained models such as BERT have achieved success in learning sequence representations, but they tend to learn representations that are covariant with the noise of pre-training. |
| Approach: | They propose to train self-trained models to learn noise invariant sequence representations . they encourage consistency between original sequence and corrupted version via unsupervised instance-wise training signals. |
| Outcome: | The proposed model improves on 11 natural language understanding and cross-modal tasks and achieves 0.6% gain on GLUE benchmarks and 0.8% increment on NLVR2 . |
Copied to clipboard
| Challenge: | Existing explanation methods are inefficient when explaining a static black-box model. |
| Approach: | They propose a Lifelong Explanation approach that continuously trains a student explainer under the supervision of a teacher on different tasks undertaken in LL. |
| Outcome: | The proposed approach can be extended to include a teacher and maintain the same level of faithfulness to the black-box model as the student explainer while being up to 102 times faster at test time. |
Copied to clipboard
| Challenge: | Existing studies show that pairs of words that occur together are likely to stand in a linguistic dependency. |
| Approach: | They propose to use large pretrained language models to compute contextualized estimates of the pointwise mutual information between words (CPMI) |
| Outcome: | The proposed model maximizes CPMI and achieves an unlabelled undirected attachment score of 0.5 . |
Copied to clipboard
| Challenge: | Existing literature is agnostic about a parsing strategy of hierarchical models . a recent study showed that hierarchically model hierarchic structures capture grammatical dependencies much better than RNNs in targeted syntactic evaluations. |
| Approach: | They evaluated three LMs with head-final left-branching structures and Recurrent Neural Network Grammars with top-down and left-corner parsing strategies as hierarchical models. |
| Outcome: | The proposed model outperforms top-down and left-corner models against human reading times in Japanese. |
Copied to clipboard
| Challenge: | Recent studies suggest that relative position encodings provide better performance than absolute position coding. |
| Approach: | They propose a mechanism to encode position and segment information into Transformer models using relative position encodings. |
| Outcome: | The proposed method achieves faster training and inference time while achieving competitive performance on GLUE, XTREME and WMT benchmarks. |
Copied to clipboard
| Challenge: | Experimental results on nine authoritative datasets demonstrate the effectiveness of our methods empirically. |
| Approach: | They propose to use a method to explicitly encode position information into Transformer models by using absolute position embedding. |
| Outcome: | The proposed methods improve Shaw-RPE and XL-R PE and achieve the best overall performance among five different RPEs. |
Copied to clipboard
| Challenge: | Experiments on five text classification benchmarks and five backbone models have shown that our methods reduce the error rate over Mixup variants in a significant margin (up to 31.3%), especially in low-resource conditions (upto 17.5%). |
| Approach: | They propose to add a small adversarial perturbation to the mixing coefficients rather than the examples to relax locally linear constraints. |
| Outcome: | Experiments on five text classification benchmarks and five backbone models show that the proposed methods reduce the error rate over Mixup variants by 31.3%, especially in low-resource conditions. |
Copied to clipboard
| Challenge: | Existing evaluations of grammatical error correction systems use reference-based metrics, but they are limited because of multiple correct outputs. |
| Approach: | They propose a system that uses commonly available tools to evaluate grammatical error correction (GEC) systems. |
| Outcome: | The proposed system solves the issues related to the use of a reference and does not need another annotated dataset for fine-tuning. |
Copied to clipboard
| Challenge: | Existing language models do not produce suitable representations at the discourse level. |
| Approach: | They propose to augment BERT-style language models with a mechanism that allows them to learn suitable discourse-level representations by incorporating top-down connections that operate at the intermediate layers of the network. |
| Outcome: | The proposed approach improves in 6 out of 11 tasks by detecting discourse relationship detection. |
Copied to clipboard
| Challenge: | Pre-trained models can be maliciously poisoned with certain triggers, causing a security threat. |
| Approach: | They propose a stronger weight poisoning attack method that introduces a layerwise weight poison strategy to plant deeper backdoors. |
| Outcome: | The proposed method can be widely applied and provide hints for future models robustness studies. |
Copied to clipboard
| Challenge: | Existing approaches to improve the early exiting of natural language processing (NLP) are notoriously gigantic and slow in both training and inference. |
| Approach: | They propose a framework for improving the early exiting of BERT by asking each exit to distill knowledge from each other. |
| Outcome: | The proposed framework outperforms the state-of-the-art (SOTA) BERT early exiting methods on the GLUE benchmark. |
Copied to clipboard
| Challenge: | Unlike discrete text prompts used by GPT-3, soft prompts are learned through backpropagation and can be tuned to incorporate signals from any number of labeled examples. |
| Approach: | They propose a mechanism for learning "soft prompts" to condition frozen language models to perform specific downstream tasks. |
| Outcome: | The proposed method outperforms fewshot learning using GPT-3 and matches the quality of model tuning as models exceed billions of parameters. |
Copied to clipboard
| Challenge: | a recent study has shown that fonts with a large number of missing glyphs are difficult to model due to the relative sparsity of most fonts. |
| Approach: | They propose a deep generative model that performs typography analysis and font reconstruction by learning disentangled manifolds of both font style and character shape. |
| Outcome: | The proposed model scales up the number of character types we can model compared to previous methods . it can generalize to characters that were not observed during training time, and it compares favorably to other models . |
Copied to clipboard
| Challenge: | Text-based games are important testbeds for reinforcement learning in the natural language domain. |
| Approach: | They propose a method that learns interpretable action policy rules from symbolic abstractions of textual observations for improved generalization. |
| Outcome: | The proposed method outperforms existing methods in RL using 5-10x fewer training games. |
Copied to clipboard
| Challenge: | In spite of impressive results of neural networks, the huge model size has hindered their applications in cases where computation and memory resources are limited. |
| Approach: | They propose a method for layer-wise pruning using mutual information based feature selection in SVMs and logistic regression. |
| Outcome: | The proposed pruning strategy offers greater speedup and higher performance than weight-based pruning methods. |
Copied to clipboard
| Challenge: | Short text classification is a fundamental task in natural language processing. |
| Approach: | They propose a new method called SHINE which is based on graph neural network for short text classification. |
| Outcome: | The proposed method outperforms state-of-the-art methods on benchmark short text datasets. |
Copied to clipboard
| Challenge: | Existing studies studying OOD detection in NLP often rely on external data to diversify model predictions. |
| Approach: | They propose a framework which mimics OOD detection behavior without external data . they take text classification as an archetype and compare them to existing datasets . |
| Outcome: | The proposed framework can resolve in- and out-distribution examples in a natural way using existing datasets. |
Copied to clipboard
| Challenge: | Masked language modeling (MLM) is widely used in natural language processing for self-supervised learning of text representations. |
| Approach: | They propose to use token-level classification tasks as main pretraining objectives instead of Masked language modeling (MLM) . Empirical results show that pretraining a model with 41% of the BERT-BASE’s parameters, BERT MEDIUM results in only a 1% drop in GLUE scores with their best objective. |
| Outcome: | Empirical results show that the proposed methods achieve comparable or better performance to MLM using a BERT-BASE architecture. |
Copied to clipboard
| Challenge: | Large pre-trained language models (PLMs) have shown overwhelming performances on many tasks, but their large size and slow inference speed have hindered practical deployments. |
| Approach: | They propose a hierarchical relational knowledge distillation method to capture hierarchic and domain relational information. |
| Outcome: | The proposed method outperforms existing methods on multi-domain datasets and is highly reproducible. |
Copied to clipboard
| Challenge: | Existing methods to defend against adversarial word-substitution attacks have not been evaluated or compared in a systematic manner. |
| Approach: | They propose to compare different defense methods under representative adversarial attacks . they propose a method that improves the robustness of neural text classifiers against such attacks a . |
| Outcome: | The proposed method improves robustness of neural text classifiers against such attacks by a significant margin. |
Copied to clipboard
| Challenge: | Existing frameworks for imbalanced text classification can generate anchor instances for difficult samples . difficult samples are hard to classify as they are embedded into an overlapping semantic region with the majority class. |
| Approach: | They propose a Mutual Information constrained Semantically Oversampling framework that generates anchor instances for difficult samples to help the backbone network determine the re-embedding position of a non-overlapping representation. |
| Outcome: | The proposed framework can generate anchor instances to help classifiers achieve significant improvements over baselines on a variety of imbalanced text classification tasks. |
Copied to clipboard
| Challenge: | Existing methods for multi-label document classification ignore the heterogeneous graphical structures of metadata and labels. |
| Approach: | They propose a neural network based approach to multi-label document classification that uses two heterogeneous graphs to model metadata and labels. |
| Outcome: | The proposed approach outperforms state-of-the-art models on two benchmark datasets. |
Copied to clipboard
| Challenge: | Recent research has focused on quantum-inspired algorithms for NLP and quantum-based algorithms for cognition. |
| Approach: | They propose to categorize quantum-inspired algorithms according to quantum theory, linguistic targets that are modeled, and the downstream application. |
| Outcome: | The proposed methods are categorized according to the use of quantum theory, the linguistic targets that are modeled, and the downstream application. |
Copied to clipboard
| Challenge: | Sequence labeling aims to predict fine-grained sequences of labels for text, but lack of token-level annotated data hinders the effectiveness of supervised methods. |
| Approach: | They propose a Meta Teacher-Student (MetaTS) Network to alleviate data scarcity by leveraging large multilingual unlabeled data. |
| Outcome: | The proposed meta learning method alleviates data scarcity by leveraging large multilingual unlabeled data. |
Copied to clipboard
| Challenge: | Existing NMT models can handle meaning ambiguities based on local contexts, but it remains a challenge to translate words in implicit collocations. |
| Approach: | They propose heterogeneous ways of embedding topic information into NMT models . they propose to incorporate topic knowledge embeddable into the NMT model . |
| Outcome: | The proposed methods outperform baselines on English -> German and English => French translation tasks. |
Copied to clipboard
| Challenge: | Existing models require a more expressive vocabulary to represent all languages . however, increasing the vocabulary size significantly slows down the pre-training speed . |
| Approach: | They propose an algorithm VoCap to determine the desired vocabulary capacity of each language. |
| Outcome: | The proposed algorithm improves cross-lingual model pre-training while reducing side effects of increasing vocabulary size. |
Copied to clipboard
| Challenge: | Recent research questions the importance of dot-product self-attention in Transformer models and shows that most attention heads learn simple positional patterns. |
| Approach: | They propose a novel mechanism to replace dot-product self-attention with a recurrent atteNtion mechanism that directly learns attention weights without token-to-token interaction. |
| Outcome: | The proposed model outperforms the Transformer model on translation tasks with fewer parameters and inference time. |
Copied to clipboard
| Challenge: | Existing approaches to scale out spoken language understanding to low-resource languages are noisy. |
| Approach: | They propose a method for mitigating noise in augmented data by training models with augmented datasets. |
| Outcome: | The proposed method outperforms state-of-the-art methods on two benchmark datasets. |
Copied to clipboard
| Challenge: | Existing methods to improve multi-head self-attention are lacking in many languages. |
| Approach: | They propose a redundant head enlivening method to identify redundant heads and vitalize their potential by learning syntactic relations and prior knowledge in the text. |
| Outcome: | The proposed method can identify and vitalize redundant heads without sacrificing the roles of important heads. |
Copied to clipboard
| Challenge: | Unsupervised machine translation relies on parallel corpora for training, but performance still lags behind traditional supervised machine translators. |
| Approach: | They propose to leverage shared grammar clues to provide more explicit language parallel signals to enhance the training of unsupervised machine translation models. |
| Outcome: | The proposed models improve on a common language pair training task in English and german, and use embedding alignments and pretrained language models to synthesize pseudo parallel corpora. |
Copied to clipboard
| Challenge: | Experimental results show document-level neural machine translation improves lexical consistency . inconsistent translations tend to confuse readers in some cases . |
| Approach: | They propose to use a word link to obtain a document word link and an auxiliary loss function to constrain that their translation should be consistent. |
| Outcome: | The proposed approach improves translation consistency on ChineseEnglish and EnglishFrench translation tasks. |
Copied to clipboard
| Challenge: | Experimental results show that bidirectional training pushes the SOTA neural machine translation performance significantly higher. |
| Approach: | They propose a bidirectional training strategy that updates model parameters at the early stage and tunes it normally. |
| Outcome: | The proposed approach pushes the SOTA neural machine translation performance significantly higher on 15 translation tasks on 8 language pairs. |
Copied to clipboard
| Challenge: | Neural machine translation models are trained to maximize the likelihood of next token given previous golden tokens as inputs, but at the inference stage, golden token is unavailable. |
| Approach: | They propose to use scheduled sampling to replace ground-truth tokens with predicted tokens to bridge the gap between training and inference. |
| Outcome: | The proposed methods outperform the Transformer baseline and vanilla scheduled sampling on three large-scale WMT tasks. |
Copied to clipboard
| Challenge: | Existing non-autoregressive neural machine translations have poor inference speed but weak recognition of erroneous translation pieces. |
| Approach: | They propose an architecture to explicitly learn to rewrite the erroneous translation pieces. |
| Outcome: | The proposed architecture can achieve better performance while significantly reducing decoding time. |
Copied to clipboard
| Challenge: | Existing position representations suffer from a lack of generalization to test data with unseen lengths or high computational cost. |
| Approach: | They propose to achieve shift invariance by randomly shifting absolute positions during training by a SHAPE algorithm that is empirically comparable to its counterpart. |
| Outcome: | The proposed method outperforms existing representations on sequence-to-sequence tasks due to extrapolation, i.e., the ability to generalize to sequences that are longer than those observed during training. |
Copied to clipboard
| Challenge: | Training QE models require massive parallel data with hand-crafted quality annotations, which are time-consuming and labor-intensive to obtain. |
| Approach: | They propose a self-supervised method to evaluate machine-translated sentences without references by recovering masked target words. |
| Outcome: | The proposed method outperforms previous unsupervised methods on several QE tasks in different language pairs and domains. |
Copied to clipboard
| Challenge: | Existing work on unsupervised domain adaptation of neural machine translation assumes access to monolingual text in either the source or target language in the new domain. |
| Approach: | They propose a method to extract in-domain sentences from a large generic monolingual corpus from 'missing' text. |
| Outcome: | The proposed method outperforms baselines up to +1.5 BLEU score on five diverse domains in three language pairs and a real-world translation scenario. |
Copied to clipboard
| Challenge: | Existing models for text classification are limited in performance, resulting in poor rumor detection. |
| Approach: | They propose to use Chinese microblogs to detect rumors using pre-trained language models and auxiliary features such as comments to mask co-attention. |
| Outcome: | The proposed model outperforms the state-of-the-art on Weibo20 and three existing social media datasets. |
Copied to clipboard
| Challenge: | Existing approaches to combining knowledge Graphs (KGs) are incomplete but complementary to each other. |
| Approach: | They propose a novel Active Learning framework for neural EA that creates highly informative seed alignments to obtain more effective models with less annotation cost. |
| Outcome: | The proposed framework significantly improves sampling quality with good generality across different datasets, EA models and amount of bachelors. |
Copied to clipboard
| Challenge: | a real-world information extraction system for semi-structured document images often involves a long pipeline of multiple modules, which can lead to unstable performance if not designed carefully. |
| Approach: | They propose to use a sequence generation task to build an end-to-end IE system . they propose to combine three manually engineered modules with one data-driven module . |
| Outcome: | The proposed system can be easily replaced and deployed in large-scale production. |
Copied to clipboard
| Challenge: | Existing algorithms for math word problems only capture word-level relationship and ignore to build hierarchical reasoning like the human being. |
| Approach: | They propose a Reasoning with Pre-trained Knowledge and Hierarchical Structure network that uses outside knowledge to build hierarchical reasoning like the human being. |
| Outcome: | The proposed method outperforms state-of-the-art methods on two large-scale datasets and boosts performance. |
Copied to clipboard
| Challenge: | Existing studies have shown the effectiveness of sequence-to-sequence (Seq2Seque) on mathematics solving. |
| Approach: | They propose a graph-to-sequence neural network which can learn hierarchical information of graphs inputs to solve mathematical problems and speculate answers. |
| Outcome: | The proposed neural network outperforms other neural networks in hidden information learning and mathematics resolving. |
Copied to clipboard
| Challenge: | GPT-3 has been used to train large-scale language models on hundreds of billion scale data. |
| Approach: | They propose a Korean variant of GPT-3 that uses Korean tokens to train in-context models. |
| Outcome: | The proposed method shows state-of-the-art zero-shot and few-shot learning on downstream tasks in Korean. |
Copied to clipboard
| Challenge: | API recommendation tools can help programmers use APIs by recommending which APIs to be used next given the APIs that have been written. |
| Approach: | They propose a cross-library API recommendation approach that uses BPE to split API calls in each sequence and pre-train a GPT based language model. |
| Outcome: | The proposed APIRecX can recommend APIs that are previously regarded as OOV . it can migrate knowledge of existing libraries to a new library and recommend API that is previously viewed as OVO . |
Copied to clipboard
| Challenge: | Knowledge graphs are incomplete with many facts missing, causing performance bottlenecks in many applications. |
| Approach: | They propose a general multi-hop reasoning task that can be formulated as a search process and can be extended to long-distance reasoning scenarios. |
| Outcome: | The proposed model improves on baselines in short and long distance reasoning scenarios. |
Copied to clipboard
| Challenge: | Backchannel (BC) is a short and quick reaction signal of a listener to a speaker's utterances. |
| Approach: | They propose a model that utilizes lexical information in utterances to enhance backchannel (BC) prediction. |
| Outcome: | The proposed model showed 14.24% performance improvement compared to baseline in the four BC categories: continuer, understanding, empathic response, and No BC. |
Copied to clipboard
| Challenge: | Lack of large-scale terminology definition dataset hinders definition generation . lack of precise terminology definitions poses great challenges in scientific communication . |
| Approach: | They propose a large-scale terminology definition dataset Graphine that exploits the graph structure of terminologies to generate graph-aware text generation models. |
| Outcome: | The proposed model outperforms existing models by exploiting graph structure of terminologies. |
Copied to clipboard
| Challenge: | Existing approaches to tag recommendation neglect orderlessness and inter-dependency . Empirical results on Instagram and Stack Overflow show that our method is significantly superior to the previous approaches. |
| Approach: | They propose a sequence-oblivious generation method for tag recommendation . the next tag to be generated is independent of the order of the generated tags . they also propose regressive generation methods that take orderlessness into account . |
| Outcome: | Empirical results show that the proposed method is superior to previous approaches . the proposed system is based on two domains, Instagram and Stack Overflow . |
Copied to clipboard
| Challenge: | a new study proposes a conversational search system that integrates product attributes and dialog with search . but it faces two real world challenges: imperfect product schema/knowledge and lack of training dialog data . |
| Approach: | They propose an end-to-end conversational search system that integrates search with text . they propose an utterance transfer approach that generates dialogue utterations from other domains . |
| Outcome: | The proposed system outperforms the best tested baseline in a conversational search dataset for online shopping. |
Copied to clipboard
| Challenge: | Current approaches to SEC typically leverage a pre-training then fine-tuning procedure that treats data equally. |
| Approach: | They propose a self-supervised curriculum learning approach to improve model performance and model learning. |
| Outcome: | The proposed approach improves the model training and improves CL measurement. |
Copied to clipboard
| Challenge: | Existing approaches for bug fixing lack generality and use only textual or structured information. |
| Approach: | They propose an intuitive yet effective general framework called Fix-Filter-Fix for bug fixing that connects models with their filter mechanism to filter out the last model’s unchanged fix to the next. |
| Outcome: | The proposed framework can quantify and accurately calculate the lifting effect of the model. |
Copied to clipboard
| Challenge: | Existing deep reinforcement learning methods require many trials before convergence and no direct interpretability of trained policies is provided. |
| Approach: | They propose a novel RL method which can learn symbolic and interpretable rules in their differentiable network. |
| Outcome: | The proposed method can learn symbolic and interpretable rules in their differentiable network. |
Copied to clipboard
| Challenge: | Biomedical Concept Normalization (BCN) is widely used in biomedical text processing . despite numerous surface variants of biomedically-defined concepts, it remains challenging and unsolved. |
| Approach: | They propose a framework that uses hypernyms and synonyms to facilitate BCN . they use list-wise training to make use of both hypernies and synonym entities . |
| Outcome: | The proposed framework outperforms the state-of-the-art model on the NCBI dataset. |
Copied to clipboard
| Challenge: | Existing methods to integrate knowledge into text can confuse the representation and import unexpected noises. |
| Approach: | They propose to leverage capsule routing to associate knowledge with medical literature hierarchically . they extract two fragments from medical literature and encode them into fragment representations . |
| Outcome: | The proposed method can more accurately associate knowledge with medical literature than mainstream methods. |
Copied to clipboard
| Challenge: | Recent approaches to identify metaphors ignore extra information from data, such as contextual information and broader discourse information. |
| Approach: | They propose a model augmented with hierarchical contextualized representation to extract more information from both sentence-level and discourse-level. |
| Outcome: | The proposed model outperforms state-of-the-art methods on two tasks using a VUA dataset. |
Copied to clipboard
| Challenge: | Chinese Spelling Check is a nontrivial task because of the nature of ideographic language. |
| Approach: | They propose a pretrained model with graph-based extra features that captures erroneous patterns . they use a graph neural network to introduce radical and pinyin information as visual and phonetic features. |
| Outcome: | The proposed model can show competitive performance on OCR datasets where most errors are not covered by existing confusion set. |
Copied to clipboard
| Challenge: | Existing medical report generation efforts focus on producing human-readable reports, yet the generated text may not be well aligned to the clinical facts. |
| Approach: | They propose to automate the generation of medical reports from chest X-ray image inputs . medical reports are the primary medium, which physicians communicate findings from scans - authors say . |
| Outcome: | The proposed method achieves fluency and clinical accuracy on common metrics. |
Copied to clipboard
| Challenge: | Document Retrieval (DR) requires the machine to retrieve and rank documents according to their relevance with the query. |
| Approach: | They propose a ranking model DR-BERT which improves the Document Retrieval task by a task-adaptive training process and a Segmented Token Recovery Mechanism. |
| Outcome: | The proposed ranking model keeps in the top three on the MS MARCO leaderboard since 2020. |
Copied to clipboard
| Challenge: | Existing scientific claim verification models have problems of error propagation among modules and lack of sharing valuable information among modules. |
| Approach: | They propose an approach that jointly learns the modules for the three tasks with a machine reading comprehension framework by including claim information. |
| Outcome: | The proposed approach outperforms existing models on the SciFact dataset on the three tasks of abstract retrieval, rationale selection and stance prediction. |
Copied to clipboard
| Challenge: | Experimental results show that joint models of word segmentation and POS tagging can lead to better performance because they are closely related. |
| Approach: | They propose a domain adaption method for Chinese word segmentation and POS tagging that uses a simple metric to model the gaps between target and target domains. |
| Outcome: | The proposed method can gain significant performance improvements over baselines on a benchmark dataset. |
Copied to clipboard
| Challenge: | a new benchmark is developed to answer open-domain questions from text . the system uses a single multi-task transformer model to perform all the necessary subtasks . |
| Approach: | They develop a unified system to answer directly from open-domain questions . they use a single multi-task transformer model to perform all the necessary subtasks . |
| Outcome: | The proposed system can answer open-domain questions on any text collection without prior knowledge of reasoning complexity. |
Copied to clipboard
| Challenge: | Existing iterative approaches to open-domain question answering use predefined strategies . e.g., BM25, DPR, and hyperlink are defined as actions . |
| Approach: | They propose a novel adaptive information-seeking strategy for open-domain question answering . they propose to use a partially observed Markov decision process to select a proper retrieval action . |
| Outcome: | Experiments on SQuAD Open and HotpotQA fullwiki show that AISO outperforms baseline methods with predefined strategies in retrieval and answer evaluations. |
Copied to clipboard
| Challenge: | a recent paper addresses the problem of solving math word problems automatically . a number of approaches have been proposed for solving word problems . |
| Approach: | They employ a sequence-to-sequence model to generate intermediate representations for word problems . they then use a probabilistic programming system to provide the answer . their best performing model incorporates general-domain contextualised word representations . |
| Outcome: | The proposed model is the best performing on a declarative language and a probabilistic programming system. |
Copied to clipboard
| Challenge: | Multiple-choice MRC is one of the most studied tasks in MRC due to the convenience of evaluation and the flexibility of answer format. |
| Approach: | They propose to use multiple-choice MRC to explain a trained model and reveal how it arrives at the prediction by punishing illogical attributions. |
| Outcome: | The proposed method improves model performance without external information and model structure change without any external information. |
Copied to clipboard
| Challenge: | Existing KBQA methods focus on the natural language but ignore textual information carried by the nodes and edges. |
| Approach: | They propose to perform relation extraction, relation matching, and relation reasoning tasks to align the natural language expressions to the relations in the KB and reason over the missing connections. |
| Outcome: | Experiments on WebQSP show that the proposed model outperforms baselines even when the KB is incomplete. |
Copied to clipboard
| Challenge: | Dense retrieval methods have shown great promise over sparse methods in a range of NLP problems. |
| Approach: | They propose to use dense phrase retrieval to learn coarse-level retrieval including passages . they show phrase retrievals can be fine-tuned for more coarse-grained retrieval units . |
| Outcome: | The proposed method improves passage retrieval accuracy and QA performance with fewer passages. |
Copied to clipboard
| Challenge: | Existing question answering models are based on textual entailment tasks . prior work has focused on QA on premise-based questions . |
| Approach: | They propose a neural-symbolic QA approach that integrates natural logic reasoning within deep learning architectures towards developing effective question answering models. |
| Outcome: | The proposed model outperforms previous work on multiple-choice science questions . it integrates natural logic reasoning within deep learning architectures to build proof paths . |
Copied to clipboard
| Challenge: | Existing studies train independent or pipeline systems for the two subtasks but are trivial by using hard-label decisions to activate question generation. |
| Approach: | They propose a method to smooth two dialogue states in one decoder and bridge decision making and question generation to provide a richer dialogue state reference. |
| Outcome: | The proposed method achieves state-of-the-art on the OR-ShARC dataset. |
Copied to clipboard
| Challenge: | Popular, large, pre-trained models fall far short of expert humans in acquiring finance knowledge and in complex multi-step numerical reasoning on that knowledge. |
| Approach: | They propose a large-scale dataset with Question-Answering pairs over financial reports written by financial experts to facilitate analytical progress. |
| Outcome: | The proposed dataset is the first of its kind and is available on github. |
Copied to clipboard
| Challenge: | Pre-trained sequence to sequence models are effective in making and generating NL explanations, but they have many shortcomings. |
| Approach: | They propose a model that uses sentence markers to eliminate explanation fabrication . they use fusion-in-decoder architecture to handle long input contexts . |
| Outcome: | The proposed model significantly improves on the ERASER explainability benchmark. |
Copied to clipboard
| Challenge: | Recent named entity recognition models have great performance on many conventional benchmarks, but it is not reliable in realistic applications. |
| Approach: | They propose a method to create natural adversarial examples using Wikidata and pre-trained language models. |
| Outcome: | The proposed method produces natural adversarial examples with a shifted distribution from training data. |
Copied to clipboard
| Challenge: | Existing studies have focused on diagnosing LMs' reasoning abilities in natural language understanding tasks. |
| Approach: | They propose a diagnostic method for first-order logic reasoning with a proposed benchmark, LogicNLI. |
| Outcome: | The proposed method disentangles the target FOL reasoning from commonsense inference and can be used to diagnose LMs from four perspectives: accuracy, robustness, generalization, and interpretability. |
Copied to clipboard
| Challenge: | Psychometric dimensions are important for understanding user behavior in various contexts including health, security, e-commerce, and finance. |
| Approach: | They propose to construct a corpus for psychometric natural language processing related to important dimensions such as trust, anxiety, numeracy, and literacy, in the health domain. |
| Outcome: | The proposed corpus includes 8,502 user-generated responses from 8,502-person survey datasets and includes self-reported demographic information, including race, sex, age, income, and education. |
Copied to clipboard
| Challenge: | 16K FAQ items scraped from 55 credible websites . 32 human-annotated FAQ items for each query. |
| Approach: | They present a large, challenging dataset for FAQ retrieval for COVID-19 . they use a FAQ bank, Query Bank and Relevance Set to evaluate the dataset . |
| Outcome: | The proposed model achieves 48.8 under P@5 and is compared with other datasets. |
Copied to clipboard
| Challenge: | Existing datasets for word prediction with long-range context have not been tested. |
| Approach: | They propose automatic and manual selection strategies tailored to Chinese to ensure that target words can only be predicted with long-term context. |
| Outcome: | The proposed model is 45 points behind human in terms of top-1 word prediction accuracy. |
Copied to clipboard
| Challenge: | Recent success of neural language models on the Winograd Schema Challenge has called for further investigation of commonsense reasoning ability of these models. |
| Approach: | They propose a logic-based framework that focuses on high-quality commonsense knowledge. |
| Outcome: | The proposed framework focuses on high-quality commonsense knowledge. |
Copied to clipboard
| Challenge: | Masked language models have contributed to drastic performance improvements with regard to zero anaphora resolution (ZAR). |
| Approach: | They propose a pretraining task that trains MLMs on anaphoric relations with explicit supervision and a finetuning method that remedies a notorious discrepancy. |
| Outcome: | The proposed method improves zero anaphora resolution in Japanese ZAR . it uses a pretrain task and finetuning task to correct the discrepancy . |
Copied to clipboard
| Challenge: | Existing work uses sentences within the same batch as negatives, which suffers from easy negatives. |
| Approach: | They propose to align sentence representations from different languages into a unified embedding space . they adapt MoCo to further improve the quality of alignment . |
| Outcome: | The proposed model achieves state-of-the-art on several tasks. |
Copied to clipboard
| Challenge: | Existing methods for continual learning for semantic parsing fail to account for special properties of structured outputs . retraining from scratch is not feasible due to the fast growing number of tasks . |
| Approach: | They propose a continual learning method that uses sequential learning to learn tasks without accessing full training data from previous tasks. |
| Outcome: | The proposed method achieves a 3-6 times speedup compared to re-training from scratch. |
Copied to clipboard
| Challenge: | Existing studies on pronoun coreference resolution focus on anaphora and cataphores . exophoric pronounos are common in daily communications, but can be disambiguated by general topics of the dialogue. |
| Approach: | They propose to leverage local context and global topics of dialogues to solve out-of-text PCR problem by adding topic regularization. |
| Outcome: | Extensive experiments show that topic regularization can be used to solve the out-of-text PCR problem. |
Copied to clipboard
| Challenge: | Existing models focus on word-level local matching and neglect the importance of contextual information. |
| Approach: | They propose a context-aware interaction network to properly align two sequences and infer their semantic relationship by using gate fusion layers. |
| Outcome: | The proposed model can accurately align two sequences and infer their semantic relationship on two question matching datasets. |
Copied to clipboard
| Challenge: | Existing taxonomies are unable to maintain coverage due to the rising of new concepts . TEMP uses pre-trained contextual encoders to predict the position of new ideas . |
| Approach: | They propose a self-supervised taxonomy expansion method that ranks taxonomies by ranking them . they use pre-trained contextual encoders to train the model with dynamic margin loss . |
| Outcome: | The proposed method outperforms state-of-the-art taxonomy expansion methods by 14.3% and 15.8% on public benchmarks. |
Copied to clipboard
| Challenge: | Existing studies focus on frame semantic parsing as a graph construction problem. |
| Approach: | They propose an end-to-end neural model to tackle frame semantic parsing jointly. |
| Outcome: | The proposed model is highly competitive and performs better than pipeline models on two benchmark datasets. |
Copied to clipboard
| Challenge: | Recent studies have shown that powerful pre-trained language models can be fooled by small perturbations or intentional attacks. |
| Approach: | They propose a framework for fine-tuning PLMs using a masked language model and Gaussian noise to augment semantically relevant examples with sufficient diversity. |
| Outcome: | The proposed framework improves the robustness of pre-trained language models and alleviates performance degradation under adversarial attacks. |
Copied to clipboard
| Challenge: | Existing models for metaphor detection require a large amount of labeled data and are not linguistically-based. |
| Approach: | They propose a ContrAstive pre-Trained modEl (CATE) for metaphor detection with semi-supervised learning using a pre-trained model to obtain a contextual representation of target words. |
| Outcome: | The proposed model outperforms existing models on several benchmark datasets and achieves better performance against state-of-the-art models. |
Copied to clipboard
| Challenge: | Dependency parsers are not designed for capturing interaction between opinion words and aspect words. |
| Approach: | They propose to learn an aspect-centric tree structure to shorten distance between aspects and opinion words. |
| Outcome: | The proposed model outperforms baselines on five aspect-based sentiment datasets. |
Copied to clipboard
| Challenge: | Existing models focus on aspect term extraction, opinion term extraction and sentiment polarity classification but ignore the difference. |
| Approach: | They propose a joint aspect-based sentiment analysis task that focuses on the difference between the two tasks to improve the model's robustness. |
| Outcome: | Empirical results show that the proposed model outperforms the previous state-of-the-art on four benchmark datasets. |
Copied to clipboard
| Challenge: | Existing studies on argumentation mining focus on monological argumentation and dialogical argumentation. |
| Approach: | They propose a mutual guidance framework that could guide arguments in one passage . they propose an inter-sentence relation graph to effectively model the inter-relations between two sentences . |
| Outcome: | The proposed method outperforms the current state-of-the-art model. |
Copied to clipboard
| Challenge: | Empirical studies on three different benchmark conversation datasets demonstrate the effectiveness of the proposed model over several strong baselines. |
| Approach: | They propose an addressee-aware module to automatically learn whether the participant keeps the historical emotional state or is affected by others in the next upcoming turn. |
| Outcome: | The proposed model can predict the participant's emotion in the next upcoming turn without knowing the participant’s response yet. |
Copied to clipboard
| Challenge: | Aspect-based sentiment analysis (ABSA) predicts sentiment polarity for aspect term in sentences . labeled data stored at different locations and inaccessible due to privacy or legal concerns . |
| Approach: | They propose a model with federated learning to combine labeled data across different domains . they incorporate topic memory to take data from diverse domains into consideration . |
| Outcome: | The proposed model outperforms baselines on a simulated environment with three nodes. |
Copied to clipboard
| Challenge: | Comparative opinion mining is an important task in opinion mining. |
| Approach: | They propose a task to extract comparative opinion quintuples from product reviews . they propose supplementary annotations and construct three datasets for the task . |
| Outcome: | The proposed method outperforms baseline systems on three datasets and represents a strong benchmark for COQE. |
Copied to clipboard
| Challenge: | Existing audio-language task-specific predictive approaches focus on building complicated late-fusion mechanisms. |
| Approach: | They propose a cross-modal transformer for audio-and-language that learns inter-modal connections between audio and language through two proxy tasks on a large amount of audio- and-language pairs. |
| Outcome: | The proposed model improves on multiple audio-and-language tasks and can be used in fine-tuning phase. |
Copied to clipboard
| Challenge: | Existing methods for temporal language grounding in videos are boundary regression and span extraction tasks. |
| Approach: | They propose a Relation-aware Network to localize a temporal span relevant to a given query sentence. |
| Outcome: | The proposed framework selects a video moment choice from the predefined answer set with the aid of coarse-and-fine choice-query interaction and choice-choice relation construction. |
Copied to clipboard
| Challenge: | Existing approaches to end-to-end speech translation (E2E) models only allow one way knowledge transfer, which is limited by the performance of the teacher model. |
| Approach: | They propose a one-way knowledge transfer paradigm where the MT and ST models are collaboratively trained and considered as peers rather than teacher/student. |
| Outcome: | The proposed model improves the performance of end-to-end speech translation (ST) task by combining knowledge from two models with peer models. |
Copied to clipboard
| Challenge: | Existing MAS models cannot leverage GPLMs’ powerful generation ability. |
| Approach: | They propose a method to construct vision guided (VG) GPLMs that incorporate visual information while maintaining their original text generation ability. |
| Outcome: | The proposed model outperforms the previous state-of-the-art model by 5.7 ROUGE-1, 5.3 ROUGe-2, and 5.1 ROUGEL-L scores on the How2 dataset and contributes 83.6% of the overall improvement. |
Copied to clipboard
| Challenge: | Existing methods for video moment localization have poor performance due to predefined rules. |
| Approach: | They propose a model with a fixed set of learnable moment proposals with 'border-aware loss' they propose to localize the video moment corresponding to the query by locating the start and end timestamps in an untrimmed video. |
| Outcome: | The proposed model outperforms state-of-the-art models on two challenging benchmarks. |
Copied to clipboard
| Challenge: | Prior work on Vision-and-Language Navigation (VLN) tasks do not measure how much of a language instruction the agent is able to follow. |
| Approach: | They propose a language-aligned supervision scheme that measures the number of sub-instructions the agent has completed during navigation. |
| Outcome: | The proposed method is based on the previous work on the Vision-and-Language Navigation task, which assumes a discrete navigation graph (navgraph) but not on the current work. |
Copied to clipboard
| Challenge: | Using deep learning to improve healthcare is challenging due to the complexity of EHR data. |
| Approach: | They propose a method to integrate clinical notes from EHR and combine them with different data to improve prediction performance. |
| Outcome: | The proposed model outperforms the state-of-the-art method without clinical notes on two prediction tasks. |
Copied to clipboard
| Challenge: | Sentence extractive summarization shortens a document by selecting sentences for a summary while preserving its important contents. |
| Approach: | They propose a nested tree-based extractive summarization model on RoBERTa that uses syntactic and discourse trees to represent sentences in a given document. |
| Outcome: | The proposed model outperforms baseline models on the CNN/DailyMail dataset and achieves significantly better scores than the baseline models in terms of coherence and comparable scores to the state-of-the-art models. |
Copied to clipboard
| Challenge: | Sentence-level extractive text summarization is difficult to model the importance of sentences. |
| Approach: | They propose a Frame Semantic-Enhanced Sentence Modeling for Extractive Summarization that leverages Frame semantics to model sentences from both intra-sentence level and inter-sentent level. |
| Outcome: | The proposed model outperforms six state-of-the-art methods on two benchmark corpus datasets. |
Copied to clipboard
| Challenge: | Existing methods for code summarization do not capture rich information in ASTs . existing methods are labor-intensive and time-consuming to document code with good summaries manually. |
| Approach: | They propose a model that hierarchically splits and reconstructs ASTs by a neural network . they propose to use AST embeddings and a vanilla code token encoder to generate the model . |
| Outcome: | The proposed model splits and reconstructs ASTs into subtrees and then aggregates embeddings of subtreas to get the complete AST. |
Copied to clipboard
| Challenge: | Existing extractive multi-document summarization methods score each sentence individually and extract salient sentences one by one. |
| Approach: | They propose a novel framework for extractive multi-document summarization that selects a sub-graph as the summary instead of selecting salient sentences. |
| Outcome: | The proposed framework improves on existing methods on multi-document datasets and human evaluations show it produces more coherent and informative summaries. |
Copied to clipboard
| Challenge: | Sentence fusion is a conditional generation task that merges related sentences into a coherent text. |
| Approach: | They propose to build an event graph from the input sentences to capture related events in a structured way and use the constructed event graph to guide sentence fusion. |
| Outcome: | The proposed method achieves state-of-the-art on two datasets . it is based on the input sentences and shows that it is effective . |
Copied to clipboard
| Challenge: | Existing automatic headline generation methods cannot include a given phrase in the generated headline. |
| Approach: | They propose a Transformer-based method that guarantees to include a given phrase in a generated headline. |
| Outcome: | The proposed method achieves ROUGE scores comparable to previous methods with Japanese news corpus. |
Copied to clipboard
| Challenge: | Existing methods for abstractive summarization use encoder-decoder attention, but this leads to incomplete copying. |
| Approach: | They propose a copying scheme that takes advantage of prior copying distributions and explicitly encourages the model to copy the input word that is relevant to the previously copied one. |
| Outcome: | The proposed scheme achieves state-of-the-art on summarization benchmarks . it takes advantage of prior copying distributions and explicitly encourages copying . |
Copied to clipboard
| Challenge: | Abstractive summarization models often produce inconsistent statements or false facts. |
| Approach: | They propose an efficient weak-supervised adversarial data augmentation approach to generate factual consistency datasets by backpropagating gradients on token embeddings. |
| Outcome: | The proposed model can make interpretable factual errors tracing on public datasets and is cost-effective. |
Copied to clipboard
| Challenge: | Current sentence encoders are word order sensitive, resulting in poor performance . Adapting word order from one language to another is key in cross-lingual structured prediction. |
| Approach: | They propose a new module to organize words following the source language order . they build structured prediction models with bag-of-words inputs and introduce a module to do this . |
| Outcome: | The proposed model significantly improves target language performance for languages that are distant from the source language. |
Copied to clipboard
| Challenge: | Existing approaches to encode dynamic structures only encode partial information of structures. |
| Approach: | They propose a new attention-based encoder unifying all structures in a transition system. |
| Outcome: | The proposed method significantly improves the test speed and achieves the best transition-based model. |
Copied to clipboard
| Challenge: | Question Generation (QG) is the production of meaningful questions given a set of input passages and corresponding answers. |
| Approach: | They propose a method which uses questions generated heuristically from news summaries as a source of training data for a QG system. |
| Outcome: | The proposed method outperforms previous unsupervised models on three in-domain datasets and three out-of-domain ones. |
Copied to clipboard
| Challenge: | Existing models infer the answer by predicting the sequential relation path or aggregating the hidden graph features. |
| Approach: | They propose a model which jumps between entities at multiple steps . they demonstrate that TransferNet surpasses state-of-the-art models by a large margin . |
| Outcome: | The proposed model surpasses state-of-the-art models on MetaQA and on other datasets. |
Copied to clipboard
| Challenge: | Weakly-supervised table question-answering (TableQA) models have achieved state-of-art performance by using pre-trained BERT transformer to jointly encoding a question and a table to produce structured query for the question. |
| Approach: | They propose a framework for TableQA that incorporates topic-specific vocabulary injection into BERT, a novel text-to-text transformer generator and a logical form re-ranker. |
| Outcome: | The proposed framework provides a reasonably good baseline for topic shift benchmarks. |
Copied to clipboard
| Challenge: | Using a web page and a question, a machine can't understand the contents of web pages. |
| Approach: | They propose a novel dataset for web-based structural reading comprehension that consists of 400K question-answer pairs and a dataset of 6.4K web pages. |
| Outcome: | The proposed dataset consists of 400K question-answer pairs, collected from 6.4K web pages with corresponding HTML source code, screenshots, and metadata. |
Copied to clipboard
| Challenge: | Current datasets targeting ambiguity can be solved by a native speaker with relative ease. |
| Approach: | They present a large-scale dataset based on cryptic crosswords with a cryptical clue. |
| Outcome: | The proposed dataset is based on cryptic crosswords with 523K examples. |
Copied to clipboard
| Challenge: | End-to-end (E2E) trained models for question answering over knowledge graphs (KGQA) are effective, but training a weakly supervised dataset is difficult. |
| Approach: | They extend the boundaries of E2E learning for KGQA to include the training of an ER component. |
| Outcome: | The proposed model is fully differentiable thanks to a recent method for building differentiably KGs. |
Copied to clipboard
| Challenge: | Existing Knowledge-based Question Answering methods use a query graph to find the answer to a question. |
| Approach: | They propose a method that starts with the entire knowledge base and gradually shrinks it to the desired query graph. |
| Outcome: | Experimental results show that the proposed method achieves state-of-the-art performance on ComplexWebQuestion dataset. |
Copied to clipboard
| Challenge: | Generating long passages that maintain long-range coherence is a long-standing problem in natural language generation (NLG). |
| Approach: | They propose a discourse-aware discrete variational Transformer that learns a latent variable sequence that summarizes the global structure of the text and then applies it to guide the generation process at each decoding step. |
| Outcome: | The proposed model can generate long texts with better long-range coherence by learning a latent variable sequence with each latent code. |
Copied to clipboard
| Challenge: | Existing models for generating mathematical word problems are lacking in educational assessment. |
| Approach: | They propose an end-to-end neural model to generate diverse mathematical word problems from commonsense knowledge graph and equations. |
| Outcome: | The proposed model outperforms the SOTA models in terms of evaluation metrics and topic relevance. |
Copied to clipboard
| Challenge: | Text style transfer is a task aimed at converting a text of one style into another while preserving its content. |
| Approach: | They propose a multi-step procedure which builds on a generic pre-trained sequence-to-sequence model and an iterative back-translation approach to train two models in a transfer direction. |
| Outcome: | The proposed method outperforms existing unsupervised approaches on the two most popular style transfer tasks: formality transfer and polarity swap. |
Copied to clipboard
| Challenge: | Existing methods for paraphrase generation rely on language as the pivot . however, there is no evidence that parallel data of paraphrases is needed for paraphrasing. |
| Approach: | They propose to use semantic and syntactic representations as pivot for paraphrase generation. |
| Outcome: | The proposed method can generate paraphrases with better quality than using language as pivot. |
Copied to clipboard
| Challenge: | Pretrained language models (PLMs) have advanced graph-to-text generation, but efficient encoding of graph structure is challenging because of the nature of the data. |
| Approach: | They propose a method to encode graph structure into pretrained language models by training only graph structure-aware adapter parameters. |
| Outcome: | The proposed method outperforms the state-of-the-art on two AMR-to-text datasets, training only 5.1% of the adapter parameters. |
Copied to clipboard
| Challenge: | Existing work on data-to-text generation relies on retrieved "neighbors" but instead generates text token-by-token, left-to right. |
| Approach: | They propose to splice together retrieved segments of text from "neighbor" source-target pairs to generate text token-by-token, left-to-right. |
| Outcome: | The proposed method performs on par with strong baselines in terms of automatic and human evaluation, but allows for more interpretable and controllable generation. |
Copied to clipboard
| Challenge: | Existing approaches to integrate knowledge bases into end-to-end task-oriented dialogue systems are limited in their ability to properly represent the entity of KB. |
| Approach: | They propose a framework that dynamically perceives all relevant entities and dialogue history . it uses a Memory Mask to enforce the entity to focus on its relevant entities . |
| Outcome: | The proposed framework can achieve superior performance over the state of the arts. |
Copied to clipboard
| Challenge: | Existing methods for training dialogue policies rely on a single learning system, but it requires many rounds of interaction. |
| Approach: | They propose a complementary policy learning framework which exploits the complementary advantages of the episodic memory (EM) policy and the deep Q-network (DQN) policy. |
| Outcome: | The proposed framework outperforms existing methods relying on a single learning system on three dialogue datasets. |
Copied to clipboard
| Challenge: | Existing conversational recommender systems (CRS) do not track the deep shift of user interest in conversations due to the complex of high-order and incomplete paths. |
| Approach: | They propose a conversational context-based reinforcement learning model which does explicit multi-hop reasoning on KGs with a contextual context-driven reinforcement learning framework. |
| Outcome: | Extensive experiments show that CRFR improves on paths of interest shift in knowledge graphs (KGs) . |
Copied to clipboard
| Challenge: | Existing datasets for conversational recommendation are limited to English and Chinese . |
| Approach: | They propose a bilingual parallel human-to-human recommendation dialog dataset . the data item is annotated in two languages, both English and Chinese . |
| Outcome: | The proposed dataset provides a testbed for future studies of multilingual and cross-lingual conversational recommendation. |
Copied to clipboard
| Challenge: | Existing systems that use human-to-human dialogs to help users with specific tasks are still unexplored. |
| Approach: | They propose a problem in which a dialog system mimics a troubleshooting agent . they use a dataset grounded on 12 different troubleshooking flowcharts to train the agent a neural model . |
| Outcome: | The proposed model can do zero-shot transfer to unseen flowcharts and sets a strong baseline for future research. |
Copied to clipboard
| Challenge: | Using a model to predict fine-grained emotions along the continuous dimensions of valence, arousal, and dominance (VAD) with a corpus with categorical emotion annotations, we show that our approach reaches comparable performance to that of the state-of-the-art classifiers in categorial emotion classification and shows significant positive correlations with the ground truth VAD scores. |
| Approach: | They propose to train a model to predict fine-grained emotions along the continuous dimensions of valence, arousal, and dominance with a corpus with categorical emotion annotations. |
| Outcome: | The proposed model can predict emotions along the continuous dimensions of valence, arousal, and dominance (VAD) with a corpus with categorical emotion annotations. |
Copied to clipboard
| Challenge: | Fine-grained classification tasks involve distinguishing between classes with subtle differences between them. |
| Approach: | They analyse fine-grained text classification tasks by embedding class relationships into a contrastive objective function to help differently weigh the positives and negatives. |
| Outcome: | The proposed model outperforms previous contrastive methods on emotion classification and sentiment analysis. |
Copied to clipboard
| Challenge: | Existing studies on aspect-level sentiment analysis focus on extracting aspect terms and sentiment polarities separately. |
| Approach: | They propose a multi-modal joint learning approach with auxiliary cross-modal relation detection for multi-dimensional aspect-level sentiment analysis. |
| Outcome: | The proposed approach can obtain all aspect-level sentiment polarities dependent on the jointly extracted specific aspects. |
Copied to clipboard
| Challenge: | Existing methods for Aspect category sentiment analysis use pre-trained language models to learn aspect category-specific representations. |
| Approach: | They propose to make use of pre-trained language models by casting the ACSA tasks into natural language generation tasks, using natural language sentences to represent the output. |
| Outcome: | The proposed method gives the best reported results, having large advantages in few-shot and zero-shot settings. |
Copied to clipboard
| Challenge: | Existing methods for data augmentation address data deficiencies and semantic consistency, but they ignore the second issue. |
| Approach: | They propose a semantics-preserving data augmentation approach that preserves the semantics of a textual sequence. |
| Outcome: | The proposed method achieves better performance on publicly available datasets and stock price/risk movement prediction scenarios. |
Copied to clipboard
| Challenge: | Sentiment analysis systems exhibit sensitivity to protected attributes, while round-trip translation has been shown to normalize text. |
| Approach: | They propose to use round-trip translation to normalize text to reduce the fairness gap between groups in sentiment analysis. |
| Outcome: | The proposed method reduces the fairness gap between groups by up to 47%. |
Copied to clipboard
| Challenge: | Humor detection is difficult due to individualistic and cultural differences in humor perception . authors propose a framework to generate perceived humor labels on Facebook posts . |
| Approach: | They propose a framework to generate perceived humor labels on Facebook posts . they use the naturally available user reactions to these posts to generate the labels . |
| Outcome: | The proposed framework generates perceived humor labels on Facebook posts with no manual annotation needed. |
Copied to clipboard
| Challenge: | Existing summarization methods are prone to generate redundant and incoherent summaries, causing the performance to be worse. |
| Approach: | They propose a Chinese dataset for Customer Service Dialogue Summarization (CSDS) that provides role-oriented summaries to acquire different speakers' viewpoints. |
| Outcome: | The proposed dataset improves the abstractive summaries in two aspects . it also provides role-oriented summary to acquire different speakers’ viewpoints . |
Copied to clipboard
| Challenge: | Existing relation extraction methods focus on extracting relational facts between entity pairs within single sentences or documents. |
| Approach: | They present a problem of cross-document relation extraction (CRE) using human annotations. |
| Outcome: | The proposed dataset is the first human-annotated cross-document RE dataset . it shows that it is challenging to existing RE methods including strong BERT-based models. |
Copied to clipboard
| Challenge: | Recent advances on neural approaches to natural language processing have triggered a renaissance in end-to-end neural open-domain chatbots. |
| Approach: | They propose to use offline and online steps to evaluate the quality of clarifying questions in various open-domain dialogues to improve the quality and accuracy of the system response. |
| Outcome: | The proposed pipeline is suitable as a foundation for further research. |
Copied to clipboard
| Challenge: | Standard train-dev-test splits used to benchmark multiple models are now used in NLP . comparing multiple versions of the same model on the test data leads to overfitting and "expiration" of test sets. |
| Approach: | They propose to use a tune-set when developing neural network methods to do model picking. |
| Outcome: | The proposed model picker is more robust against the evaluated hyperparameter ranges than the standard split split. |
Copied to clipboard
| Challenge: | We present a high-quality and large-scale Vietnamese-English parallel dataset . our dataset is 2.9M pairs larger than the benchmark Vietnamese- English corpus . |
| Approach: | They present a large-scale Vietnamese-English parallel dataset with 3.02M sentence pairs . they compare strong neural baselines and well-known automatic translation engines . |
| Outcome: | The proposed dataset is 2.9M pairs larger than the benchmark Vietnamese-English corpus IWSLT15. |
Copied to clipboard
| Challenge: | Existing studies on verbal leakage cues do not address their impact on models' validity. |
| Approach: | They propose to use LIWC to show verbal leakage cues in lie detection datasets to understand their effect on data collection and examine their validity. |
| Outcome: | The proposed models with more strong verbal leakage cue categories perform better than models trained on a dataset with only a greater number of strong cues. |
Copied to clipboard
| Challenge: | Existing adversarial attack models are vulnerable to adversarials crafted by human-imperceptible perturbations. |
| Approach: | They propose a multi-granularity adversarial attack model that generates high-quality adversarials with fewer queries to victim models. |
| Outcome: | The proposed model generates high-quality adversarial samples with fewer queries to victim models compared to baseline models . the proposed model also reduces query times for black-box models that only output labels without confidence scores . |
Copied to clipboard
| Challenge: | Similarity measures are a vital tool for understanding how language models represent and process language. |
| Approach: | They propose to use cosine similarity and Euclidean distance to understand how words cluster in semantic space. |
| Outcome: | The proposed measures show that rogue dimensions dominate similarity measures and reveal representational quality. |
Copied to clipboard
| Challenge: | Transformer architecture is composed of multi-head attention, which has been extensively analyzed. |
| Approach: | They extended the scope of the analysis of Transformers from solely the attention patterns to the whole attention block, i.e., multi-head attention, residual connection, and layer normalization. |
| Outcome: | The proposed method incorporates the whole attention block, i.e., multi-head attention, residual connection, and layer normalization into the analysis. |
Copied to clipboard
| Challenge: | Experimental results show that popular NLP models are vulnerable to both adversarial and backdoor attacks based on text style transfer. |
| Approach: | They propose to conduct adversarial and backdoor attacks based on text style transfer . the authors propose to use text style to alter the style of a sentence . |
| Outcome: | The proposed methods show that popular models are vulnerable to both attacks based on text style transfer . the results show that the proposed methods perform better than baselines in many aspects . |
Copied to clipboard
| Challenge: | Using data from English cloze tests, we demonstrate wide performance gaps across demographic groups and show that pretrained language models disfavor young non-white male speakers. |
| Approach: | They use data from English cloze tests to examine performance differences of pretrained language models across demographic groups. |
| Outcome: | The models disfavor young non-white male speakers, but larger models reduce performance gaps between majority and minority groups. |
Copied to clipboard
| Challenge: | Existing studies on whether multilingual embeddings can be aligned in a shared space across languages are lacking. |
| Approach: | They propose to learn a projection based on monolingual annotated datasets and evaluate syntactic and lexical information encoded in a shared cross-lingual embedding space. |
| Outcome: | The proposed model can be used to learn representations for languages with low resources. |
Copied to clipboard
| Challenge: | Recent studies have shown that unsupervised sentence representations of neural networks encode syntactic information by observing that neural language models are able to predict the agreement between a verb and its subject. |
| Approach: | They propose to take an alternative look at these results by studying whether neural networks are able to build an abstract sentence representation rather than capture surface statistical regularities. |
| Outcome: | The proposed model can achieve high accuracy on the long-range French object-verb agreement, indicating a possible flaw in the model's syntactic ability. |
Copied to clipboard
| Challenge: | Existing approaches to fine-grained entity typing are based on independent classification paradigms, which make them difficult to recognize inter-dependent, long-tailed and fine-granular entities. |
| Approach: | They propose a label reasoning network that exploits label dependencies knowledge entailed in the data. |
| Outcome: | The proposed network can model, learn and reason complex labels in a sequence-to-set, end-to end manner. |
Copied to clipboard
| Challenge: | Existing approaches to extract text spans from plain text do not fully exploit label knowledge. |
| Approach: | They propose a model to integrate label knowledge into text representations by encoding texts and annotations independently and then integrating label knowledge with an elaborate-designed semantics fusion module. |
| Outcome: | The proposed model achieves state-of-the-art performance on four benchmarks and reduces training time and inference time by 76% and 77% on average compared with the existing paradigm. |
Copied to clipboard
| Challenge: | Existing methods for extracting interpersonal relationships from dialogues are limited to end-to-end learning. |
| Approach: | They propose a neural multi-label classifier that infers relationships from dialogues by external knowledge about speaker features and conversation style. |
| Outcome: | The proposed method outperforms the state-of-the-art methods on large-scale datasets with directed relationships of conversation participants. |
Copied to clipboard
| Challenge: | Existing approaches focus on high-level description of how research is carried out . instead, we focus on the subtleties of how experimental associations are presented . |
| Approach: | They propose a transformer-based approach to relational scientific information extraction that captures associations over experimental variables and their qualifications, subtypes, and evidence. |
| Outcome: | The proposed schema captures causal, comparative, predictive, statistical, and proportional associations over experimental variables along with qualifications, subtypes, and evidence. |
Copied to clipboard
| Challenge: | Existing models for cross-document coreference resolution have been used for within-document entity coreference but have been relatively limited. |
| Approach: | They propose a model that extends the efficient sequential prediction paradigm for coreference resolution to cross-document settings and achieves competitive results for both entity and event coreference. |
| Outcome: | The proposed model achieves competitive results for entity and event coreference while minimizing error propagation in complex reasoning tasks. |
Copied to clipboard
| Challenge: | Infusing factual knowledge into pre-trained models is fundamental for many knowledge-intensive tasks. |
| Approach: | They propose an infusion approach that partitions a large knowledge graph into smaller sub-graphs and infuses their specific knowledge into various BERT models using lightweight adapters. |
| Outcome: | The proposed approach improves the underlying BERTs and achieves new SOTA performance on six downstream tasks. |
Copied to clipboard
| Challenge: | cuneiform clay tablets were written in 2500 BCE - 100 CE and are a target of extensive transcription and transliteration efforts due to their deterioration. |
| Approach: | They propose to use a masked language modelling task to complete missing text given cuneiform clay tablets written on cuniform signswedges (2500 BCE - 100 CE) they develop models which automatically complete these missing signs based on contextual cues and greedy decoding schemes. |
| Outcome: | The proposed models perform well on missing token prediction (89% hit@5) despite data scarcity (1M tokens), and human evaluations show that they are able to transcribe texts in extinct languages. |
Copied to clipboard
| Challenge: | Existing methods to fine-tune a language model with a large corpus in a general domain are suboptimal for downstream data when domain discrepancy exists. |
| Approach: | They propose to consider the pretrained vocabulary as an optimizable parameter . they add domain specific vocabulary based on a tokenization statistic . their method achieved consistent performance improvements on diverse domains . |
| Outcome: | The proposed method achieves consistent performance improvements on diverse domains. |
Copied to clipboard
| Challenge: | Recent research has explored how models rely on spurious correlations and how counterfactual data augmentation (CDA) can mitigate such issues. |
| Approach: | They propose a context-aware methodology which takes into account the impact of secondary attributes on the model’s predictions and increases sensitivity for secondary attributes over reweighted counterfactually augmented data. |
| Outcome: | The proposed approach improves sliced accuracy on the original dataset by 7% compared to existing methods and provides guidelines to extend this to other tasks. |
Copied to clipboard
| Challenge: | Statistical language modeling and translation with transformers have found many successful applications in program understanding and generation tasks. |
| Approach: | They propose an architecture-independent approach for leveraging syntactic hierarchies of source code . they use syntax trees to extract syntak hierarchical structures and integrate them into context window . |
| Outcome: | The proposed approach achieves state-of-the-art in code completion and summarization for Python in the CodeXGLUE benchmark. |
Copied to clipboard
| Challenge: | Existing studies have focused on probing LMs in the general domain but little attention has been given to whether they can be used as domain knowledge bases. |
| Approach: | They propose to use 49K biomedical factual knowledge triples to probe LMs for biomedically . they find that biomedic LM can achieve up to 18.51% Acc@5 on retrieving biomedcial knowledge. |
| Outcome: | The proposed biomedical factual knowledge probing benchmark achieves 18.51% Acc@5 on biomedically-relevant knowledge retrieval. |
Copied to clipboard
| Challenge: | Existing methods for reading order detection are too laborious to annotate large datasets. |
| Approach: | They propose to use a large-scale dataset to annotate reading order information for document images . they use XML metadata to capture the reading order of WORD documents . |
| Outcome: | The proposed model performs almost perfectly in reading order detection and improves both open-source and commercial OCR engines in ordering text lines in their results. |
Copied to clipboard
| Challenge: | Visual Dialog is assumed to require the dialog history to generate correct responses during a dialog. |
| Approach: | They propose an interpretable representation that visually grounds dialog history by constraining the image’s spatial features according to a semantic representation inspired by Question under Discussion. |
| Outcome: | The proposed representation constrains the image’s spatial features according to a semantic representation of the history inspired by the information structure notion of Question under Discussion. |
Copied to clipboard
| Challenge: | Existing methods to learn semantic representations from text are limited to words in context and words in isolation. |
| Approach: | They propose a method to learn meaning representations of words as low-dimensional node embeddings on an underlying graph hierarchy. |
| Outcome: | The proposed model outperforms the state-of-the-art in terms of similarity judgments and concept categorization. |
Copied to clipboard
| Challenge: | Existing systems for action recognition rely on pattern memorization and do not understand the action. |
| Approach: | They propose a multimodal model that leverages visual and textual information to automatically infer the reasons corresponding to an action presented in the video. |
| Outcome: | The proposed model leverages visual and textual information to automatically infer the reasons corresponding to an action presented in the video. |
Copied to clipboard
| Challenge: | Recent work shows that monolingual masked language models learn to represent data-driven notions of language variation. |
| Approach: | They harness genre metadata as a weak supervision signal for targeted data selection in zero-shot dependency parsing. |
| Outcome: | The proposed method outperforms baseline and embedding-based methods for 12 low-resource language treebanks and three of these target languages. |
Copied to clipboard
| Challenge: | Recent advances in cross-lingual transfer methods have enabled significant advances in grammatical processing tasks. |
| Approach: | They examine the extent to which syntactic relations are preserved in translation and parsability in a zero-shot setting. |
| Outcome: | The proposed model is based on a translation task in English and a subset of a standard English RE benchmark translated to Russian and Korean. |
Copied to clipboard
| Challenge: | Distant supervision is not a practical way to perform unsupervised syntactic parsing. |
| Approach: | They propose a technique that uses distant supervision to improve unsupervised constituency parsing by using phrase bracketing. |
| Outcome: | The proposed method improves constituency parsing on English WSJ Penn Treebank by more than 5 F1 compared with full parse tree annotations. |
Copied to clipboard
| Challenge: | Existing work on understanding worldviews and ideological distinctions focuses on political polarization . et al., 2018: a novel method for uncovering complex ideological and worldview characteristics of communities. |
| Approach: | They propose a method to uncover multifaceted ideological differences across multiple axes . they use comments from the largest communities on reddit.com to train word embedding models . |
| Outcome: | The proposed method can uncover complex ideological differences across multiple axes of polarization using over 1B comments from the largest communities on reddit.com representing 40% of Reddit activity. |
Copied to clipboard
| Challenge: | despite progress toward data-driven conversational agents, dialogue models still suffer from issues surrounding safety and offensive language. |
| Approach: | They analyze reddit threads and reddits to determine the stance of offensive dialogue models . they find 42% of human responses agree with toxic comments, compared to 13% with safe comments . |
| Outcome: | The proposed model produces 29% fewer offensive replies than the baseline model. |
Copied to clipboard
| Challenge: | Recent work in open-domain conversational agents has demonstrated that significant improvements in humanness and user preference can be achieved via massive scaling in both pre-training data and model size. |
| Approach: | They combine open-domain dialogue agents with vision models to investigate human preferences and humanness. |
| Outcome: | The proposed model outperforms existing models in multi-modal dialogue while performing as well as its predecessor (text-only) BlenderBot. |
Copied to clipboard
| Challenge: | Existing systems for speech-based dialogs have found the inadequacy of relying on simple classification techniques to accomplish the automation task. |
| Approach: | They propose a Label-Aware BERT Attention Network (LABAN) for zero-shot multi-intent detection by encoding input utterances with BERT and building a label embedded space by considering embedded semantics in intent labels. |
| Outcome: | The proposed approach can detect many unseen intent labels correctly on a few/zero-shot setting, and achieves state-of-the-art performance on five multi-intent datasets in normal cases. |
Copied to clipboard
| Challenge: | a zero-shot dialogue disentanglement solution is difficult due to the need for manual annotation. |
| Approach: | They propose a zero-shot dialogue disentanglement solution using a web dataset . they train a model on the data and fine-tune the model using labeled data . |
| Outcome: | The proposed model achieves a cluster F1 score of 25 without labeling data . it can be used to analyze discourses and to perform response selection . |
Copied to clipboard
| Challenge: | Existing task-oriented dialog datasets do not situate the dialog in the user’s multimodal context. |
| Approach: | They propose to use a dataset to study multimodal task-oriented dialogs in the shopping domain to situate them in the user’s multimodal context. |
| Outcome: | The proposed dataset includes 11K task-oriented user->assistant dialogs (117K utterances) in the shopping domain, grounded in immersive and photo-realistic scenes. |
Copied to clipboard
| Challenge: | Existing models for dialogue rewriting suffer from the robustness issue, i.e., performances drop dramatically when testing on a different dataset. |
| Approach: | They propose a sequence-tagging-based approach that reduces the search space while preserving the core of the task. |
| Outcome: | The proposed model significantly reduces the search space while still covering the core of the task. |
Copied to clipboard
| Challenge: | Existing approaches to deep learning for open-domain dialogue include training end-to-end models to learn various conversational features like emotional content of response, symbolic transitions of dialogue contexts and persona of the agent and the user, among others. |
| Approach: | They propose a probabilistic approach using Markov Random Fields to augment existing deep-learning methods for improved next utterance prediction. |
| Outcome: | The proposed approach significantly improves the performance of existing state-of-the-art retrieval models for open-domain conversational agents. |
Copied to clipboard
| Challenge: | Task-oriented conversational systems often use dialogue state tracking to represent the user’s intentions, which involves filling in values of pre-defined slots. |
| Approach: | They propose a schema-driven prompting approach that provides task-aware history encoding that is used for both categorical and non-categorical slots. |
| Outcome: | The proposed system achieves state-of-the-art performance on MultiWOZ 2.2 and competitive performance on two other benchmarks: MultiWOz 2.1 and M2M. |
Copied to clipboard
| Challenge: | Sign Language Processing is based on linguistic theories of spoken languages and expect either speech or written text as input. |
| Approach: | They propose a new challenge for coreference modeling and Sign Language Processing to solve this problem. |
| Outcome: | The proposed models will be linguistically informed and can address the complexities of the challenge effectively. |
Copied to clipboard
| Challenge: | Amortized or approximate computational methods increase efficiency, but can result in unpredictable performance costs. |
| Approach: | They propose a method that increases computational efficiency while guaranteeing a specifiable degree of consistency with the original model with high confidence. |
| Outcome: | The proposed method improves on four classification and regression tasks and can be used to predict the performance of the proposed model. |
Copied to clipboard
| Challenge: | Recent studies have shown that pre-trained language models can learn well when primed with only a few labeled examples. |
| Approach: | They propose a method that uses task-specific unlabeled data to provide denser supervision during fine-tuning. |
| Outcome: | The proposed approach outperforms GPT-3 on SuperGLUE without any unlabeled data. |
Copied to clipboard
| Challenge: | Unsupervised Data Augmentation (UDA) is a semisupervised learning method that penalizes differences between a model's predictions on unlabeled examples and corresponding 'noised' examples produced via data augmentation. |
| Approach: | They propose to use a consistency loss to penalize differences between models' predictions on unlabeled and unlabed examples to enforce consistency between models and their perturbed counterparts. |
| Outcome: | The proposed method is able to penalize differences between models' outputs on unlabeled and unlabed examples without complex data augmentation. |
Copied to clipboard
| Challenge: | Recent work shows that pre-training in-domain language models can boost performance when adapting to a new domain. |
| Approach: | They propose to combine annotation and pre-training to maximize performance under budget constraints. |
| Outcome: | The proposed approach is based on the annotation cost of three procedural text datasets and pre-training cost of 3 in-domain language models. |
Copied to clipboard
| Challenge: | Commonsense knowledge bases are mostly human-generated and reflect societal biases . a filtering-based approach can reduce the issues in both resources and models but leads to a performance drop . |
| Approach: | They propose a filtering-based approach to mitigating representational harms in ConceptNet and GenericsKB . they propose filtered-based approaches can reduce issues in both resources and models but leads to performance drop . |
| Outcome: | The proposed approach reduces issues in resources and models but leads to performance drop . the paper proposes a filtering-based approach that reduces biases but leaves room for future work . |
Copied to clipboard
| Challenge: | Existing methods to mitigate stereotypical biases by linear projection are too aggressive . existing methods remove bias, but they also erase valuable information from word embeddings . |
| Approach: | They propose a bias-mitigating method that disentangles biased associations between concepts instead of removing concepts wholesale. |
| Outcome: | The proposed method disentangles biased associations between concepts rather than eliminating concepts wholesale. |
Copied to clipboard
| Challenge: | Existing models produce similar contents from homogenized contexts due to the fixed left-to-right sentence order. |
| Approach: | They propose a framework permuting sentence orders to improve content diversity of multi-sentence paragraphs by permutating the sentence orders. |
| Outcome: | The proposed framework produces more diverse outputs with higher quality than existing models. |
Copied to clipboard
| Challenge: | Existing models for text-to-text generation do not explicitly focus on important concepts in the input and output. |
| Approach: | They propose a framework to automatically extract, denoise, and enforce important input concepts as lexical constraints. |
| Outcome: | The proposed framework performs comparably or better than its unconstrained counterpart on automatic metrics and receives better ratings in the human evaluation. |
Copied to clipboard
| Challenge: | Using neural models, paraphrase generation research has shifted to neural methods . a recent study focused on paraphrases, which are used in language understanding tasks . |
| Approach: | They propose to use neural methods to generate fluent, diverse paraphrases from a sentence . they propose to combine large pretrained language models with other mechanisms to generate more advanced paraphrase generation models. |
| Outcome: | This paper examines various approaches to paraphrase generation with a main focus on neural methods. |
Copied to clipboard
| Challenge: | Exposure bias is a central problem for auto-regressive language models (LM) it is believed that teacher forcing would cause test-time generation to be incrementally distorted due to the training-generation discrepancy. |
| Approach: | They propose to quantify the impact of exposure bias in quality, diversity, consistency and consistency by using ground-truth data prefixes instead of prefix generated by the model. |
| Outcome: | The proposed model performs better when the training-generation discrepancy is removed . the model is more robust and self-recovery ability is shown to counter exposure bias. |
Copied to clipboard
| Challenge: | a proposed model for question-answer pairs with self-contained, summary-centric questions and length-constrained, article-summarizing answers is based on suggested question generation in conversational news recommendation systems. |
| Approach: | They propose a model for generating question-answer pairs with self-contained, summary-centric questions and length-constrained, article-summarizing answers. |
| Outcome: | The proposed model captures the central gists of the articles and achieves high answer accuracy. |
Copied to clipboard
| Challenge: | Paraphrase generation has benefited from recent advances in the design of training objectives and model architectures, but previous studies focused on supervised methods that require a large amount of labeled data that is costly to collect. |
| Approach: | They propose a transfer learning approach that enables pre-trained language models to generate high-quality paraphrases in an unsupervised setting. |
| Outcome: | The proposed model performs state-of-the-art on the Quora Question Pair and ParaNMT datasets and is robust to domain shift between the two datasets. |
Copied to clipboard
| Challenge: | a recent study shows that inappropriate language can cause models to output profanity . authors propose a training framework to prevent such outputs from hurting the usability of models . |
| Approach: | proposed training framework eliminates the causes that trigger the generation of profanity . authors propose a framework that leverages a short list of profans to prevent this . |
| Outcome: | a proposed training framework can prevent models from generating profanity . the proposed framework leverages a short list of profanities examples . |
Copied to clipboard
| Challenge: | Experimental results show that JoGANIC outperforms state-of-the-art methods for image caption generation. |
| Approach: | They propose a method to generate descriptive and informative captions for news article images . they leverage the structure of captions to improve the generation quality and guide their representation . |
| Outcome: | The proposed method outperforms state-of-the-art methods on two large-scale datasets. |
Copied to clipboard
| Challenge: | Existing models for paraphrase generation use fixed syntactic structures for all input sentences. |
| Approach: | They propose to add syntactical control to a pretrained language model to generate fluent paraphrases using a retrieval-based selection module. |
| Outcome: | The proposed model achieves state-of-the-art on semantic preservation and syntactic conformation on two benchmark datasets with ground-truth syntaktic control from human-annotated exemplars. |
Copied to clipboard
| Challenge: | a number of negative effects exist when NLG systems are not grounded to a specific input text. |
| Approach: | They argue that NLG systems should focus on making use of additional context . they argue that relevance should be thought of as a crucial tool for user-oriented text-generating tasks . |
| Outcome: | The proposed approach is more of the rule than the exception, the authors argue . they argue that value-sensitive design represents a crucial path forward . |
Copied to clipboard
| Challenge: | Event schemas encode knowledge of stereotypical structures of events and their connections . previous work on event schema induction focuses on atomic events or linear temporal sequences . |
| Approach: | They propose a Temporal Complex Event Schema: a graph-based schema representation that encompasses events, arguments, temporal connections and argument relations. |
| Outcome: | The proposed model outperforms existing models on HITS@1 by 17.8%. |
Copied to clipboard
| Challenge: | Event mentions in text correspond to real-world events of varying degrees of granularity . task of subevent detection aims to resolve this granulem issue by recognizing membership of events . |
| Approach: | They propose a task of event-based text segmentation as an auxiliary task to improve learning for subevent detection. |
| Outcome: | The proposed method outperforms baseline methods on subevent detection, HiEve and IC datasets while achieving decent performance on EventSeg prediction. |
Copied to clipboard
| Challenge: | Named entity recognition (NER) is a fundamental step in scientific literature analysis to build AI-driven systems for molecular discovery, synthetic strategy designing, and manufacturing. |
| Approach: | They propose an ontology-guided method for fine-grained named entity recognition (NER) it leverages the chemistry type ontologies to generate distant labels with flexible KB-matching . |
| Outcome: | The proposed method significantly outperforms the state-of-the-art methods with a .25 absolute F1 improvement. |
Copied to clipboard
| Challenge: | Academic neural models for coreference resolution (coref) are typically trained on OntoNotes and model improvements are benchmarked on that dataset. |
| Approach: | They aim to quantify transferability of coref models based on the number of annotated documents available in the target dataset. |
| Outcome: | The proposed model improvements are consistent with the state-of-the-art results on PreCo. |
Copied to clipboard
| Challenge: | Document-level entity-based extraction (EE) tasks extract entity-centric information from unstructured text across multiple sentences. |
| Approach: | They propose a generative framework for two document-level EE tasks: role-filler entity extraction (RE) and relation extraction ( RE). |
| Outcome: | The proposed framework captures cross-entity dependencies and avoids exponential computation complexity of identifying N-ary relations. |
Copied to clipboard
| Challenge: | Existing training data for event detection are too expensive to achieve in real applications where novel event types emerge . Typical ED systems require labeled data for each predefined event type, but only a few examples are available. |
| Approach: | They propose to introduce cross-task prototypes to model relationships between training tasks in few-shot learning for event detection. |
| Outcome: | The proposed model improves on three few-shot learning datasets. |
Copied to clipboard
| Challenge: | Traditional supervised Information Extraction (IE) methods can extract structured knowledge elements from unstructured data, but it is limited to a pre-defined target ontology. |
| Approach: | They propose a new lifelong event detection framework that is generalizable to other IE tasks and updates old knowledge with new event types’ mentions using a self-training loss. |
| Outcome: | The proposed framework outperforms baselines with a 5.1% gain in the F1 score and can boost the F2 score for over 30% on some new long-tail rare event types with few training instances. |
Copied to clipboard
| Challenge: | Prior work on information extraction tends to focus on binary relations within sentences . practical applications often require extracting complex relations across large text spans . |
| Approach: | They propose to decompose document-level relation extraction into relation detection and argument resolution, taking inspiration from Davidsonian semantics. |
| Outcome: | The proposed method outperforms state-of-the-art methods in biomedical machine reading for precision oncology by 20 absolute F1 points. |
Copied to clipboard
| Challenge: | Existing methods augment input sequence with token replacement, assuming annotations on the replaced positions are unchanged. |
| Approach: | They propose to use paraphrasing to enhance unsupervised consistency training by replacing tokens with augmented data. |
| Outcome: | The proposed method is especially effective when annotations are limited. |
Copied to clipboard
| Challenge: | Existing work on fine-grained entity typing (FET) relies on knowledge bases as distant supervision, but lack of or incompleteness of KB can hinder training. |
| Approach: | They propose a two-step framework that trains FET models without accessing any knowledge base. |
| Outcome: | The proposed framework achieves competitive performance with respect to the models trained on the original KB-supervised datasets. |
Copied to clipboard
| Challenge: | Existing studies on cross-lingual entity alignment under adversarial attacks have not been conducted. |
| Approach: | They propose to use adversarial attack techniques to perturb cross-lingual entity alignment under adversarials. |
| Outcome: | The proposed model hides the attacked entities in dense regions in two KGs, and reduces the gradient vanishing issues in the process of adversarial attacks for further improving the attack effectiveness. |
Copied to clipboard
| Challenge: | Recent studies have shown that few-shot relation classification models can be used to extract any relation of interest from a collection of text with only a few example instances. |
| Approach: | They propose to modify the training routine to encourage models to better discriminate between relations involving similar entity types. |
| Outcome: | The proposed models outperform human models on relation extraction tasks while relying on entity type information. |
Copied to clipboard
| Challenge: | Existing methods for named entity recognition focus on augmenting in-domain data in low-resource scenarios where annotated data is limited. |
| Approach: | They propose a neural architecture to transform data from high-resource to low-resourced domains by learning the patterns in the text that differentiate them. |
| Outcome: | The proposed approach improves on high-resource domain representations over high- and low-resourced domains. |
Copied to clipboard
| Challenge: | Pre-trained language models (PLMs) are used for diverse NLP tasks such as Information Extraction, Sentiment Analysis and Question/Answering. |
| Approach: | They propose to add medical knowledge to pre-trained language models to facilitate clinical relation extraction using a large text corpus. |
| Outcome: | The proposed model outperforms the state-of-the-art systems on the benchmark i2b2/VA 2010 clinical relation extraction dataset. |
Copied to clipboard
| Challenge: | Pre-trained language models (PTLMs) have achieved noticeable success on many NLP tasks, but struggle for tasks that require event temporal reasoning. |
| Approach: | They propose a continual pre-training approach that equips PTLMs with targeted knowledge about event temporal relations by focusing on masked-out event and temporal indicators and discriminating sentences from their corrupted counterparts. |
| Outcome: | The proposed framework improves the PTLMs’ fine-tuning performances across five relation extraction and question answering tasks and achieves new or on-par state-of-the-art in most of our downstream tasks. |
Copied to clipboard
| Challenge: | Recent information extraction approaches can easily overfit noisy labels and suffer from performance degradation. |
| Approach: | They propose a co-regularization framework for entity-centric information extraction that optimizes neural models with task-specific losses and regularizes them to generate similar predictions based on agreement loss. |
| Outcome: | The proposed framework is optimized with task-specific losses and generates similar predictions based on agreement loss. |
Copied to clipboard
| Challenge: | a lack of large training datasets hampers machine learning-based prediction of material properties . relevant measurements and information exist only in unstructured formats such as the published literature . |
| Approach: | They propose a framework for automatic property extraction using material solubility as the target property. |
| Outcome: | The proposed framework extracts solubility data from scientific literature and compares it with other frameworks. |
Copied to clipboard
| Challenge: | Existing methods for Event Detection (ED) do not encode long-range document-level context . e.g., BERT cannot encode long text-level contextual information . |
| Approach: | They propose a method to model document-level context for Event Detection using transformer-based language models. |
| Outcome: | The proposed model can predict event prediction of target sentence in document-level context . the proposed model is effective on multiple benchmark datasets . |
Copied to clipboard
| Challenge: | Existing approaches to crosslingual Relation and Event Extraction (REE) suffer from monolingual bias due to training of models on source language data. |
| Approach: | They propose to use unlabeled data in target language to aid alignment of crosslingual representations by fooling a language discriminator. |
| Outcome: | The proposed method significantly advances the state-of-the-art in crosslingual REE tasks. |
Copied to clipboard
| Challenge: | Existing event extraction methods require predefined event types and their annotations to learn event extractors. |
| Approach: | They propose to represent each event type as a cluster of predicate sense, object head> pairs. |
| Outcome: | The proposed method can discover salient and high-quality event types on three datasets from different domains. |
Copied to clipboard
| Challenge: | Existing approaches to Named Entity Recognition (NER) are limited in labeled resources and domain shift. |
| Approach: | They propose a progressive domain adaptation knowledge distillation approach to adapt high-resource domains to low-resourced target domains by employing three components to achieve superior domain adaptability. |
| Outcome: | The proposed approach can adapt high-resource domains to low-resourced target domains even if they are diverse in terms and writing styles. |
Copied to clipboard
| Challenge: | Document retrieval systems often use two styles of neural network models . dual encoder models are used for retrieval and deep re-ranking, while cross-attention models are typically used for shallow reranking. |
| Approach: | They propose a dual encoder and cross-attention neural network architectures that combine query and document representations to optimize retrieval accuracy. |
| Outcome: | The proposed architecture trades off retrieval accuracy with joint computation and offline document storage cost. |
Copied to clipboard
| Challenge: | Existing question answering datasets lack diversity in gender, profession, and nationality. |
| Approach: | They focus on how well QA models generalize across demographic subsets . english-language QA datasets mostly ask about US men from a few professions - this is problematic because most English speakers are not from the US or UK . |
| Outcome: | The proposed model accuracy is lower for people based on gender, profession, and nationality, but there is more variation on professions (question topic) and question ambiguity. |
Copied to clipboard
| Challenge: | Recent work proposes lightweight updates to improve commonsense reasoning models . fine-tuning can cause models to overfit to task-specific data and forget knowledge gained during training . |
| Approach: | They propose to use lightweight models to update pre-trained language models to learn commonsense background knowledge. |
| Outcome: | The proposed models learn from commonsense reasoning datasets, but they are overfitted and limited generalized. |
Copied to clipboard
| Challenge: | Using feed-forward layers, we show that the learned patterns are human-interpretable, and that lower layers tend to capture shallow patterns, while upper layers learn more semantic ones. |
| Approach: | They propose that feed-forward layers in transformer-based language models operate as key-value memories where each key correlates with textual patterns in the training examples and each value induces a distribution over the output vocabulary. |
| Outcome: | The proposed model is based on key-value memories with a key-level correlation with the training examples and a distribution over the output vocabulary. |
Copied to clipboard
| Challenge: | Recent research in interpretability of neural models has yielded numerous token attribution techniques, but it is hard to evaluate whether these explanations are faithful. |
| Approach: | They propose to use pairwise attributions to connect outputs to high-level model behavior to examine how well different attribution techniques align with this assumption on realistic counterfactuals in the case of reading comprehension (RC). |
| Outcome: | The proposed methods are better suited to RC than token-level attributions across different RC settings, and the best performance comes from a modification that was proposed to an existing pairwise attribution method. |
Copied to clipboard
| Challenge: | Using RNN and transformer language models, we show consistent generalization in out-of-distribution contexts. |
| Approach: | They propose two idealized models of generalization in next-word prediction . they show that neural language models interpolate between these two forms of generalisation . |
| Outcome: | The proposed models exhibit consistent generalization in out-of-distribution contexts. |
Copied to clipboard
| Challenge: | Recent advances in NLP have been made by learning representations that transform complex tasks into simple classification tasks. |
| Approach: | They propose a method to evaluate the compatibility between representations and tasks by fitting text features to specific characteristics of text datasets. |
| Outcome: | The proposed model provides a calibrated, quantitative measure of the difficulty of a classification-based NLP task. |
Copied to clipboard
| Challenge: | In this work, we present a method for imparting human-like rationalization to a stance detection model using crowdsourced annotations on a small fraction of the training data. |
| Approach: | They propose a method for imparting human-like rationalization to a stance detection model using crowdsourced annotations on a small fraction of the training data. |
| Outcome: | The proposed method improves the reasoning of a state-of-the-art classifier in a data-scarce setting at no cost in predictive performance. |
Copied to clipboard
| Challenge: | Multi-task learning with transformer encoders (MTL) has emerged as a powerful technique to improve performance on closely-related tasks for both accuracy and efficiency. |
| Approach: | They propose a multi-task learning technique that uses transformer encoders to improve performance on closely-related tasks. |
| Outcome: | The proposed method performs better on five NLP tasks than single-task learning on similar tasks. |
Copied to clipboard
| Challenge: | Using latent optimization and Shapley values, we generate a set of minimal modifications to the text to change the classifier's prediction. |
| Approach: | They propose to generate a counterfactual by making minimal modifications to the text to change the model's prediction. |
| Outcome: | The proposed approach achieves favorable performance compared to white-box and black-box baselines using human and automatic evaluations. |
Copied to clipboard
| Challenge: | Contextualized representations have been used in various NLP tasks, but their nature remains a mystery. |
| Approach: | They propose to use a property to estimate the power of contextualized representations . they show that the average representation shares almost the same direction as the first principal component . |
| Outcome: | The proposed representations share the same direction as the first principal component . the results suggest that the property is intrinsic to the distribution of representations . |
Copied to clipboard
| Challenge: | Prior work has shown that structural supervision helps English language models learn generalizations about syntactic phenomena such as subject-verb agreement. |
| Approach: | They train LSTMs, Recurrent Neural Network Grammars, Transformer language models, and Transformer-parameterized generative parsing models on Mandarin Chinese datasets. |
| Outcome: | The proposed models learn aspects of Mandarin Chinese grammar that assess syntactic and semantic relationships. |
Copied to clipboard
| Challenge: | A key problem in multi-task learning (MTL) research is how to select high-quality auxiliary tasks automatically. |
| Approach: | They propose an automatic auxiliary task selection method based on gradient calculation in Transformer-based models that improves MT-DNN performance. |
| Outcome: | The proposed method improves MT-DNN performance on 8 natural language understanding (GLUE) tasks, while costing less than AUTOSEM and comparable GPU consumption. |
Copied to clipboard
| Challenge: | Recent advances in knowledge base construction techniques focus on the acquisition of positive (true) KB statements, but negative (false) statements are important for discriminative reasoning. |
| Approach: | They propose a framework that ranks potential negatives in commonsense KBs using a contextual language model. |
| Outcome: | The proposed framework ranks negatives in commonsense KBs using a language model . it yields positives that are more grammatical, coherent, and informative . |
Copied to clipboard
| Challenge: | Several noise-robust losses have been proposed and evaluated on tasks in computer vision, but they use a single dataset-wise hyperparamter to control the strength of noise resistance. |
| Approach: | They propose to change single dataset-wise hyperparameters of noise resistance to be instance-wise. |
| Outcome: | The proposed frameworks increase noise-robustness on noisy and corrupted NLP datasets. |
Copied to clipboard
| Challenge: | Multilingual neural machine translation (MNMT) learns to translate multiple language pairs with a single model, but the data imbalance hinders it from performing uniformly across language pairs. |
| Approach: | They propose a distributionally robust optimization objective which minimizes the worst-case expected loss over the set of language pairs. |
| Outcome: | The proposed learning objective outperforms baseline methods on three sets of languages and shows that it is cost-effective and efficient. |
Copied to clipboard
| Challenge: | Existing work has relied on English dev data to select among models that are fine-tuned with different learning rates, number of steps and other hyperparameters, often resulting in suboptimal choices. |
| Approach: | They propose a machine learning approach that uses the fine-tuned model’s internal representations to predict its cross-lingual capabilities. |
| Outcome: | The proposed model selects better than English validation data across twenty five languages, including eight low-resource languages, and often achieves comparable results to model selection using target language development data. |
Copied to clipboard
| Challenge: | a large number of end-to-end systems are needed for many tasks in natural language processing. |
| Approach: | They propose a continual few-shot learning task where a system is asked to correct mistakes with a few training examples. |
| Outcome: | The proposed task compares two NLI and one sentiment analysis datasets with baselines from diverse paradigms. |
Copied to clipboard
| Challenge: | Non-parametric neural language models (NLMs) learn text distributions by memorizing training data points. |
| Approach: | They propose to use an external datastore to learn from a non-parametric language model. |
| Outcome: | The proposed methods achieve up to a 6x speed-up in inference speed while retaining comparable performance. |
Copied to clipboard
| Challenge: | Recent advances in NLP demonstrate the effectiveness of applying large-scale pre-trained language models to downstream tasks. |
| Approach: | They propose a method that uses task augmentation to fine-tune unlabeled data. |
| Outcome: | The proposed approach improves sample efficiency across 12 few-shot benchmarks. |
Copied to clipboard
| Challenge: | Existing approaches to solve domain shifts in NLP tasks require additional pre-training . current approaches focus on the downstream corpus when it is small, but are not effective . |
| Approach: | They propose a task-adapted pre-training framework that can be used when the downstream corpus is too small for additional pre-tuning. |
| Outcome: | The proposed framework outperforms baseline methods on biomedical, computer science, news, and movie reviews tasks. |
Copied to clipboard
| Challenge: | Existing methods for obtaining adversarial examples are difficult with text data. |
| Approach: | They propose a gradient-based adversarial attack against transformer models that searches for a distribution of adversarials parameterized by a continuous-valued matrix. |
| Outcome: | The proposed attack outperforms existing methods on a variety of natural language tasks with matching imperceptibility. |
Copied to clipboard
| Challenge: | Currently, the Transformer is the de facto architecture of choice for processing sequential data. |
| Approach: | They evaluate the Transformer architecture and its modifications in a shared experimental setting . they conjecture that performance improvements may strongly depend on implementation details . |
| Outcome: | The proposed improvements do not significantly improve performance, the authors find . the proposed improvements are either developed in the same codebase or are minor changes . |
Copied to clipboard
| Challenge: | a new method to learn compositional structured models is needed . end-task supervision provides only a weak indirect signal on values the latent decisions should take. |
| Approach: | They propose a way to leverage paired examples that provide stronger cues for learning latent decisions . they use a DROP dataset to acquire paired questions that provide strong cue signals . |
| Outcome: | The proposed approach improves compositional question answering on a DROP dataset. |
Copied to clipboard
| Challenge: | Recent efforts to improve sentence representation learning have a common weakness . siamese or triplet loss only learns from individual sentence pairs or tripletes . |
| Approach: | They propose a discrimination-based approach to bridge entailment and contradiction understanding with categorical concept encoding. |
| Outcome: | The proposed method outperforms the state-of-the-art method on downstream tasks . it improves 10%–13% on clustering tasks and 5%–6% on STS tasks compared with the previous method . |
Copied to clipboard
| Challenge: | Recent work shows gains from pre-training and fine-tuning that are multi-task . but it can be difficult to know which intermediate tasks will best transfer . |
| Approach: | They propose a large-scale learning stage for pre-finetuning between pre-training and fine-tun. |
| Outcome: | The proposed model improves performance on pretrained discriminators and generation models on a wide range of tasks while improving sample efficiency during fine-tuning. |
Copied to clipboard
| Challenge: | Meta-learning considers learning as an efficient learning process that can leverage its past experience to accurately solve new tasks. |
| Approach: | They propose to provide task distributions for meta-learning by considering self-supervised tasks automatically proposed from unlabeled text to enable large-scale meta- learning in NLP. |
| Outcome: | The proposed distributions show that human learning models perform better on the few-shot benchmark than previous methods. |
Copied to clipboard
| Challenge: | Language agnostic and semantic-language information isolation is an emerging research direction for multilingual representations models. |
| Approach: | They propose a method that factors out language identity information from semantic related components in multilingual representations pre-trained on monolingual data. |
| Outcome: | The proposed method improves cross-lingual transfer performance on weak alignment models. |
Copied to clipboard
| Challenge: | Cross-lingual language models house representations for many different languages in the same space. |
| Approach: | They investigate linguistic and non-linguistic factors affecting sentence-level alignment in cross-lingual pretrained language models for 101 languages and 5,050 language pairs. |
| Outcome: | The results show that word order agreement and agreement in morphological complexity are strongest predictors of cross-linguality. |
Copied to clipboard
| Challenge: | Existing training data is limited for languages other than English, so is the performance of the developed parsers. |
| Approach: | They propose to apply a pre-trained multilingual model to Italian, German and Dutch parsers where only a small number of manually annotated parses are available. |
| Outcome: | The proposed model improves on six parsers in English and Italian, German and Dutch, with the addition of universal dependency relations and universal POS tags as model-agnostic features. |
Copied to clipboard
| Challenge: | Existing systems for simultaneous translation are still trained on full-sentence bitexts due to the abundance of unnecessary long-distance reorderings. |
| Approach: | They propose to rewrite target side of existing full-sentence corpora into simultaneous-style translation by adding generated pseudo-references to the target side. |
| Outcome: | Experiments on ZhEn and JaEn simultaneous translation show that the proposed method improves on existing full-sentence corpora. |
Copied to clipboard
| Challenge: | Sentence-level Quality estimation (QE) is traditionally a regression task . but large multilingual contextualized language models are expensive and infeasible for real-world applications. |
| Approach: | They evaluate several model compression techniques for QE and find they are inefficient . they argue that a full model parameterization is required to achieve SoTA results . |
| Outcome: | The proposed models are poorly expressive in a regression task, the authors argue . they show that reframing QE as a classification problem and evaluating models would improve their performance in real-world applications. |
Copied to clipboard
| Challenge: | a large corpus covering 22 Turkic languages is included in this paper . low-resource MT evaluation has traditionally focused on European languages due to limitations of available technology and resources. |
| Approach: | They present a case study of the practical application of MT in the Turkic language family . they propose to realize the gains of NMT for Turkic languages under high-resource to extremely low-resourced scenarios. |
| Outcome: | The proposed study shows that the new methods can be used in the Turkic language family . the results highlight bottlenecks in building competitive systems . |
Copied to clipboard
| Challenge: | Word embeddings are powerful representations that form the foundation of many natural language processing architectures. |
| Approach: | They explore word embedding stability in a wide range of languages to gain insight into their stability. |
| Outcome: | The proposed results provide insights into word embedding stability in English and other languages. |
Copied to clipboard
| Challenge: | Current approaches to incorporating terminology constraints in machine translation (MT) typically assume that the constraint terms are provided in their correct morphological forms. |
| Approach: | They propose a framework for incorporating lemma constraints in machine translation . they use a cross-lingual inflection module that inflects the target lemmo constraints based on the source context. |
| Outcome: | The proposed framework outperforms existing methods with lower training costs and linguistic knowledge in domain adaptation and low-resource MT settings. |
Copied to clipboard
| Challenge: | Recent work shows that supervised neural machine translation models scale like a power law with the amount of training data and number of non-embedding parameters in the model. |
| Approach: | They show that cross-entropy loss of supervised neural machine translation models scales like a power law with the amount of training data and number of non-embedding parameters in the model. |
| Outcome: | The proposed model can predict BLEU and ROI of labeling data in low-resource language pairs. |
Copied to clipboard
| Challenge: | GE3 is a data augmentation protocol that can be used to increase text examples from one class onto another. |
| Approach: | They propose a data augmentation protocol that extrapolates the hidden space distribution of text examples from one class onto another to investigate whether this bias is valid for data augmented. |
| Outcome: | The proposed protocol improves on three text classification datasets for various data imbalance scenarios. |
Copied to clipboard
| Challenge: | Existing approaches to generate paraphrases with weak supervision are limited in real-world scenarios due to the lack of coherent and controllable generated paraphrase. |
| Approach: | They propose a method to generate high-quality paraphrases with weak supervision . they obtain abundant weakly-labeled parallel sentences via retrieval-based pseudo paraphrase expansion . |
| Outcome: | The proposed approach achieves significant improvements over existing methods and is even comparable in performance with supervised state-of-the-arts. |
Copied to clipboard
| Challenge: | a large number of medical encounters need to be coded everyday due to long document sets and large label set. |
| Approach: | They propose a convolutional attention network for multi-label document classification problem . they use convolution-based encoders and convolution networks to aggregate information across documents . |
| Outcome: | The proposed model outperforms prior best model and multilingual Transformer model on a widely used dataset in the medical domain. |
Copied to clipboard
| Challenge: | Recent work learns contextual representations of source code by reconstructing tokens from their context. |
| Approach: | They propose a contrastive pre-training task that learns code functionality, not form . they propose scalable compilers that can generate variants of a program . |
| Outcome: | The proposed task outperforms RoBERTa on an adversarial code clone detection benchmark by 39% AUROC. |
Copied to clipboard
| Challenge: | Pretrained language models have improved writing assistance functions such as autocomplete, but more complex and controllable writing assistants have yet to be explored. |
| Approach: | They build an intent-guided authoring assistant that follows fine-grained author directives by specifying different writing intents. |
| Outcome: | The proposed system generates output satisfying the author's intent and can be rephrased to their liking. |
Copied to clipboard
| Challenge: | Existing approaches to generate arithmetic math word problems are invalid or have unsatisfactory language quality. |
| Approach: | They propose a method for automatically generating arithmetic math word problems from equations and context. |
| Outcome: | The proposed approach improves language quality and mathematical validity on three real-world MWP datasets. |
Copied to clipboard
| Challenge: | Various deep learning models have been successfully employed for this type of NLP task of text classification. |
| Approach: | They propose a mixed-domain transfer learning approach that only captures local context and exhibits poor generalization. |
| Outcome: | The proposed model captures local and global contexts, but lacks generalization . a combination of shallow network-based domain-specific models and convolutional neural networks can extract local and globally context directly from the target data in a hierarchical fashion, enabling it to offer a more generalizable solution. |
Copied to clipboard
| Challenge: | Health and medical researchers often give clinical and policy recommendations to inform health practice and public health policy. |
| Approach: | They developed a BERT-based prediction model that can predict whether a sentence gives strong advice, weak advice, or not. |
| Outcome: | The proposed model can predict whether a sentence gives strong advice, weak advice, or not with a macro-averaged F1 score of 0.93. |
Copied to clipboard
| Challenge: | Multiple choice questions can be graded automatically, but automated short answer grading is time-consuming and has bias and errors. |
| Approach: | They propose a Semantic Feature-wise transformation Relation Network that captures relational knowledge among the questions, reference answers or rubrics and labeled student answers. |
| Outcome: | The proposed model has up to 11% performance improvement over state-of-the-art approaches on the benchmark SemEval-2013 datasets, and surpasses custom approaches designed for a Kaggle challenge. |
Copied to clipboard
| Challenge: | Scientific, engineering, and technological (SET) innovations drive many positive advances in our modern economy, society, and life. |
| Approach: | They propose a new metric that uses the content of the paper as a source of distant-supervision to quantify how much the cited-node informs the citing-n node. |
| Outcome: | The proposed method achieves up to 103% improvement over the second-best method. |
Copied to clipboard
| Challenge: | Existing methods to improve NLU are laborintensive and expensive. |
| Approach: | They propose a scalable and automatic approach to improving NLU in a large-scale conversational AI system by leveraging implicit user feedback. |
| Outcome: | The proposed framework improves NLU in a large-scale conversational AI system across 10 domains. |
Copied to clipboard
| Challenge: | Recent approaches to multi-hop Reading Comprehension (RC) have greatly improved its explainability, models ability to explain their own answers. |
| Approach: | They propose to generate a question-focused abstractive summary of input paragraphs and feed it to an RC system. |
| Outcome: | The proposed explanation generates more compact explanations than an extractive explainer with limited supervision while maintaining sufficiency. |
Copied to clipboard
| Challenge: | Existing pre-trained models need fine-tuning on tens of thousands of examples to achieve good results. |
| Approach: | They propose a framework that leverages pre-trained text-to-text models and aligns them with their pre-training framework. |
| Outcome: | The proposed framework outperforms the XLM-Roberta-large on multiple QA benchmarks and is applicable to multilingual situations. |
Copied to clipboard
| Challenge: | Existing neural firststage retrieval models overcome lexical gap issue by projecting query and document to a shared dense space. |
| Approach: | They propose a multi-stage framework for neural passage retrieval using synthetic data, negative sampling, and fusion techniques. |
| Outcome: | The proposed framework improves retrieval accuracy and enhances the negative contrast in both stages. |
Copied to clipboard
| Challenge: | Taking the exam closed book, but having read the textbook, yields at best minor improvement (56%), suggesting that the PTLM may not have “understood” the textbook (or perhaps misundersttoo the questions). |
| Approach: | They propose to use pre-trained language models to answer questions from introductory college textbooks and hundreds of true/false statements based on review questions written by the authors. |
| Outcome: | The proposed task includes two college-level introductory texts in the social sciences (American Government 2e) and humanities (U.S. History). |
Copied to clipboard
| Challenge: | Existing pre-training methods only harvest learning signals from local contexts of naturally occurring texts . ReasonBert provides a method for reasoning over long-range relations and multiple, possibly hybrid contexts. |
| Approach: | They propose a method that augments language models with the ability to reason over long-range relations and multiple, possibly hybrid contexts. |
| Outcome: | The proposed method significantly improves sample efficiency over strong baselines. |
Copied to clipboard
| Challenge: | Prior work has focused on training one network on multiple datasets to build a model that performs well on all of the training datasets and generalizes and transfers better to new datasets. |
| Approach: | They combine multiple reading comprehension datasets to build a multi-dataset question answering model with an ensemble of single-data set experts. |
| Outcome: | The proposed model outperforms baseline models in in-distribution accuracy and generalization and transfer performance. |
Copied to clipboard
| Challenge: | Open-domain question answering has exploded in popularity due to the success of dense retrieval models. |
| Approach: | They construct a set of simple, entity-rich questions based on facts from Wikidata and test their models against supervised datasets. |
| Outcome: | The proposed model outperforms sparse retrieval methods on open-domain question answering datasets by a large margin. |
Copied to clipboard
| Challenge: | Question Answering models typically use retrieval and reasoning components to identify relevant information for reasoning. |
| Approach: | They propose a retrieval parameterization method that marginalizes unanswerable queries . they show that marginalization allows a model to mitigate false negatives in annotations . |
| Outcome: | The proposed model improves on two multi-document question answering datasets and shows that marginalization improves performance. |
Copied to clipboard
| Challenge: | Existing work treats document-grounded dialogue modeling as a machine reading comprehension task based on a single document or passage. |
| Approach: | They propose a task and dataset for modeling goal-oriented dialogues grounded in multiple documents. |
| Outcome: | The proposed task and dataset address realistic scenarios where goal-oriented dialogues involve multiple topics and hence are grounded on different documents. |
Copied to clipboard
| Challenge: | Abstractive summarization is the process of generating a condensed version of a given conversation while preserving the most salient aspects. |
| Approach: | They propose to use a dataset to analyze code-switched conversations in Hindi and English to summarize them. |
| Outcome: | The proposed dataset contains over 6,800 code-switched conversations and their corresponding human-annotated summaries in English (En) and Hi-En. |
Copied to clipboard
| Challenge: | Several past efforts have created Split and Rephrase training sets, which consist of long, complex input sentences paired with multiple shorter sentences that preserve the meaning of the input sentence. |
| Approach: | They propose a new dataset and a model for this task by extracting 1-2 sentence alignments from bilingual parallel corpora and using machine translation to convert both sides of the corpus into the same language. |
| Outcome: | The proposed model can perform a wider variety of split operations and improve upon previous state-of-the-art approaches in automatic and human evaluations. |
Copied to clipboard
| Challenge: | Knowledge Graph Completion (KGC) attempts to learn missing links from subsets. |
| Approach: | This survey/position paper discusses ways to improve coverage of resources such as WordNet. |
| Outcome: | The proposed method improves WordNet coverage by reducing the number of words in the sample and reducing unbalanced corpora. |
Copied to clipboard
| Challenge: | Existing methods to learn sentence representations on unlabeled corpora are difficult and expensive to obtain, making it hard to cover many domains and languages. |
| Approach: | They propose a method to train sentence representations on large unlabeled corpora by conditioning on the encoded vectors of adjacent sentences. |
| Outcome: | The proposed method outperforms existing models on SentEval and can be extended to a broad range of languages and domains. |
Copied to clipboard
| Challenge: | Recent advances in neural architectures and pre-trained representations have greatly improved the performance of fully-supervised semantic role labeling (SRL) but there are limitations in the availability of supervised training data. |
| Approach: | They propose to leverage syntactic dependencies to facilitate cross-lingual transfer by annotating predicate-argument structures in text. |
| Outcome: | The proposed model can be extended to other languages with limited training data. |
Copied to clipboard
| Challenge: | In argumentation theory, an enthymeme is defined as incomplete argument found in discourse . encoding discourse-aware commonsense improves the quality of the generated implicit premises . |
| Approach: | They propose a task that generates an implicit premise in an enthymeme using commonsense . they use a narrative text dataset to analyze the quality of the generated premises . |
| Outcome: | The proposed model outperforms baseline models on three datasets. |
Copied to clipboard
| Challenge: | Existing neural models lack systematic compositionality in learning symbolic structures . existing models lack this ability in learning symbols, despite being able to understand complex structures. |
| Approach: | They propose to use auxiliary sequence prediction tasks to train a Transformer model to understand compositional symbolic structures of input data. |
| Outcome: | The proposed model improves on the SCAN compositionality challenge, with only 418 (5%) training instances, and achieves 97.8% accuracy on the MCD1 split. |
Copied to clipboard
| Challenge: | Developing models that can make useful inferences from natural language premises has been a core goal in artificial intelligence since the field's early days. |
| Approach: | They propose a method for building models to generate deductive inferences from diverse natural language inputs without direct human supervision. |
| Outcome: | The proposed model is more accurate and flexible than baseline systems. |
Copied to clipboard
| Challenge: | Recent work shows that pre-trained sequence-to-sequence Transformer models are effective in predicting linearized Abstract Meaning Representation graphs. |
| Approach: | They propose a structure-aware transition-based approach to AMR parsing that integrates general pre-trained sequence-to-sequence language models with a structured transition set. |
| Outcome: | The proposed approach retains the desirable properties of previous approaches while reaching the new parsing state of the art for AMR 2.0. |
Copied to clipboard
| Challenge: | Existing literature suggests that a person forms a mental model of the problem scenario before answering questions. |
| Approach: | They propose to have a model first create a graph of relevant influences and leverage that graph as an additional input when answering a defeasible query. |
| Outcome: | The proposed model achieves state-of-the-art on three different defeasible reasoning datasets. |
Copied to clipboard
| Challenge: | Existing aspects target sentiment classification models are not trainable if annotated data are not available. |
| Approach: | They propose an approach that solves ATSC with natural language prompts by 24.13 accuracy points and 33.14 macro F1 points. |
| Outcome: | The proposed model outperforms supervised SOTA approaches under few-shot scenarios and under supervised settings, especially for few-shot cases. |
Copied to clipboard
| Challenge: | Using pre-trained models, people use different styles to express their interpersonal goal and attitude in their communication. |
| Approach: | They use a dataset to collect lexicon usages across styles using two lenses: human perception and machine word importance. |
| Outcome: | The proposed model can predict human perception and machine word importance based on a popular style classifier like BERT . human- and machine-identified words share significant overlap for some styles . |
Copied to clipboard
| Challenge: | stance detection is a method to determine whether a text author is in favor of, against or neutral toward a specific target. |
| Approach: | They propose a method that applies instance-specific temperature scaling to the teacher and student predictions. |
| Outcome: | The proposed method outperforms the state-of-the-art on all datasets and on multiple datasets. |
Copied to clipboard
| Challenge: | Existing methods address this issue by introducing an auxiliary task such as visual grounding, cycle consistency, or debiasing. |
| Approach: | They propose a data augmentation pipeline to turn “known” knowledge into training examples for VQA. |
| Outcome: | The proposed model can handle multi-modal information and is based on human-annotated examples. |
Copied to clipboard
| Challenge: | Existing studies have focused on the phrase grounding ability of pretrained vision-and-language models, but it is unclear how they can be used for phrase ground. |
| Approach: | They propose to extract phrase-region pairs from pre-trained vision-and-language embeddings and propose four fine-tuning objectives to improve model phrase grounding ability using image-caption data without any supervised grounding signals. |
| Outcome: | The proposed model outperforms baseline models in weakly-supervised and supervised phrase grounding settings on two representative datasets and shows that it is possible to achieve better phrase groundability without sacrificing representation generality. |
Copied to clipboard
| Challenge: | Existing, naive defenses against adversarial attacks are lagging . a new paper aims to break these defenses with adaptive noise ensembling . |
| Approach: | They propose a randomized smoothing paradigm that can be used to break adversarial attacks . they use speech enhancement methods and a novel use for ASR output ensembling methods . |
| Outcome: | The proposed model is robust to all attacks that use inaudible noise and can only be broken with very high distortion. |
Copied to clipboard
| Challenge: | Current literature mostly considers textual content while assessing argument quality, and it is limited to datasets containing short text sequences (18-48 words). |
| Approach: | They propose a set of interpretable debate centric features that are inspired by theories of argument quality and propose MARQ model that summarizes the multimodal signals on long debate videos. |
| Outcome: | The proposed model outperforms baseline models with an error rate reduction of 22.7% on the argument quality prediction task and achieves 81.91% accuracy. |
Copied to clipboard
| Challenge: | Prior implementations of NMN use pre-defined and fixed textual inputs in their module instantiation. |
| Approach: | They propose to parameterize the module arguments to reduce the number of modules in NMN by up to 75% without any loss in performance. |
| Outcome: | The proposed model outperforms the state-of-the-art model on CLEVR-Ref+ dataset with +8.1% improvement in accuracy and +4.3% on full test set. |
Copied to clipboard
| Challenge: | Existing knowledge-based visual question answering systems rely on Concept-Net and Wikipedia to obtain external knowledge. |
| Approach: | They propose a visual retriever-reader pipeline that uses a natural language knowledge base and a Visual retriever to retrieve relevant knowledge. |
| Outcome: | The proposed method significantly improves the visual retriever-reader pipeline on the OK-VQA benchmark. |
Copied to clipboard
| Challenge: | Vision-and-Dialogue Navigation is one of the tasks that evaluate the agent’s ability to interact with humans for assistance and navigate based on natural language responses. |
| Approach: | They propose a vision-and-dialogue navigation task which evaluates the agent's ability to interact with humans and navigate based on natural language responses. |
| Outcome: | The proposed model performs well on the Navigation from Dialogue History task, but it is not evaluated by the primary metric Goal Progress. |
Copied to clipboard
| Challenge: | Existing methods for timeline summarization ignore the events’ intra-structures and inter-structure connections. |
| Approach: | They propose to represent news articles as an event-graph, thus compressing the whole graph to its salient sub-graph. |
| Outcome: | The proposed method significantly improves on the state-of-the-art on three real-world datasets, including two public benchmarks and a Timeline100 dataset. |
Copied to clipboard
| Challenge: | StreamHover is a framework for annotating and summarizing livestream transcripts . the problem is that there is n't enough annotated datasets to summarize livestreams based on the informal nature of spoken language . |
| Approach: | They propose a framework for annotating and summarizing livestream transcripts using a text preview. |
| Outcome: | The proposed model generalizes better and improves over strong baselines. |
Copied to clipboard
| Challenge: | Part of speech (POS) tagging models are underperforming on headlines due to differences in the register of English news headlines and long-form text. |
| Approach: | They propose to annotate news headlines with POS tags by projecting predicted tags from corresponding sentences in news bodies. |
| Outcome: | The proposed model reduces errors by 23% and 19% on a newly-annotated corpus of over 5,248 English news headlines from the Google sentence compression corpus. |
Copied to clipboard
| Challenge: | KnowledgeEditor can be used to edit factual knowledge stored in Language Models without the need for expensive retraining or fine-tuning. |
| Approach: | They propose a method which edits factual knowledge implicitly stored in Language Models and uses it to fix 'bugs' and 'obvious errors' they train a hyper-network with constrained optimization to modify a fact without affecting the rest of the knowledge; the hyper-netzwork is then used to predict the weight update at test time. |
| Outcome: | The proposed method can be used to edit factual knowledge without retraining or fine-tuning and can fix 'bugs' or unexpected predictions without the need for expensive re-training or meta-learning. |
Copied to clipboard
| Challenge: | Recent studies have suggested that sparse attention mechanisms can be made more interpretable by replacing the softmax activation with its sparser variants. |
| Approach: | They propose a method to replace softmax activation with a ReLU to achieve sparsity in attention by layer normalization with either a specialized initialization or an additional gating function. |
| Outcome: | The proposed model is easy to implement and more efficient than previously proposed sparse attention mechanisms. |
Copied to clipboard
| Challenge: | Existing knowledge bases are tedious and require a large amount of labor to build. |
| Approach: | They propose a method that allows for transfer of knowledge from one collection of facts to another without entity or relation matching. |
| Outcome: | The proposed method is the most impactful on small datasets, showing a 6% increase in rank and 65% decrease in rank over the previous best method. |
Copied to clipboard
| Challenge: | Sparse attention mechanisms are a deterministic alternative, but they lack a way to regularize rationale extraction. |
| Approach: | They propose a framework for deterministic extraction of structured explanations via constrained inference on a factor graph, forming a differentiable layer. |
| Outcome: | The proposed framework outperforms previous studies on performance and plausibility of extracted rationales. |
Copied to clipboard
| Challenge: | Knowledge distillation (KD) is a common knowledge transfer algorithm used for model compression across a variety of deep learning based natural language processing (NLP) solutions. |
| Approach: | They propose to use teacher training data for model compression . they investigate six tasks and find they can achieve between 75% and 92% of the teacher’s classification score while compressing the model 30 times. |
| Outcome: | The proposed solution achieves between 75% and 92% of the teacher’s classification score while compressing the model 30 times. |
Copied to clipboard
| Challenge: | Existing approaches to adversarial regularization treat adversarials and defending players equally, which is undesirable because only the defending player contributes to the generalization performance. |
| Approach: | They propose a method which formulates adversarial regularization as a Stackelberg game and induces a competition between a leader and a follower. |
| Outcome: | The proposed method outperforms existing adversarial regularization baselines on a set of machine translation and natural language understanding tasks. |
Copied to clipboard
| Challenge: | Recent work on opinion summarization produces general summaries based on reviews and popularity of opinions expressed in them. |
| Approach: | They propose an approach that generates customized opinion summaries based on aspect queries. |
| Outcome: | The proposed model outperforms the current state of the art and generates personalized summaries by controlling the number of aspects discussed in them. |
Copied to clipboard
| Challenge: | Existing evaluation metrics for summarization evaluation are limited and do not correlate well with human judgments. |
| Approach: | They propose to extend existing evaluation metrics to include question answering models to assess whether a summary contains all relevant information in its source document. |
| Outcome: | The proposed framework significantly improves the correlation with human judgments over four evaluation dimensions. |
Copied to clipboard
| Challenge: | Abstractive conversation summarization models heavily rely on human-annotated summaries. |
| Approach: | They propose a set of Conversational Data Augmentation methods for semi-supervised abstractive conversation summarization that use random swapping/deletion to perturb the discourse relations inside conversations and dialogue-acts-guided insertion to interrupt the development of conversations. |
| Outcome: | The proposed methods over several state-of-the-art datasets show that they are more efficient than previous methods. |
Copied to clipboard
| Challenge: | Automated summarization metrics are reliable but often poorly correlated with human judgment. |
| Approach: | They propose a semi-automatic to automatic summary evaluation metrics, following the Pyramid human evaluation method. |
| Outcome: | The proposed metrics are semi-automatic to automatic summary evaluation metrics, following the Pyramid human evaluation method. |
Copied to clipboard
| Challenge: | Existing methods for generating abstractive summarization are inconsistent and rely on heuristically created data for error handling. |
| Approach: | They propose a contrastive learning formulation that leverages both positive and negative summaries to train summarization systems that are better at distinguishing between them. |
| Outcome: | The proposed learning framework produces more factual summaries than strong comparisons with post error correction, entailment-based reranking, and unlikelihood training. |
Copied to clipboard
| Challenge: | Multilingual unsupervised machine translation is a computationally expensive and hard to tune approach . auxiliary parallel data is used to train translation systems from monolingual data . |
| Approach: | They propose to use auxiliary parallel language pairs to train unsupervised machine translations . they propose to add auxiliary languages to pre-trained mBART-50 models with denoising adapters . |
| Outcome: | The proposed approach is on-par with back-translation and allows adding unseen languages incrementally. |
Copied to clipboard
| Challenge: | Existing methods for incorporating pre-trained models into NMT systems are non-trivial and lack a comparison of the impact that other pre-trainers may have on translation performance. |
| Approach: | They propose to use the input of a bilingual pre-trained language model as the input for NMT encoders and a stochastic layer selection approach to ensure sufficient utilization of contextualized embeddings. |
| Outcome: | The proposed bilingual pre-trained language model outperforms all other pre-train models on the IWSLT’14 dataset and the proposed dual-directional translation model. |
Copied to clipboard
| Challenge: | A standard approach for exerting control in MT is to prepend the input with a special tag to signal the desired output attribute. |
| Approach: | They propose a vector-valued approach which allows for fine-grained control over multiple attributes simultaneously via a weighted linear combination of the corresponding vectors. |
| Outcome: | The proposed approach achieves better control over a wider range of tasks than tagging and even fine-tuning a model trained without annotations. |
Copied to clipboard
| Challenge: | Existing approaches use a fixed number of source words to translate or learn dynamic policies for the number of sources by reinforcement learning. |
| Approach: | They propose a generative framework that uses a latent variable to model read or translate actions at every time step and integrates out to consider all possible translation policies. |
| Outcome: | The proposed framework achieves the best BLEU scores on benchmark datasets. |
Copied to clipboard
| Challenge: | Existing siMT systems are trained and evaluated on offline translations . however, evaluation gap remains notable, calling for constructing large-scale interpretation corpora . |
| Approach: | They propose a translation-to-interpretation transfer method which converts offline translations into interpretation-style data. |
| Outcome: | The proposed interpretation test set shows that SiMT models improve on translation vs interpretation data. |
Copied to clipboard
| Challenge: | Recent pre-trained language models have achieved remarkable zero-shot performance . we propose a self-learning framework that utilizes unlabeled data of target languages . |
| Approach: | They propose a self-learning framework that utilizes unlabeled data of target languages to select silver labels for cross-lingual transfer tasks. |
| Outcome: | The proposed framework outperforms baseline models on two cross-lingual tasks by 10 F1 on average and 2.5 accuracy on natural language inference (NLI). |
Copied to clipboard
| Challenge: | a novel scheme to perform word-level quality estimation is proposed for word-based quality estimation . authors propose a two-stage transfer learning procedure on augmented and human data . a Levenshtein Transformer can learn to post-edit without explicit supervision. |
| Approach: | They propose a novel scheme to use a Levenshtein Transformer to perform word-level quality estimation. |
| Outcome: | The proposed method performs better under data-constrained and unconstrained conditions than existing methods. |
Copied to clipboard
| Challenge: | Extensive experiments on iSQuAD suggest that graph representations can result in significant performance improvements for RL agents. |
| Approach: | They propose to use graph representations to build and update graphs during information gathering and neural models to encode graph representation in RL agents. |
| Outcome: | Extensive experiments on iSQuAD show that graph representations can improve performance for RL agents. |
Copied to clipboard
| Challenge: | Automatic Speech Recognition systems perform poorly on atypical speech and heavily accented speech. |
| Approach: | They add a residual adapter to the encoder layer to improve model adaptation . they show that the residual adapters update only a tiny fraction of the model parameters . |
| Outcome: | The proposed model fine-tuning improves performance on atypical and accented speech . the system can update only a tiny fraction of the model parameters . |
Copied to clipboard
| Challenge: | Visual News Captioner is an entity-aware model for news image captioning . Unlike standard image captions, news images depict situations where people, locations, and events are of paramount importance. |
| Approach: | They propose a visual news captioner model that integrates visual and textual features to generate captions with richer information such as events and entities. |
| Outcome: | The proposed model can generate captions with richer information such as events and entities. |
Copied to clipboard
| Challenge: | Existing work on text-to-image synthesis does not explore the use of linguistic structure of the input text. |
| Approach: | They propose to use constituency parse trees to encode structured input and a Transformer-based recurrent architecture to augment commonsense information to generate visual story. |
| Outcome: | The proposed model improves visual quality and consistency of images from a target domain without fine-tuning. |
Copied to clipboard
| Challenge: | Recent work adopts a "pre-training + fine-tuning" approach for zero-shot transfer to end tasks without fine- tuning. |
| Approach: | They propose a contrastive approach to pre-train a transformer model for zero-shot video and text understanding without using any labels on downstream tasks. |
| Outcome: | The proposed model outperforms supervised approaches on downstream tasks and outperformed previous approaches. |
Copied to clipboard
| Challenge: | a threat scenario where an image is used out of context to support a narrative is proposed. |
| Approach: | They propose a dataset where both image and text are unmanipulated but mismatched . they benchmark several state-of-the-art multimodal models on their dataset . |
| Outcome: | The proposed dataset shows that machine-driven image repurposing is now a realistic threat . it provides samples that represent challenging instances of mismatch between text and image . |
Copied to clipboard
| Challenge: | Comparative Preference Classification (CPC) is a natural language processing task that predicts whether a preference comparison exists between two entities in a given sentence . |
| Approach: | They propose a sentiment analyzer that learns sentiments to individual entities via domain adaptive knowledge transfer. |
| Outcome: | Experiments on the CompSent-19 dataset present a significant improvement on the F1 scores over the best existing CPC approaches. |
Copied to clipboard
| Challenge: | a new neural architecture can be used to classify stances on social media without relying on linguistic features. |
| Approach: | They propose a neural architecture where the input also includes automatically generated negated perspectives over a given claim. |
| Outcome: | The proposed model improves on the original input and removes doubtful predictions over the retained information. |
Copied to clipboard
| Challenge: | Crypto markets are forums where goods and services are exchanged between parties who use encryption to conceal their identities. |
| Approach: | They propose a stylometry-based multitask learning approach for natural language and model interactions using graph embeddings. |
| Outcome: | The proposed approach outperforms existing methods in four darknet forums with a lift of up to 2.5X on the mean retrieval rank and 2X on recall@10. |
Copied to clipboard
| Challenge: | Existing studies on dyadic human-human interactions focus on conversations without specific business objectives. |
| Approach: | They propose a method to detect emotions in a live chat customer service . they propose 'ProtoSeq' for conversational emotion classification using different languages . |
| Outcome: | The proposed method is competitive even when applied to other ones. |
Copied to clipboard
| Challenge: | Existing studies have focused on continual learning of aspect sentiment classification (ASC) tasks in domain incremental learning (DIL) |
| Approach: | They propose a continual learning method that learns a sequence of tasks incrementally . they propose CLASSIC, which uses a domain incremental learning setting . |
| Outcome: | The proposed model is highly effective in a domain incremental learning setting. |
Copied to clipboard
| Challenge: | Existing methods for implicit sentiment analysis simply view noun phrases or entities in text as events or indirectly model events with sophisticated models. |
| Approach: | They propose an event-centric implicit sentiment analysis that utilizes the sentiment-aware event contained in a sentence to infer sentiment polarity. |
| Outcome: | The proposed model can detect sentiment in sentences without sentiment words and is compared to existing models on a benchmark dataset. |
Copied to clipboard
| Challenge: | Existing methods for learning universal sentence embeddings are based on unsupervised approaches with only dropout as noise. |
| Approach: | They propose an unsupervised approach that takes an input sentence and predicts itself in a contrastive objective with only standard dropout used as noise. |
| Outcome: | The proposed framework performs on par with previous supervised approaches and can produce superior sentence embeddings from unlabeled or labeled data. |
Copied to clipboard
| Challenge: | Using manual content to learn languages is expensive and time consuming. |
| Approach: | They propose a method for automatically identifying fine-grained lexical distinctions and extracting rules explaining them in a human- and machine-readable format. |
| Outcome: | The proposed method is able to identify fine-grained distinctions and explain them in a human- and machine-readable format. |
Copied to clipboard
| Challenge: | a recipe explains step by step how to cook a dish, but recipes differ in which cooking actions they describe explicitly, how they describe them, and in which order. |
| Approach: | They propose a recipe corpus which annotates cooking steps in recipes at sentence level . they train a neural model to predict recipes on ARA and model it for automatic understanding . |
| Outcome: | The proposed model can predict recipes with fine-grained structural information . it shows that recipes can be explained in different ways, or not at all . |
Copied to clipboard
| Challenge: | Recent approaches to obtain high-quality sentence embeddings from pretrained language models require labeled data or finetuned on large set of labeles. |
| Approach: | They propose to use generative abilities of large and high-performing PLMs to generate entire datasets of labeled text pairs from scratch and fine tune much smaller and more efficient models. |
| Outcome: | The proposed approach outperforms baselines on several semantic textual similarity datasets. |
Copied to clipboard
| Challenge: | Pretrained language models can be used to perform lexical inference in context tasks with relatively small training data. |
| Approach: | They propose to combine a pretrained language model with textual patterns to improve performance in both zero-shot and few-shot settings. |
| Outcome: | The proposed method compares pre-trained models with textual patterns on two established benchmarks for lexical inference in context (LIiC) the results show that the proposed patterns improve performance on LIiC, setting a new state of the art. |
Copied to clipboard
| Challenge: | Specialized number representations have shown improvements on numerical reasoning tasks like arithmetic word problems and masked number prediction. |
| Approach: | They propose to use six different number encoders to improve masked word prediction by avoiding conflating nominal and ordinal number occurrences. |
| Outcome: | The proposed representations improve masked word prediction accuracy and generalize to contexts without annotated numbers. |
Copied to clipboard
| Challenge: | Neural networks depend heavily on lexicalized information, which can be overfitted . this can be a problem in fact verification, which has important societal implications. |
| Approach: | They propose a knowledge distillation approach for fact verification using student models. |
| Outcome: | The proposed approach outperforms state-of-the-art classifiers on a training dataset and in supervised settings. |
Copied to clipboard
| Challenge: | MULTI-EURLEX is a dataset for topic classification of EU legal documents . fine-tuning a multilingually pretrained model in a single source language leads to catastrophic forgetting of multilingual knowledge and poor zero-shot transfer to other languages. |
| Approach: | They propose to use the dataset as a testbed for zero-shot cross-lingual transfer to exploit annotated training documents in one language to classify documents in another language. |
| Outcome: | The proposed model can be used to classify EU legal documents in other languages without a single source language and retain multilingual knowledge. |
Copied to clipboard
| Challenge: | Existing approaches to multi-answer retrieval cannot reason about the set of passages jointly. |
| Approach: | They propose a joint passage retrieval model focusing on reranking to solve multi-answer retrieval problem. |
| Outcome: | The proposed model outperforms baseline models on three multi-answer datasets. |
Copied to clipboard
| Challenge: | Recent studies have shown that discriminative training results in models that exploit these underlying biases to achieve a better held-out performance, without learning the right way to reason. |
| Approach: | They propose a generative context selection model for multi-hop QA that reasons about how the given question could have been generated given a context pair and not just independent contexts. |
| Outcome: | The proposed model outperforms the state-of-the-art model on hotpotQA while being comparable to the state of the art answering performance on adversarial held-out set. |
Copied to clipboard
| Challenge: | Existing methods to improve Question Answering performance on non-English data are expensive and limited to evaluation set. |
| Approach: | They propose a method to improve Question Answering performance without additional annotations by leveraging Question Generation models to produce synthetic samples in a cross-lingual fashion. |
| Outcome: | The proposed method outperforms baselines on four datasets in English significantly . the proposed model outperformed baselines in english and is comparable to the validation set of the original SQuAD. |
Copied to clipboard
| Challenge: | Numerical reasoning in machine reading comprehension (MRC) has shown drastic improvements over the past few years. |
| Approach: | They propose an E-digit number form that alleviates the lack of extrapolation in numerical MRC models. |
| Outcome: | The proposed model can't extrapolate to unseen numbers, the authors say . they also show that the model needs to treat numbers differently from regular words . |
Copied to clipboard
| Challenge: | Large language models have shown promising results in zero-shot settings due to surface form competition . since probability mass is finite, this lowers the probability of the correct answer . |
| Approach: | They propose a scoring function that compensates for surface form competition by reweighing each option according to its a priori likelihood. |
| Outcome: | The proposed scoring function achieves consistent gains in zero-shot over calibrated and uncalibrated scoring functions on all GPT-2 and GPT-3 models on a variety of multiple choice datasets. |
Copied to clipboard
| Challenge: | Knowledge-dependent tasks typically use two sources of knowledge: parametric, learned at training time, and contextual, given as a passage at inference time. |
| Approach: | They propose a method to mitigate over-reliance on parametric knowledge, which minimizes hallucination, and improves out-of-distribution generalization by 4% - 7%. |
| Outcome: | The proposed method minimizes hallucination and improves generalization to evolving information by 4% - 7%. |
Copied to clipboard
| Challenge: | Using self-training to train unsupervised domains can be expensive, resulting in poor generalization due to distributional shift. |
| Approach: | They propose to use unaligned data to train unsupervised domain adaptation models using cheap synthetically generated labeled data. |
| Outcome: | The proposed method significantly outperforms self-training on question generation and passage retrieval domains and on MLQuestions and PubMedQA. |
Copied to clipboard
| Challenge: | Existing methods for graded contextual word meaning annotation have not been implemented yet. |
| Approach: | They propose a multi-round incremental annotation process and a clustering algorithm to group usages into senses to create a large-scale dataset. |
| Outcome: | The proposed method is the largest resource of graded contextualized, diachronic word meaning annotation in four different languages, based on 100,000 human semantic proximity judgments. |
Copied to clipboard
| Challenge: | Using machine translation, counterfactual statements are often found in natural languages. |
| Approach: | They annotate a multilingual CFD dataset from Amazon product reviews covering counterfactuals written in English, German, and Japanese languages. |
| Outcome: | The proposed dataset is robust against selection biases due to cue phrase-based sentence selection. |
Copied to clipboard
| Challenge: | linguistic style is an integral part of natural language, but evaluation methods for style measures are rare, often task-specific and usually do not control for content. |
| Approach: | They propose a modular, fine-grained and content-controlled similarity-based STyle EvaLuation framework to test the performance of any model that can compare two sentences on style. |
| Outcome: | The proposed model outperforms simple versions of commonly used style measures like 3-grams, punctuation frequency and LIWC-based approaches. |
Copied to clipboard
| Challenge: | Text generation systems are ubiquitous in natural language processing applications, but evaluation of these systems remains a challenge, especially in multilingual settings. |
| Approach: | They propose a metric to evaluate the morphosyntactic well-formedness of text using its dependency parse and morphologically-rich rules of the language. |
| Outcome: | The proposed metric can evaluate the morphosyntactic well-formedness of text using its dependency parse and morphologically-rich rules of the language. |
Copied to clipboard
| Challenge: | Existing multilingual evaluation datasets that evaluate lexical semantics "in-context" have various limitations, including limited coverage of high-resource languages and superficial cues. |
| Approach: | They propose to use a set of pretrained language models to evaluate lexical semantics in context. |
| Outcome: | The proposed set shows that current models lag behind human performance in interpreting word meaning in cross-lingual contexts. |
Copied to clipboard
| Challenge: | We study whether and how cross-task generalization ability can be acquired . we use CrossFit to standardize seen/unseen task partitions and evaluation protocols . |
| Approach: | They propose a problem setup for studying cross-task generalization ability which standardizes seen/unseen task partitions and data access during different learning stages. |
| Outcome: | The proposed model can be used to build few-shot learners across diverse tasks. |
Copied to clipboard
| Challenge: | Existing studies show that inserting an intermediate pre-training stage improves performance of masked language models. |
| Approach: | They propose methods to automate the discovery of optimal masking policies via direct supervision or meta-learning. |
| Outcome: | The proposed method outperforms the heuristic of masking named entities on TriviaQA and can be generalizable beyond that task. |
Copied to clipboard
| Challenge: | Word embeddings learn implicit biases from word co-occurrence statistics . valNorm is a new intrinsic evaluation task and method to quantify affect in word embedded word sets . |
| Approach: | They propose a method to quantify valence dimension of affect in human-rated word sets . they apply ValNorm to embeddings from seven languages and 200 years of text . |
| Outcome: | The proposed method achieves a high accuracy in quantifying the valence of non-discriminatory, non-social group word sets. |
Copied to clipboard
| Challenge: | Existing evaluation metrics for natural language generation are inadequate . existing metrics are not robust against simple perturbations and disagree with scores assigned by humans to perturbed output. |
| Approach: | They propose to propose checks which perturb the output and target a specific criteria and then use them to refine their evaluation. |
| Outcome: | The proposed templates show that existing evaluation metrics are not robust against simple perturbations and disagree with human scores on the perturbed output. |
Copied to clipboard
| Challenge: | MT models have discrete vocabularies and often use subword segmentation to achieve an ‘open vocabulary’. |
| Approach: | They propose to use visual text representations to create continuous vocabularies by processing visually rendered text with sliding windows. |
| Outcome: | The proposed models achieve 25.9 BLEU on character permuted German–English task, compared with traditional models on smaller and larger datasets. |
Copied to clipboard
| Challenge: | despite improvements in machine translation quality, automatic poetry translation remains a challenging problem . et al., a study of automatic poetry translators shows that multilingual fine-tuning on poetic data outperforms bilingual fine-timing on non-poetic text . |
| Approach: | They propose to use poetic parallel corpora for 6 languages to study poetry translation . they find that multilingual fine-tuning on poetic data outperforms bilingual fine-uning . |
| Outcome: | The proposed model outperforms bilingual and multilingual models on poetic data . the proposed model is based on a parallel dataset of poetry translations for several languages . |
Copied to clipboard
| Challenge: | Multilingual Neural Machine Translation models often produce low quality translations, often failing to produce outputs in the right target language. |
| Approach: | They propose a joint approach to regularize NMT models at both representation-level and gradient-level to reduce off-target translation occurrences and improve zero-shot translation performance. |
| Outcome: | The proposed approach reduces off-target translation occurrences and improves zero-shot translation performance by +5.59 and +10.38 BLEU on WMT and OPUS datasets. |
Copied to clipboard
| Challenge: | Existing methods to update deployed models are prone to overfit . however, non-parametric methods are liable to over-fit the retrieved examples . |
| Approach: | They propose to learn Kernel-Smoothed Translation with Example Retrieval (KSTER) this approach allows users to adapt models to emerging cases without retraining . |
| Outcome: | The proposed approach achieves 1.1 to 1.5 BLEU scores over existing methods without retraining . the proposed model is released on https://github.com/jiangqn/KSTER. |
Copied to clipboard
| Challenge: | MultiUAT dynamically adjusts training data usage based on model’s uncertainty on a small set of trusted clean data for multi-corpus machine translation. |
| Approach: | They propose an approach that dynamically adjusts the training data usage based on the model’s uncertainty on a small set of trusted clean data for multi-corpus machine translation. |
| Outcome: | The proposed approach outperforms baselines on 16 languages and 2 domains on English-German translation. |
Copied to clipboard
| Challenge: | Existing methods for simultaneous machine translation require multiple models for different latency levels, resulting in large computational costs. |
| Approach: | They propose a universal SiMT model with Mixture-of-Experts Wait-k Policy to achieve the best translation quality under arbitrary latency with only one model. |
| Outcome: | The proposed model outperforms all the strong baselines under different latency levels including the state-of-the-art adaptive policy. |
Copied to clipboard
| Challenge: | a new reasoning challenge is proposed to help AI systems to solve real-world problems . Fermi Problems are questions whose answers can only be approximated because their computation is either impossible or impossible. |
| Approach: | They propose a new reasoning challenge, Fermi Problems, which asks questions whose answers can only be approximated because their computation is either impractical or impossible. |
| Outcome: | The proposed datasets show that even fine-tuned large-scale language models perform poorly on these datasets. |
Copied to clipboard
| Challenge: | Existing methods to improve QA efficiency do not take specific answers into account. |
| Approach: | They propose a transformer-based approach to improve QA efficiency by filtering out questions that will not be answered by the system. |
| Outcome: | The proposed model can approximate the Precision/Recall curves of the target QA system. |
Copied to clipboard
| Challenge: | a study shows that training reading comprehension models assumes that the training instances are independent and identically distributed . however, this assumption can cause the learner to ignore distinguishing cues between related or minimally different questions . |
| Approach: | They propose to normalize question-answer scores across neighborhoods of closely contrasting questions and/or answers by adding a cross entropy loss term to the supervision signal. |
| Outcome: | The proposed methods show up to 9% absolute gains in accuracy on two datasets. |
Copied to clipboard
| Challenge: | ENTAILMENTBANK is the first dataset to contain multistep entailment trees. |
| Approach: | They propose to generate explanations in the form of entailment trees, a tree of multipremise entanglements steps from facts that are known to the hypothesis of interest. |
| Outcome: | The proposed model can generate explanations in the form of entailment trees . this is a tree of multipremise enttailment steps from facts known to the hypothesis of interest. |
Copied to clipboard
| Challenge: | Existing models struggle with producing answers that are frequently updated or from uncommon locations. |
| Approach: | They propose an open-retrieval QA dataset where systems must produce the correct answer given the context. |
| Outcome: | The proposed dataset shows that existing models struggle with producing answers that are frequently updated or from uncommon locations. |
Copied to clipboard
| Challenge: | Existing studies on abusive language towards conversational AI systems are not conclusive as they are not performed with live systems nor with real users due to the lack of reliable abuse detection tools. |
| Approach: | They propose to use a convAI dataset to account for the complexity of the task and to bench-mark existing models against this data. |
| Outcome: | The proposed model shows that abuse distribution is different compared to other datasets, with sexual tinted aggression towards the virtual persona of the systems. |
Copied to clipboard
| Challenge: | Currently, conversational agents lack commonsense reasoning, preventing them from engaging in rich conversations with humans. |
| Approach: | They propose a commonsense reasoning system that uncovers unstated presumptions from user commands satisfying a general template of if-(state), then-(action), because-(goal) They propose to use a transformer-based generative commons sense knowledge base as its source of background knowledge to extract multi-hop reasoning chains from the neural KB. |
| Outcome: | The proposed model achieves a 35% higher success rate than existing methods with human users. |
Copied to clipboard
| Challenge: | Existing methods for evaluation of dialog systems are expensive and not scalable . a framework for estimating human evaluation scores is proposed to bridge this gap . |
| Approach: | They propose a framework for estimating human evaluation scores based on off-policy evaluation . they use language quality metrics for single-turn response generation given a fixed context . |
| Outcome: | The proposed framework outperforms existing methods in terms of correlation with human evaluation scores. |
Copied to clipboard
| Challenge: | Existing continuous learning systems are not designed to add new domains and functionalities through time without incurring the high cost of retraining the whole system. |
| Approach: | They propose a first-ever continual learning benchmark for task-oriented dialogue systems . they propose 'architecture' method based on residual adapters to implement continual training . |
| Outcome: | The proposed architectural method performs better than multitask learning while being 20X faster in learning new domains. |
Copied to clipboard
| Challenge: | a systematic study on multilingual and cross-lingual intent detection from spoken data is presented . current work on intent detection is limited to English, and standard benchmarks exist only in English. |
| Approach: | They present a systematic study on multilingual and cross-lingual intent detection from spoken data. |
| Outcome: | The proposed resource is called MInDS-14, and it provides strong intent detection in most target languages. |
Copied to clipboard
| Challenge: | Existing dialog models are unable to handle popular figurative language constructs like metaphor and simile when faced with dialog contexts containing figurativ language. |
| Approach: | They propose lightweight solutions to help existing dialog models become more robust to figurative language by simply using an external resource to translate figurativ language to literal (non-figurative) forms while preserving the meaning to the best extent possible. |
| Outcome: | The proposed models show large drops in performance when faced with dialog contexts consisting of figurative language compared to contexts without figurativ language . |
Copied to clipboard
| Challenge: | Using Sequence-to-Sequence models for dialogue state tracking remains an understudied topic. |
| Approach: | They propose to use a pre-training objective and a dialogue context representation to investigate this problem. |
| Outcome: | The proposed model is more effective than auto-regressive language modeling, the authors show . the proposed model may have a hard time recovering from earlier mistakes, they say . |
Copied to clipboard
| Challenge: | Existing datasets for multi-document summarization (MDS) are either in the general domain, such as WikiSum, or very small such as DUC 1 or TAC 2011 . Existing systems for summarizing biomedical literature take 1-2 years to complete . |
| Approach: | They propose to use a multi-document summarization system based on BART to assess the quality of the summarized biomedical literature. |
| Outcome: | The proposed system has high summarization quality, but significant work remains to achieve it. |
Copied to clipboard
| Challenge: | Image captioning relies on reference-based automatic evaluations, but references are expensive to collect and comparing against multiple human-authored captions is insufficient. |
| Approach: | They propose a reference-free metric that can be used for automatic caption evaluation without references. |
| Outcome: | The proposed model outperforms existing metrics on image-text compatibility and a reference-augmented version achieves even higher correlation with human judgements. |
Copied to clipboard
| Challenge: | a large corpus of domain-expert relevance ratings augments a corpus for compositional explanations . a writer's study shows that the evaluations of compositional inference models underestimate performance . |
| Approach: | They construct a corpus of 126k domain-expert relevance ratings that augment explanations to standardized science exam questions. |
| Outcome: | The results show that evaluations underestimate performance of compositional explanations . they show that models regularly discover and produce valid explanations that are different than gold explanations. |
Copied to clipboard
| Challenge: | Recent event-centric reading comprehension datasets focus mostly on event arguments or temporal relations. |
| Approach: | They propose a machine reading comprehension dataset that leverages natural language queries to reason about the five most common event semantic relations. |
| Outcome: | The proposed dataset shows that current SOTA systems achieve 22.1%, 63.3% and 83.5% for token-based exact-match, **F1** and event-based **HIT@1** scores. |
Copied to clipboard
| Challenge: | Pre-trained language models have impressive performance on commonsense inference benchmarks, but their ability to make robust inferences is debated. |
| Approach: | They propose a challenge that evaluates robust commonsense inference despite textual perturbations using commonsensical knowledge bases and probe PTLMs across two different evaluation settings. |
| Outcome: | The proposed procedure evaluates robust commonsense inference despite textual perturbations using commonsensense knowledge bases and probe PTLMs across two evaluation settings. |
Copied to clipboard
| Challenge: | Natural language generation (NLG) tasks have complex nature and require manual evaluation. |
| Approach: | They propose a unifying perspective based on the nature of information change in NLG tasks . they propose 'information alignment' metrics that can be used to evaluate different aspects of NLG . |
| Outcome: | The proposed metrics achieve stronger or comparable correlations with human judgement compared to state-of-the-art metrics in diverse tasks. |
Copied to clipboard
| Challenge: | Tables are ubiquitous on the web, and are rich in information. |
| Approach: | They propose a sparse-attention Transformer architecture for modeling documents that contain large tables. |
| Outcome: | The proposed architecture scales linearly with respect to speed and memory, and can handle documents containing more than 8000 tokens with current accelerators. |
Copied to clipboard
| Challenge: | a lack of annotator agreement can hinder training of NLP systems . we propose a learning algorithm that can learn from training examples with zero, one, or multiple labels. |
| Approach: | They propose an annotation distribution scheme that assigns multiple labels to training examples . they propose a learning algorithm that can learn from training examples with different amount of annotation . |
| Outcome: | The proposed method achieves consistent gains in two tasks, suggesting distributing labels unevenly among training examples can be beneficial for many NLP tasks. |
Copied to clipboard
| Challenge: | Large language models are difficult to train because of the growing computation time and cost. |
| Approach: | They propose a highly-efficient architecture that combines fast recurrence and attention for sequence modeling. |
| Outcome: | The proposed model achieves state-of-the-art on a Wiki-103 and Billion Word datasets using 1.6 days of training on an 8-GPU machine. |
Copied to clipboard
| Challenge: | Existing methods for intermediate layer matching are limited due to huge over-parameterization . |
| Approach: | They propose to match intermediate layers of teacher and student in output space via attention-based layer projection. |
| Outcome: | The proposed method outperforms existing methods on GLUE tasks. |
Copied to clipboard
| Challenge: | Existing approaches to EL have been shown to be effective for both Entity Disambiguation and Entity Linking, but they suffer from high computational cost due to a complex (deep) decoder and the need for training on a large amount of data. |
| Approach: | They propose a method that parallelizes autoregressive linking across all potential mentions and relies on a shallow and efficient decoder. |
| Outcome: | The proposed method outperforms state-of-the-art approaches on the English dataset AIDA-CoNLL and is >70 times faster and more accurate than the previous generative method. |
Copied to clipboard
| Challenge: | Recent coreference resolution models rely heavily on span representations to find coreference links between word spans. |
| Approach: | They propose to consider coreference links between individual words rather than word spans and reconstruct the word span. |
| Outcome: | The proposed model outperforms existing models on the OntoNotes benchmark while being highly efficient. |
Copied to clipboard
| Challenge: | Existing FL frameworks require a trusted aggregator or require heavy-weight cryptographic primitives, which makes the performance significantly degraded. |
| Approach: | They propose a framework that is federated and efficient for NLP . they propose to eliminate the need for trusted entities and achieve better model accuracy . |
| Outcome: | The proposed framework achieves better model accuracy and model accuracy than existing FL frameworks. |
Copied to clipboard
| Challenge: | a mechanism for enacting behavior changes without expensive model re-training would be preferable. |
| Approach: | They propose a controllable semantic parser that retrieves related exemplars from a retrieval index and augments them to the query. |
| Outcome: | The proposed model can parse queries in a new domain, adapt predictions toward specified patterns, or adapt to new semantic schemas without re-training the model. |
Copied to clipboard
| Challenge: | Large pretrained language models excel at generating natural language, but they are not efficient for task specific semantic parsing. |
| Approach: | They propose to use large pretrained language models as few-shot semantic parsers . they paraphrase inputs into a controlled sublanguage resembling English . |
| Outcome: | The proposed model can generate surprisingly accurate models on multiple tasks with minimal code and data. |
Copied to clipboard
| Challenge: | Current commonsense-reasoning tasks are discriminative in nature, where a model answers a multiple-choice question for a certain context. |
| Approach: | They propose a generative task that generates a commonsense-augmented graph for stance prediction by using a create-verify-and-refine graph collection framework. |
| Outcome: | The proposed model is able to generate a graph that serves as non-trivial, complete, and unambiguous explanation for the predicted stance. |
Copied to clipboard
| Challenge: | Existing supervised models struggle to make correct predictions on rare word senses due to limited training data. |
| Approach: | They propose a gloss alignment algorithm that can align definition sentences with the same meaning from different sense inventories to collect rich lexical knowledge. |
| Outcome: | The proposed method outperforms previous methods on both frequent and rare word senses. |
Copied to clipboard
| Challenge: | Recent work casts GEC as a translation problem using encoder-decoder models to map bad (ungrammatical) sentences into good (grammatically) sentences. |
| Approach: | They propose to use a pretrained language model to define an LM-Critic that judges a sentence to be grammatical if the LM assigns it a higher probability than its local perturbations. |
| Outcome: | The proposed method outperforms existing methods in both the unsupervised and supervised setting. |
Copied to clipboard
| Challenge: | Existing methods to extract language-specific information from multilingual sentence embeddings are remarkably successful in cross-lingual and multilingual NLU tasks. |
| Approach: | They propose to extract language-specific information from the original embedding and use it to retrieve an embeddable that fully represents the sentence’s meaning. |
| Outcome: | The proposed method outperforms baselines on cross-lingual sentences even in low-resource language pairs where only tens of thousands of parallel sentence pairs are available. |
Copied to clipboard
| Challenge: | Existing research examines the origins of militarized conflict by examining bi-lateral relationships between entity pairs and multi-lateral relations among multiple entities. |
| Approach: | They propose to use Wikipedia to model dyadic and systemic causes to compare their correlations with conflict between two entities. |
| Outcome: | The proposed graphs show that Wikipedia articles of allies are semantically more similar than enemies. |
Copied to clipboard
| Challenge: | Prior efforts in POI type prediction focus on text without taking visual information into account. |
| Approach: | They propose to use multimodal information from text and images to infer the type of a place from where a social media post was shared. |
| Outcome: | The proposed method outperforms the state-of-the-art method for POI type prediction based on text-only methods and sheds light on cross-modal interactions and limitations. |
Copied to clipboard
| Challenge: | In this paper, we decompose the task of recognizing from the news coverage leading up to an election the (un)willingness of political parties to form a coalition into two related, but distinct tasks. |
| Approach: | They propose a task of recognizing from news coverage the (un)willingness of political parties to form a coalition from text and a sub-task of predicting the polarity of the signal. |
| Outcome: | The proposed approach improves over a strong monolingual transfer learning baseline. |
Copied to clipboard
| Challenge: | Existing methods based on latent topics cannot capture user interests and thus can't be used to predict how likely a user will post with a hashtag. |
| Approach: | They propose a personalized topic attention model that captures salient contents to personalize hashtag contexts by predicting how likely a user will post with a hashtag. |
| Outcome: | The proposed model significantly outperforms the state-of-the-art recommendation approach without exploiting latent topics. |
Copied to clipboard
| Challenge: | Recent advances in neural models have shown promising progress on this task, but key challenges remain . |
| Approach: | They propose a framework that can decouple dialogue generation from item recommendation . they use a response template generator and item selector to generate a responses template . |
| Outcome: | The proposed framework outperforms the state-of-the-art methods on the benchmark ReDial. |
Copied to clipboard
| Challenge: | Existing methods for evaluation of open-domain dialogues are expensive and require human annotators to evaluate their quality. |
| Approach: | They propose to use a deep-learning model trained on the general language understanding evaluation benchmark to serve as a quality indication of open-domain dialogues. |
| Outcome: | The proposed model can infer various quality metrics and derive a component-based overall score. |
Copied to clipboard
| Challenge: | Existing evaluation methods for factual consistency in knowledge-grounded dialogues are unreliable and limit their applicability. |
| Approach: | They propose an automatic evaluation metric for factual consistency in knowledge-grounded dialogue using automatic question generation and question answering. |
| Outcome: | The proposed evaluation metric consistently shows higher correlation with human judgements. |
Copied to clipboard
| Challenge: | Existing models for dialogue state tracking are based on Graph Attention Networks . if the relationship between slots and values is modelled explicitly, this can be improved . |
| Approach: | They propose a model architecture that augments GPT-2 with Graph Attention Networks to allow sequential prediction of slot values. |
| Outcome: | The proposed architecture improves performance against a strong GPT-2 baseline and with sparsely supervised training. |
Copied to clipboard
| Challenge: | Currently, most reinforcement learning methods for dialog policy learning train a centralized agent that selects a predefined joint action concatenating domain name, intent type, and slot name. |
| Approach: | They propose a hierarchical multi-agent framework in which each part of the action is led by a different agent and a joint optimization process that makes agents can exchange their policy information. |
| Outcome: | The proposed framework reduces labor costs for action templates and decreases the size of the action space for each agent. |
Copied to clipboard
| Challenge: | Existing approaches to training a dialogue state tracking model require extensive annotated dialogue data. |
| Approach: | They propose to transfer cross-task knowledge from general question answering corpora to QA model that can handle zero-shot DST. |
| Outcome: | The proposed model improves existing zero-shot and few-shot results on MultiWoz and shows better generalization ability in unseen domains. |
Copied to clipboard
| Challenge: | Neural dialogue belief trackers that take uncertainty into account are often overconfident in their decisions and therefore less robust. |
| Approach: | They propose to use different uncertainty measures in neural belief tracking to integrate uncertainty into the feature space of the policy and train policies through interaction with a user simulator. |
| Outcome: | The proposed approach improves both performance and robustness of the downstream dialogue policy. |
Copied to clipboard
| Challenge: | a pretrained language encoder can predict derailment in online conversations . this is a useful task for detecting and preventing abusive language . |
| Approach: | They extend a task to predict derailment in online conversations by using a pretrained language encoder. |
| Outcome: | The proposed task outperforms previous approaches in terms of performance and quality. |
Copied to clipboard
| Challenge: | Knowledge graph embedding is a new form of knowledge graphing that allows for better link prediction. |
| Approach: | They propose to use relational embedding to fit symmetry/antisymmetry and combination relationships. |
| Outcome: | The proposed model can fit symmetry/antisymmetry and combination relationships. |
Copied to clipboard
| Challenge: | Recent approaches to transformer models are expensive to fine-tune, slow for inference, and have large storage requirements. |
| Approach: | They propose a method to remove adapters from transformer layers during training and inference . they show that AdapterDrop can dynamically reduce computational overhead . |
| Outcome: | The proposed approach reduces computational overhead while maintaining performance over multiple tasks with minimal loss of performance. |
Copied to clipboard
| Challenge: | Recent advances in transformer quantization have shown remarkable improvement in many Natural Language Processing tasks and beyond. |
| Approach: | They propose a novel quantization scheme for transformers that can be quantized to ultra-low bit-widths, leading to significant memory savings with a minimum accuracy loss. |
| Outcome: | The proposed methods achieve state-of-the-art results on the GLUE benchmark using BERT, while preserving memory and accuracy. |
Copied to clipboard
| Challenge: | Existing methods to obtain text representations or embeddings with these models encoding personally identifiable information may lead to privacy leaks. |
| Approach: | They propose a novel approach which combines differential privacy and adversarial learning to preserve privacy during training of embeddings. |
| Outcome: | The proposed approach reduces private information leakage by 3% over the current method. |
Copied to clipboard
| Challenge: | Existing studies on text detoxification cast this task as style transfer . text detox requires better preservation of the original meaning, authors argue . |
| Approach: | They propose two unsupervised methods for eliminating toxicity in text . they use a paraphraser guided by style-trained language models to keep the text content . |
| Outcome: | The proposed methods yield new SOTA results. |
Copied to clipboard
| Challenge: | Text simplification is a valuable technique, but research on it is limited. |
| Approach: | They propose a document-level simplification task using Wikipedia dumps as a dataset and propose an automatic evaluation metric called D-SARI. |
| Outcome: | The proposed metric is more suitable for document-level simplification task. |
Copied to clipboard
| Challenge: | Using a pretrained sequence-to-sequence language model, we explore speaker name substitution, negation scope highlighting, multi-task learning with relevant tasks, and pretraining on in-domain data. |
| Approach: | They propose a pretrained sequence-to-sequence language model that can handle different parts of dialogue belonging to multiple speakers and combine them to produce a coherent monologue summary. |
| Outcome: | The proposed techniques outperform baseline models on a dialogue summarization dataset. |
Copied to clipboard
| Challenge: | Nominalizations are difficult to interpret because of ambiguous semantic relations between deverbal noun and its arguments. |
| Approach: | They propose to generate clausal paraphrases for nominalizations by mapping arguments to verbs . they use a contextualized language model to re-rank nominalization candidates . |
| Outcome: | The proposed task is based on a pre-trained model to re-rank paraphrase candidates identified by a textual entailment model. |
Copied to clipboard
| Challenge: | QuestEval is a metric used in text-to-text tasks, but its adaptation to Data-to Text tasks requires multimodal Question Generation and Answering systems, which are seldom available. |
| Approach: | They propose to build synthetic multimodal corpora enabling to train multimodal components for a data-QuestEval metric. |
| Outcome: | The proposed method obtains state-of-the-art correlations with human judgment on the WebNLG and WikiBio benchmarks. |
Copied to clipboard
| Challenge: | Entity linking is an important problem with many applications. |
| Approach: | They propose a method that exploits the fact that entities that are truly mentioned in a document tend to form a semantically dense subset of all candidate entities in the document. |
| Outcome: | a new method that outperforms existing methods on real-world datasets outperformed existing methods. |
Copied to clipboard
| Challenge: | Existing methods to extract entities and relations from unstructured texts are difficult to handle due to the overlapping triple problem. |
| Approach: | They propose a translation decoding schema for joint extraction of entities and relations from unstructured texts to form factual triples. |
| Outcome: | The proposed model can handle the overlapping triple problem, and is 2 times faster than the state-of-the-art models. |
Copied to clipboard
| Challenge: | Recent neural approaches to event temporal relation extraction map events to embeddings in the Euclidean space and train a classifier to detect temporal relations between event pairs. |
| Approach: | They propose to embed events into hyperbolic spaces to model hierarchical structures . they propose to use hyperbolical embeddings to directly infer event relations . |
| Outcome: | The proposed architecture is based on two approaches to encode events and their temporal relations in hyperbolic spaces. |
Copied to clipboard
| Challenge: | Recent supervised ED approaches have achieved promising performance but require large number of manually annotated event data. |
| Approach: | They propose to overfit the trigger confounder of the context and the result . they propose to intervene on the context via backdoor adjustment during training . |
| Outcome: | The proposed method significantly improves the FSED on ACE05 and MAVEN datasets. |
Copied to clipboard
| Challenge: | Term weighting schemes are widely used in Natural Language Processing and Information Retrieval. |
| Approach: | They perform an exhaustive and large-scale empirical comparison of term weighting methods in the context of keyword extraction using tf-idf. |
| Outcome: | The proposed methods have advantages over tf-idf, and qualitative differences between them. |
Copied to clipboard
| Challenge: | Various temporal knowledge graph (KG) completion models have been proposed . knowledge graphs are typically static and store facts in their current state . |
| Approach: | They propose to use temporal embeddings and a score function to model temporal knowledge graphs . they classify the temporal embedded methods into two classes: timestamp and time-dependent . |
| Outcome: | The proposed models outperform current models on ICEWS datasets with 3000 experiments and 13159 GPU hours. |
Copied to clipboard
| Challenge: | Product quantization (PQ) is a widely used technique for ad-hoc retrieval. |
| Approach: | They propose a match-oriented product quantization with a multinoulli contrastive loss objective. |
| Outcome: | The proposed method maximizes matching probability of query and ground-truth key, compared with previous methods on non-supervised datasets. |
Copied to clipboard
| Challenge: | Existing methods to generate mind-maps from text are difficult to capture the overall semantics of a document. |
| Approach: | They propose an efficient mind-map generation network that converts a document into a graph via sequence-to-graph. |
| Outcome: | The proposed network reduces inference time by thousands of times compared with existing methods and reveals key semantic structures better than plain text. |
Copied to clipboard
| Challenge: | Existing methods for text classification based on graph neural networks (GNNs) consider only one-hop neighborhoods and low-frequency information within texts, which suffer from over-smoothing issues if many graph layers are stacked. |
| Approach: | They propose a deep attention diffusion Graph Neural Network model to learn text representations by bridging the chasm of interaction difficulties between a word and its distant neighbors. |
| Outcome: | The proposed model outperforms existing methods on standard benchmark datasets on a set of textual features. |
Copied to clipboard
| Challenge: | Multi-label text classification is a challenging task because it requires capturing label dependencies. |
| Approach: | They propose to use distribution-balanced loss functions to solve label dependency problems in multi-label text classification by capturing label dependencies from a fixed-set of labels. |
| Outcome: | The proposed loss function addresses both the class imbalance and label linkage problems and outperforms other loss functions. |
Copied to clipboard
| Challenge: | a Bayesian topic regression model uses text and numerical information to model outcome variables. |
| Approach: | They propose a Bayesian Topic Regression model that uses both text and numerical information to model an outcome variable. |
| Outcome: | The proposed model recovers ground truth with lower bias than any benchmark model when text and numerical features are correlated. |
Copied to clipboard
| Challenge: | Pretrained transformer-based language models have demonstrated state-of-the-art predictive performance when adapted into a range of language understanding tasks. |
| Approach: | They propose to use salient information extracted a priori from training data to complement the task-specific information learned by the model during fine-tuning on a downstream task. |
| Outcome: | The proposed model can provide more faithful explanations across four different feature attribution methods compared to vanilla BERT. |
Copied to clipboard
| Challenge: | Existing paradigms for multi-task training involve a shared pre-trained language model and a small, thin network (head) given an input, a target head is the head that is selected for outputting the final prediction. |
| Approach: | They examine the behaviour of non-target heads when given input that belongs to a different task than the one they were trained for. |
| Outcome: | The non-target heads exhibit emergent behaviour, which may explain the target task, or generalize beyond their original task. |
Copied to clipboard
| Challenge: | Recent research has focused on adversarial text attacks on neural networks for natural language processing. |
| Approach: | They implement an algorithm inspired by zeroth order optimization-based attacks and compare it with benchmark results in TextAttack. |
| Outcome: | The proposed algorithm outperforms other black-box adversarial text attacks. |
Copied to clipboard
| Challenge: | Knowledge Graph Embeddings (KGE) are widely used for relational learning on large scale Knowledge . however, little is known about the security vulnerabilities that might disrupt their intended behaviour. |
| Approach: | They propose to use model-agnostic instance attribution methods to select adversarial deletions and a heuristic method to replace one of the two entities in each influential triple to generate adversarials. |
| Outcome: | The proposed methods outperform the state-of-the-art data poisoning attacks on KGE models and improve the MRR degradation by up to 62% over the baselines. |
Copied to clipboard
| Challenge: | Existing machine reading systems fail to answer What did Elizabeth want? correctly in the context of ‘My kingdom for a cough drop, cried Queen Elizabeth.’ Biased by co-occurrence statistics in the training data of pretrained language models, systems predict my kingdom, rather than a lung drop. |
| Approach: | They propose to use a dataset to quantify belief bias in machine reading based on pre-trained language models to examine the pervasiveness of belief bias. |
| Outcome: | The proposed dataset shows that machine reading models fail when contexts do not align with common beliefs. |
Copied to clipboard
| Challenge: | Current Transformer-based sequence-to-sequence architectures can suffer from overfitting during training. |
| Approach: | They propose to use Transformer-based sequence-to-sequence architectures to overcome overfitting problems when generating very long sequences. |
| Outcome: | The proposed model performs worse on very long sequences than previous approaches on string editing and translation tasks when faced with sequences of length diverging from the length distribution in training data. |
Copied to clipboard
| Challenge: | Recent work has raised the question of whether valid adversarial inputs are feasible. |
| Approach: | They analyze how human-generated adversarial examples compare to the best algorithms . they use crowdsourcing to modify words in an input text with immediate feedback . |
| Outcome: | The proposed algorithms are not more efficient than the best to generate natural-reading, sentiment-preserving examples. |
Copied to clipboard
| Challenge: | Evidence for the uniform information density principle has been found at many levels of language production. |
| Approach: | They propose to use the Uniform Information Density principle to test whether and within which contextual units it holds in task-oriented dialogues. |
| Outcome: | The proposed method is able to reduce fluctuations in the density of the information transmitted. |
Copied to clipboard
| Challenge: | Recent theories of language optimality have tried to justify its prevalence, arguing that homophony enables the reuse of efficient wordforms and is thus beneficial for languages. |
| Approach: | They propose a new information-theoretic quantification of a language’s homophony: the sample Rényi entropy. |
| Outcome: | The proposed method is more nuanced than either Piantadosi et al.'s or Trott and Bergen's results. |
Copied to clipboard
| Challenge: | Basic-level categories are an important psycholinguistic concept introduced by Rosch et al. . an at-scale algorithm for the automatic determination of BLC exists, but it operates without Rosch-style semantic features. |
| Approach: | They propose a method for the detection of BLC at scale that makes use of Rosch-style semantic features. |
| Outcome: | The proposed method outperforms the current SoA in detecting basic-level categories with an accuracy of 75.0% in English and 80.7% in Mandarin. |
Copied to clipboard
| Challenge: | Existing methods focus on reasoning at past timestamps to complete the missing facts, and there are only a few works of reasoning on known TKGs to forecast future facts. |
| Approach: | They propose a time-shaped reward method that captures historical knowledge graph snapshots and a new representation method for unseen entities to improve the inductive inference ability of the model. |
| Outcome: | The proposed method improves on four benchmark datasets with higher explainability, less calculation, and fewer parameters when compared with existing state-of-the-art methods. |
Copied to clipboard
| Challenge: | We introduce new pretraining losses tailored to learn generic multilingual spoken dialogue representations . goal is to expose model to code-switched language . |
| Approach: | They propose to build a pretraining corpus of multilingual conversations in five different languages from OpenSubtitles. |
| Outcome: | The proposed models perform better in monolingual and multilingual settings. |
Copied to clipboard
| Challenge: | Existing knowledge graph embeddings rely on geometric operations to model relational patterns such as symmetry and hierarchical semantics. |
| Approach: | They propose a new knowledge graph embedding model that integrates multiple geometric transformations to model multi-relational knowledge graphs. |
| Outcome: | Experiments on five datasets show that BiQUE can model symmetry, inversion, and composition. |
Copied to clipboard
| Challenge: | Existing models for temporal knowledge graphs model the temporal KGs in discrete state spaces, whereas static models model the KG in discretized state spaces. |
| Approach: | They propose a continuum model that extends the idea of neural ordinary differential equations to multi-relational graph convolutional networks and encodes both temporal and structural information into continuous-time dynamic embeddings. |
| Outcome: | The proposed model outperforms existing models on five benchmark datasets showing it can predict future links on temporal knowledge graphs. |
Copied to clipboard
| Challenge: | Backdoor attacks are a serious threat to the safety of reusing deep neural networks (DNNs). |
| Approach: | They propose an efficient online defense mechanism based on robustness-aware perturbations to distinguish poisoned and clean samples to defend against backdoor attacks on natural language processing models. |
| Outcome: | The proposed method achieves better defending performance and lower computational costs than existing defense methods. |
Copied to clipboard
| Challenge: | Recent work on word embeddings and pre-trained language models has shown the large impact of language representations on natural language processing (NLP) models across tasks and domains. |
| Approach: | They propose feature-based adversarial meta-embeddings with an attention function that is guided by word-specific properties, such as shape and frequency, to handle subword-based embeddings. |
| Outcome: | The proposed model improves performance in downstream tasks even with word embeddings from transformers. |
Copied to clipboard
| Challenge: | Existing black box search methods are inefficient as they do not consider the amount of queries required to generate adversarial attacks. |
| Approach: | They propose a query efficient attack strategy to generate plausible adversarial examples on text classification and entailment tasks. |
| Outcome: | The proposed attack reduces query count by 75% across all datasets and target models compared to prior attacks in a limited query setting. |
Copied to clipboard
| Challenge: | a new study examines whether beam search can be replaced by a more powerful metric-driven search technique. |
| Approach: | They propose a beam search method which is agnostic to the end metric and report results on a variety of metrics. |
| Outcome: | The proposed method is based on a Monte-Carlo Tree Search (MCTS) based method and shows it can be used in language applications. |
Copied to clipboard
| Challenge: | Existing document-level NMT methods fail to leverage contexts beyond a few set of previous sentences. |
| Approach: | They propose to represent a document as a graph that connects relevant contexts regardless of distances. |
| Outcome: | Experiments on IWSLT English–French, Chinese-English, WMT English–German and Opensubtitle English–Russian show that using document graphs can significantly improve translation quality. |
Copied to clipboard
| Challenge: | Recent work has highlighted several flaws of MNMT models in zero-shot scenarios where language labels are ignored and the wrong language is generated. |
| Approach: | They propose to combine explicit alignment to language labels with word alignment supervision to improve zero-shot translations. |
| Outcome: | The proposed model improves on three multilingual MT benchmarks. |
Copied to clipboard
| Challenge: | Word alignments are useful for typological research and can be used in machine translation systems. |
| Approach: | They propose to exploit the multiparallelity of parallel corpora by representing bilingual alignments as a graph and then predicting additional edges. |
| Outcome: | The proposed algorithm improves the accuracy of bilingual alignments by 28% over baseline algorithms. |
Copied to clipboard
| Challenge: | Building neural machine translation systems to perform well on a specific target domain remains a challenge. |
| Approach: | They propose to train a single NMT system per language pair that performs well across multiple domains. |
| Outcome: | The proposed approach improves the Pareto frontier on this task. |
Copied to clipboard
| Challenge: | Statistical MT decomposes the translation task into distinct components that are learned separately. |
| Approach: | They show that neural machine translation models acquire different competences over the course of training . previous work shows how to improve some of the competences in NMT by using lexical translation probabilities, phrase memories, alignment information. |
| Outcome: | The proposed model improves translation quality and word-by-word translation, while learning complex reordering patterns. |
Copied to clipboard
| Challenge: | Large scale multilingual pre-trained language models have shown promising results in zero- and few-shot cross-lingual tasks. |
| Approach: | They propose a co-tuning method that aims to learn more generalized semantic equivalences when the languages are structurally dissimilar. |
| Outcome: | The proposed method improves on cross-lingual inference and review tasks by capturing the semantic relationship in the parallel data when a few translation pairs are available. |
Copied to clipboard
| Challenge: | Existing approaches to generating additional parallel sentences are aimed at expanding the support of the empirical data distribution by generating new sentence pairs that contain infrequent words. |
| Approach: | They propose to use data augmentation techniques to generate additional parallel sentences by reversing the order of the target sentence to produce unfluent target sentences. |
| Outcome: | The proposed approach improves on six low-resource translation tasks and the baseline and over DA methods. |
Copied to clipboard
| Challenge: | Winograd schemas are well-established tools for evaluating coreference resolution and commonsense reasoning capabilities of computational models. |
| Approach: | They present a dataset of German, French, and Russian schemas aligned with their English counterparts. |
| Outcome: | The proposed model improves in English and German, while the model improve in other languages. |
Copied to clipboard
| Challenge: | Neural Machine Translation (NMT) is progressing at a rapid pace. |
| Approach: | They propose to combine two outputs so that each side depends on the other . they highlight the challenges of dual decoding and analyze the benefits of generating matched, rather than independent, translations. |
| Outcome: | The proposed system can generate matched, rather than independent, translations. |
Copied to clipboard
| Challenge: | In few-shot learning, discrete and soft prompting perform better than finetuning in multilingual cases. |
| Approach: | They show that discrete and soft prompting perform better than finetuning in crosslingual transfer and in-language training of multilingual natural language inference. |
| Outcome: | The proposed prompting model outperforms finetuning in crosslingual transfer and in-language training of multilingual natural language inference. |
Copied to clipboard
| Challenge: | Multimodal machine translation models outperform text-only models when visual context is available, but recent studies have shown that the performance of MMT models is only marginally impacted when the associated image is replaced with an unrelated image or noise. |
| Approach: | They propose to use visual data to highlight the importance of visual inputs in MMT models to enhance their leverage. |
| Outcome: | The proposed models outperform text-only models when visual context is available, but the results show that the visual context might not be exploited by the models at all. |
Copied to clipboard
| Challenge: | Multilingual NMT is an attractive solution for production, but to match bilingual quality, it comes at the cost of larger and slower models. |
| Approach: | They propose to use a shallow decoder with vocabulary filtering to speed up inference . they validate their findings with BLEU and chrF on 380 language pairs . |
| Outcome: | The proposed approach can be used in two 20-language multi-parallel settings. |
Copied to clipboard
| Challenge: | A study of multilingual fine-tuning yields better performance on downstream NLP applications . low resource languages such as Oriya and Punjabi are found to be the largest beneficiaries of multi-lingual fine tuning. |
| Approach: | They propose to leverage the relatedness of languages that belong to the same family in NLP models by multilingual fine-tuning. |
| Outcome: | The proposed approach improves performance on downstream NLP tasks by 15% compared to monolingual fine-tuning. |
Copied to clipboard
| Challenge: | Traditional hand-crafted features have been used for distinguishing between translated and original non-translated texts. |
| Approach: | They compare a feature-engineering-based approach to a features-learning-based one and use pre-trained neural word embeddings to train neural architectures. |
| Outcome: | The proposed approach outperforms other approaches by more than 20 accuracy points and the BERT-based model performs the best in both monolingual and multilingual settings. |
Copied to clipboard
| Challenge: | Neural Machine Translation suffers from a beam-search problem after a certain point, especially for long sentences. |
| Approach: | They propose a data augmentation technique that concatenates several sentences from the original dataset to make a long training example. |
| Outcome: | The proposed technique significantly reduces degradation with growing beam size and improves translation quality on the IWSTL15 En-Vi, IWStl17 En-Fr, and WMT14 En-De datasets. |
Copied to clipboard
| Challenge: | Policy compliance detection is the task of ensuring that a scenario conforms to a policy. |
| Approach: | They propose to decompose policy compliance detection into question answering . they propose to use an existing dataset to augment expert annotations . |
| Outcome: | The proposed approach improves accuracy in cross-policy setups, especially when policies are unseen in training. |
Copied to clipboard
| Challenge: | Large-scale multi-label text classification tasks often face long-tailed label distributions, where many labels have few or even no training instances. |
| Approach: | They propose a meta-learning approach that incorporates the objective of adapting to new low-resource tasks into the meta-Learning phase. |
| Outcome: | The proposed approach achieves state-of-the-art against strong baselines and can still enhance powerful BERTlike models. |
Copied to clipboard
| Challenge: | Prior work used text generation techniques or redundancy in similar passages for OCR error correction, which is not appropriate in cases of low corpus redundancies or weak document contextual information. |
| Approach: | They propose to use a pretrained language model to reconcile different OCR views in unsupervised way so that their combination contains fewer errors than each individual view. |
| Outcome: | The proposed model can reconcile multiple OCR views so that their combined version contains fewer errors than the best OCR view. |
Copied to clipboard
| Challenge: | Existing work injects lexical constraints into the output, which generates generic or ungrammatical sentences and has high computational complexity. |
| Approach: | They propose a model that incorporates pre-specified keywords into the output to control the generated text. |
| Outcome: | The proposed model decomposes the generated text into two sub-tasks and improves the sentence quality. |
Copied to clipboard
| Challenge: | Existing approaches to text moderation are reactive and do not account for user generated content. |
| Approach: | They propose a text toxicity propensity model to characterize extent to which a user generated text attracts toxic comments and introduce a beta regression model to do the probabilistic modeling. |
| Outcome: | The proposed model performs well in comprehensive experiments and is scalable. |
Copied to clipboard
| Challenge: | Existing approaches to sentence order prediction ignore the importance of document level global information, i.e., while predicting relative order of two sentences (s i , s j) other sentences sk from the same document do not play any role. |
| Approach: | They propose a framework based on graph neural networks and temporal commonsense knowledge to model global information and predict relative order of sentences. |
| Outcome: | The proposed method is naturally suitable for order prediction on five different datasets and has potential applications in the evaluation of the quality of machinegenerated documents. |
Copied to clipboard
| Challenge: | Documents as short as a single sentence may reveal sensitive information about authors . style transfer is effective but a number of current methods cause a drop in down-stream utility . |
| Approach: | They propose a method to remove sensitive information from documents by multilingual back-translation using off-the-shelf translation models. |
| Outcome: | The proposed method lowers adversarial gender and race prediction by 22% while retaining 95% of original utility on downstream tasks. |
Copied to clipboard
| Challenge: | Pre-trained models for Natural Languages (NL) like BERT and GPT have been shown to transfer well to Programming Languages. |
| Approach: | They propose a unified pre-trained encoder-decoder Transformer model that leverages the code semantics conveyed from the developer-assigned identifiers. |
| Outcome: | The proposed model outperforms existing models on understanding and generation tasks and can capture semantic information from code. |
Copied to clipboard
| Challenge: | Existing methods for detecting health outcomes from text ignore global structural correspondences between sentence-level and word-level information present in a given text. |
| Approach: | They propose a method that uses both word-level and sentence-level information to perform outcome span detection and outcome type classification. |
| Outcome: | The proposed method consistently outperforms decoupled methods, reporting competitive results. |
Copied to clipboard
| Challenge: | a multi-class grammatical error detection system can be used to improve grammamatical errors correction (GEC) for English. |
| Approach: | They develop a multi-class grammatical error detection system based on pre-trained ELECTRA and extend it to multi-Class detection using different error type tagsets. |
| Outcome: | The proposed system outperforms previous systems on the BEA-test benchmark. |
Copied to clipboard
| Challenge: | Existing language models can be refined for zero-shot commonsense reasoning . however, commons sense reasoning is still an unsolved problem . |
| Approach: | They propose a self-supervised learning approach that refines a pre-trained language model to boost conceptualization. |
| Outcome: | The proposed approach boosts conceptualization by utilizing loss landscape refinement. |
Copied to clipboard
| Challenge: | Existing methods to select transfer sources are limited by text and task similarity, which limits their application in transfer settings where both the task and the text domain change. |
| Approach: | They propose a model similarity measure that represents text and task similarity jointly to automatically determine which and how many sources to exploit. |
| Outcome: | The proposed approach improves performance by 24 F1 points for predicting promising sources across domains and tasks with similar models. |
Copied to clipboard
| Challenge: | Contextualised word embeddings are powerful tool to detect contextual synonyms, but most of the current SOTA methods are supervised and underexploit the potential of the context. |
| Approach: | They propose a self-supervised approach which detects contextual synonyms of concepts being trained on the data created by shallow matching. |
| Outcome: | The proposed approach outperforms the previous SOTA with gains of up to 4.5 and 4.0 absolute points after fine-tuning with as little as 20% of the labelled data. |
Copied to clipboard
| Challenge: | Contracts are a common type of legal document that frequent in business workflows, but there has been limited NLP research in understanding and generating them. |
| Approach: | They propose a task of clause recommendation to help automate contract authoring . they first predict if a specific clause type is relevant to be added in a contract . then they propose two-staged pipeline to recommend top clauses based on the contract context . |
| Outcome: | The proposed pipeline predicts if a clause type is relevant to be added in a contract and recommends the top clauses for the given type based on the contract context. |
Copied to clipboard
| Challenge: | Finnish is a language with multiple dialects that differ in accent, morphological forms and lexical choice. |
| Approach: | They propose an approach to automatically detect the dialect of a speaker based on a transcript and transcript with audio recording in a dataset consisting of 23 different dialects. |
| Outcome: | The proposed method achieves 57% accuracy, compared to 85% accuracy for text and audio. |
Copied to clipboard
| Challenge: | a survey of English Machine Reading Comprehension datasets is carried out . the aim is to provide a concise yet informative overview of the landscape . |
| Approach: | They survey 60 English Machine Reading Comprehension datasets to provide a resource for other researchers interested in this problem. |
| Outcome: | The proposed survey covers 60 English MRC datasets with a view to providing a resource for other researchers interested in the problem. |
Copied to clipboard
| Challenge: | Existing models that handle single-entity questions have focused on relation following . introducing intersection improves performance on multiple-entities questions by over 14% . |
| Approach: | They propose a model that explicitly handles multiple-entity questions by implementing an intersection operation. |
| Outcome: | The proposed model improves on multiple-entity questions by over 14% on two datasets . it also improves performance on questions with multiple entities by 19% . |
Copied to clipboard
| Challenge: | We present a new approach for weakly-supervised conversational Question Answering over Knowledge Graphs . |
| Approach: | They propose a Logical Form grammar that can model a wide range of queries on a Knowledge Graph while remaining sufficiently simple to generate supervision data efficiently. |
| Outcome: | The proposed grammar can model a wide range of queries while remaining simple to generate supervision data efficiently. |
Copied to clipboard
| Challenge: | a new approach to generate adversarial data is needed to improve question answering models . crowdworkers can fool a model only 8.8% of the time, compared to 17.6% for a trained model without synthetic data. |
| Approach: | They develop a pipeline that generates questions and then filters or labels them to improve quality. |
| Outcome: | The proposed approach improves state-of-the-art on a human-written adversarial dataset by 3.7F1 and improves model generalisation on nine of the twelve MRQA datasets. |
Copied to clipboard
| Challenge: | Pretrained language models can produce inconsistent answers when probed, even after specialized training. |
| Approach: | They propose to embed a pretrained language model in a broader system that includes an evolving, symbolic memory of beliefs that records but may modify the raw PTLM answers. |
| Outcome: | The proposed architecture improves belief consistency in the overall system by revising beliefs that clash with others and generating queries using known beliefs as context. |
Copied to clipboard
| Challenge: | Question Answering (QA) is a branch of QA that enables effective perceiving, accessing, and understanding complex biomedical knowledge by innovative applications. |
| Approach: | They present MLEC-QA, the largest-scale Chinese multi-choice biomedical QA dataset . they implement eight representative control methods and open-domain QA methods as baselines . |
| Outcome: | The proposed dataset is the largest-scale Chinese multi-choice biomedical QA dataset . it covers the following biomedically-relevant sub-fields: Clinic, Stomatology, Public Health, Traditional Chinese Medicine, and Traditional Chinese medicine Combined with Western Medicine. |
Copied to clipboard
| Challenge: | Lack of publicly available NLG benchmarks for low-resource languages poses a challenge . authors show that IndoBART and IndoGPT achieve competitive performance on all tasks . |
| Approach: | They propose a benchmark to measure natural language generation progress in three low-resource languages of Indonesia . they use a corpus of pretraining datasets to build their models . |
| Outcome: | The proposed benchmark measures progress in Indonesian, Javanese, and Sundanese . the results highlight the importance of pretraining on closely related, localized languages . |
Copied to clipboard
| Challenge: | Existing models for multi-hop reasoning are not able to evaluate their interpretability . a recent study found that many paths are unreasonable . |
| Approach: | They propose a framework to evaluate the interpretability of multi-hop reasoning models . they annotate all possible rules and establish a benchmark . |
| Outcome: | The proposed framework outperforms existing models in terms of performance and interpretability. |
Copied to clipboard
| Challenge: | Evaluation metrics are a key ingredient for progress of text generation systems . a class of novel evaluation metrics based on BERT and its variants has been explored . |
| Approach: | They propose to disentangle BERT-based evaluation metrics along linguistic factors . they show they are sensitive to lexical overlap, just like BLEU and ROUGE . |
| Outcome: | The proposed metrics capture all aspects but are sensitive to lexical overlap, just like BLEU and ROUGE, the authors show . |
Copied to clipboard
| Challenge: | Existing text-to-SQL models do not generalize when faced with domain knowledge that does not frequently appear in training data. |
| Approach: | They propose a human-curated dataset based on the Spider benchmark for text-to-SQL translation. |
| Outcome: | The proposed model performs better on unseen domains than existing models on public benchmarks. |
Copied to clipboard
| Challenge: | Existing studies have shown that human evaluations in NLP are under-powered because of two common factors: they treat ordinal data as interval data and operate under high variance settings. |
| Approach: | They propose to use ordinal mixed effects models to detect small differences between models, especially in high variance settings common in NLP evaluations of generated texts. |
| Outcome: | The proposed models detect small differences in high variance settings, especially in high-variance evaluations of generated texts. |
Copied to clipboard
| Challenge: | Recent years have seen an increasing need for gender-neutral and inclusive language. |
| Approach: | They propose a rule-based and a neural approach to gender-neutral rewriting for English . they use manually curated synthetic and natural data to train a rewriter . |
| Outcome: | The proposed approach improves on the rule-based approach with word error rates below 0.18% on synthetic, in-domain and out-domain test sets. |
Copied to clipboard
| Challenge: | Existing evaluations on the population task are either not accurate (automatic evaluation with randomly sampled negative examples) or of small scale (human annotation). |
| Approach: | They propose a reasoning over commonsense knowledge bases (CSKBs) that are free-text and have a human annotation set to probe commonsensical reasoning. |
| Outcome: | The proposed model is based on a human-annotated evaluation set and is compared with existing models on the population task. |
Copied to clipboard
| Challenge: | Existing similarity-based systems focus on learning sense embeddings using only the sentence where the word appears, neglecting its global context. |
| Approach: | They propose a contextoriented embedding technique that takes better advantage of both word-level and sense-level global context of an ambiguous word for disambiguation. |
| Outcome: | The proposed method improves on all-words WSD benchmarks in knowledge-based category by large margins. |
Copied to clipboard
| Challenge: | Existing approaches to parse text-to-SQL data are lacking labeled data for unseen evaluation databases. |
| Approach: | They propose a framework for enhancing SQL queries by automatically producing large numbers of SQL queries based on an abstract syntax tree grammar. |
| Outcome: | The proposed framework can produce high-quality natural language questions over strong baselines. |
Copied to clipboard
| Challenge: | Using annotated datasets is difficult as it requires query-language expertise. |
| Approach: | They propose a crowdsourcing pipeline to annotate natural language questions using intermediate question representations. |
| Outcome: | The proposed pipeline reduces the burden of annotating a large dataset with queries by using intermediate question representations. |
Copied to clipboard
| Challenge: | Existing embedding-based approaches disregard time information that exists in many large-scale knowledge graphs, leaving much room for improvement. |
| Approach: | They propose a Time-aware Entity Alignment approach that incorporates relation and time information into a vector space and uses Graph Neural Networks to learn entity representations. |
| Outcome: | The proposed approach outperforms the state-of-the-art methods on real-world TKG datasets due to the inclusion of time information. |
Copied to clipboard
| Challenge: | Stance detection is a task that focuses on the classification of a writer’s viewpoint towards a target. |
| Approach: | They propose an end-to-end unsupervised framework for out-of-domain prediction of unseen, user-defined labels. |
| Outcome: | The proposed framework shows that it can be used to predict unseen labels over strong baselines. |
Copied to clipboard
| Challenge: | Data augmentation aims to alleviate the overfitting issue in low-resource or class-imbalanced situations. |
| Approach: | They propose a framework called Text AutoAugment to enhance training samples . they use a Bayesian optimization algorithm to search for the best policy . |
| Outcome: | The proposed framework outperforms baseline methods on six benchmark datasets. |
Copied to clipboard
| Challenge: | Pre-trained language models capture a surprisingly rich amount of lexical knowledge, but it is unclear to what extent relation embeddings can be used to encode relational knowledge. |
| Approach: | They found that word vector differences capture lexical relations . relationship embeddings can be used to encode relational knowledge . |
| Outcome: | The results are highly competitive on analogy (unsupervised) and relation classification (supervised) benchmarks, even without any task-specific fine-tuning. |
Copied to clipboard
| Challenge: | Recent prompt-based approaches allow pretrained language models to achieve strong performances on few-shot finetuning by reformulating downstream task instances as a language modeling problem. |
| Approach: | They propose to reformulate downstream tasks as a language modeling problem and add a regularization that preserves pretraining weights to the model to mitigate the destructive tendency of few-shot finetuning. |
| Outcome: | The proposed model performs better on low data regimes than the standard model on few-shot finetuning. |
Copied to clipboard
| Challenge: | Abstract Meaning Representations (AMR) represents sentence meaning as a directed acyclic graph. |
| Approach: | They propose to treat alignment and segmentation as latent variables and induce them as part of end-to-end training. |
| Outcome: | The proposed model achieves significant performance gains over a 'greedy' segmentation heuristic. |
Copied to clipboard
| Challenge: | Neural Word Sense Disambiguation (WSD) uses pre-existing knowledge, but only close neighbors influence prediction. |
| Approach: | They propose to exploit WordNet graphs to improve a classification model by recomputing logits . they incorporate an online neural approximated PageRank to refine edge weights . |
| Outcome: | The proposed method improves the current state of the art in the field of Neural Word Sense Disambiguation (WSD) the proposed method exploits the global graph structure while keeping space requirements linear in the number of edges. |
Copied to clipboard
| Challenge: | Existing multilingual sentence embedding models require large parallel corpora to learn efficiently, limiting their scope. |
| Approach: | They propose a sentence embedding framework based on an unsupervised loss function . they capture semantic similarity and relatedness between sentences using a multi-task loss function. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on STS, BUCC and Tatoeba benchmarks and on a monolingual benchmark. |
Copied to clipboard
| Challenge: | Pre-training Masked Language Models (MLMs) on massive datasets is expensive, but it is performed for each domain or task individually and is resource-demanding. |
| Approach: | They propose a method for more efficient adaptation that focuses on predicting words with large weights of the Naive Bayes classifier trained for the task at hand. |
| Outcome: | The proposed method improves sentiment analysis by focusing on predicting words with large weights of the Naive Bayes classifier trained for the task at hand. |
Copied to clipboard
| Challenge: | Unlabeled data are useful for few-shot learning of language models. |
| Approach: | They propose a prompt-based few-shot learner that uses unlabeled data to fine-tune language models. |
| Outcome: | The proposed approach outperforms state-of-the-art models on six sentence classification and six sentence-pair classification benchmarking tasks. |
Copied to clipboard
| Challenge: | Recent studies suggest that it is impossible to learn meaning from surface form alone. |
| Approach: | They propose to develop triadic systems that combine neural and symbolic methods to provide a seamless information flow between them. |
| Outcome: | The proposed systems combine the strengths of neural and symbolic methods to achieve a seamless information flow between them. |
Copied to clipboard
| Challenge: | Existing approaches to modulate one modal feature to another are lacking in multimodal representation learning. |
| Approach: | They propose to use unimodal and crossmodal refinement networks to enhance uni and cross-modal representations by iterative updating of distributions with transformer-based attention layers to refine modality-specific learning. |
| Outcome: | The proposed network outperforms state-of-the-art techniques on MOSI and MOSEI datasets. |
Copied to clipboard
| Challenge: | YASO contains 2,215 English sentences from dozens of review domains, annotated with target terms and their sentiment. |
| Approach: | They propose a new TSA evaluation dataset of open-domain user reviews in English . YASO contains 2,215 English sentences annotated with target terms and their sentiment . |
| Outcome: | The proposed dataset verifies the reliability of the annotations and explores the characteristics of the collected data. |
Copied to clipboard
| Challenge: | Current methods for extracting opinion words for an aspect in text leverage position embeddings to capture relative position of word to the target. |
| Approach: | They propose to use pretrained word embeddings to extract opinion words for a given aspect in text. |
| Outcome: | The proposed methods outperform current methods on a task based on pre-trained word embeddings and position embedders. |
Copied to clipboard
| Challenge: | Existing work on multimodal sentiment analysis relies on back-propagated task loss or geometric property of feature spaces to produce favorable fusion results. |
| Approach: | They propose a framework which hierarchically maximizes the Mutual Information (MI) in unimodal input pairs and between multimodal fusion result and unimod input to maintain task-related information through multimodal integration. |
| Outcome: | The proposed framework maximizes the Mutual Information (MI) in unimodal input pairs and between multimodal fusion result and unimodulated input to maintain task-related information through multimodal integration. |
Copied to clipboard
| Challenge: | Existing approaches to Aspect-based sentiment classification ignore sequential features of context and lack syntactic knowledge of sentences. |
| Approach: | They propose a model which integrates sequential grammatical features from context and syntactic knowledge from dependency graphs to augment GCN to better encode dependency graph outputs. |
| Outcome: | The proposed model outperforms state-of-the-art models when equipped with contextual word embedding from pre-training language models. |
Copied to clipboard
| Challenge: | Using social features, we hypothesize that comments from the ambient community can either affirm the original view or implicitly exert pressure to change it. |
| Approach: | They propose a structured model to capture the ambient community’s sentiment towards the discussion and its effect on persuasion. |
| Outcome: | The proposed model captures the ambient community’s sentiment towards the discussion and its effect on persuasion. |
Copied to clipboard
| Challenge: | Existing studies focus on predicting the four elements in one shot, instead of predicting them all. |
| Approach: | They propose a task to jointly detect all sentiment elements in quads for a given opinionated sentence. |
| Outcome: | The proposed method can generate the semantics of the sentiment elements in the natural language form. |
Copied to clipboard
| Challenge: | Existing studies on Aspect-based sentiment analysis (ABSA) focus on English texts, but handling it in resource-poor languages remains a challenge. |
| Approach: | They propose an unsupervised cross-lingual transfer method for the Aspect-based sentiment analysis task . they propose an aspect code-switching mechanism to augment training data with code-linked bilingual sentences . |
| Outcome: | The proposed method preserves task-specific knowledge in the target language. |
Copied to clipboard
| Challenge: | Existing representation schemes for emotion analysis are based on label formats, natural languages, and even disparate model architectures. |
| Approach: | They propose a training scheme that learns a shared latent representation of emotion independent from different label formats, natural languages, and even disparate model architectures. |
| Outcome: | The proposed model performs well on a wide range of datasets without penalizing prediction quality. |
Copied to clipboard
| Challenge: | Existing methods for unsupervised text style transfer struggle to achieve high style conversion rate and low content loss. |
| Approach: | They propose a collaborative learning framework for unsupervised text style transfer using a pair of bidirectional decoders. |
| Outcome: | The proposed framework achieves strong empirical results on style compatibility and content preservation. |
Copied to clipboard
| Challenge: | Existing methods for text style transfer use autoregressive decoding, but they are slow and low parallelizability. |
| Approach: | They propose a base NAR model by directly adapting the common training scheme from its AutoRegressive counterpart. |
| Outcome: | The proposed model sacrifices performance due to lack of conditional dependence between output tokens . knowledge distillation, contrastive learning, and iterative decoding are employed to improve the model . |
Copied to clipboard
| Challenge: | Existing methods for tagging opinion triplets fail to capture the strong interdependence between the three opinion factors, whereas grid tabbing fails to capture span-level semantics while predicting sentiment between an aspect-opinion pair. |
| Approach: | They propose a tagging-free approach to extracting opinion triplets using a pointer network decoding framework that captures the interdependence between the three elements of an opinion triple. |
| Outcome: | The proposed architecture captures the interdependence between the aspect and opinion triplets while predicting their connecting sentiment. |
Copied to clipboard
| Challenge: | Temporal sentence localization in videos is an important yet challenging task in natural language processing. |
| Approach: | They propose an Adaptive Proposal Generation Network to maintain the segment-level interaction while speeding up the efficiency. |
| Outcome: | The proposed model outperforms state-of-the-art methods on three challenging benchmarks. |
Copied to clipboard
| Challenge: | Existing methods to learn effective alignment between vision and language features are insufficient in practice due to complicated multi-step reasoning. |
| Approach: | They propose an iterative alignment network which iterates inter- and intra-modal features within multiple steps for more accurate grounding. |
| Outcome: | The proposed model performs better than the state-of-the-arts on three challenging benchmarks. |
Copied to clipboard
| Challenge: | Pretrained language models demonstrate strong performance in most NLP tasks when fine-tuned on small task-specific datasets. |
| Approach: | They propose a two-stage procedure to learn from a small set of demonstrations and a simple reinforcement learning algorithm to improve by interacting with an environment. |
| Outcome: | The proposed method improves with only 1.2% of the demonstrations and a simple reinforcement learning algorithm over existing methods in the ALFWorld environment. |
Copied to clipboard
| Challenge: | Existing work on change captioning uses a natural language sentence to describe disagreement between two images. |
| Approach: | They propose a Relation-embedded Representation Reconstruction Network to distinguish real change from clutter and irrelevant changes. |
| Outcome: | The proposed method achieves state-of-the-art on two public datasets. |
Copied to clipboard
| Challenge: | State-of-the-art systems generate questions that sound unnatural to humans and are grammatically correct. |
| Approach: | They propose to use beam search re-ranking to generate a model that guides an effective goal-oriented strategy by asking questions that confirm the model’s conjecture about the referent. |
| Outcome: | The proposed model is more natural and effective than beam search decoding without re-ranking on the GuessWhat?! game. |
Copied to clipboard
| Challenge: | Adapting a model to target speakers requires a lot of compute and may cause catastrophic forgetting to the existing speakers. |
| Approach: | They propose a unified speaker adaptation approach consisting of feature adaptation and model adaptation. |
| Outcome: | The proposed model outperforms baseline models with 20.58% relative WER reduction and surpasses finetuning method by 2.54% on target speaker adaptation. |
Copied to clipboard
| Challenge: | Existing methods for classifying memes are difficult to perform, with human accuracy only about 85% . recent state-of-the-art models perform considerably less accurately, achieving up to 64.73% accuracy. |
| Approach: | They propose to use an off-the-shelf caption generator to capture the first image and overlayed text. |
| Outcome: | The proposed tool improves classification accuracy for unimodal and multimodal models . the proposed tool can be used to model the contrast between image content and overlayed text . |
Copied to clipboard
| Challenge: | Training and inference using large transformer models can be computationally expensive because the self-attention's time and memory grow quadratically with sequence length. |
| Approach: | They propose a modified transformer architecture that constrains the encoder-decoder attention mechanism to a subset of input sentences while maintaining system performance. |
| Outcome: | The proposed architecture can be trained and inferenced using large transformer models with expensive training and induction costs. |
Copied to clipboard
| Challenge: | Inductive transfer learning has taken the entire NLU field by storm, with models such as BERT and BART setting new state-of-the-art on countless tasks. |
| Approach: | They introduce a large-scale pretrained seq2seq model for French that is very competitive with state-of-the-art BERT-based French language models such as CamemBERT and FlauBERT. |
| Outcome: | The proposed model outperforms existing models on discriminative and generative tasks on a French summarization dataset. |
Copied to clipboard
| Challenge: | Abstractive summarization is one of the areas influenced by pre-trained language models. |
| Approach: | They propose a Transformer-based encoder-decoder model pre-trained with three novel objectives to address this issue. |
| Outcome: | The proposed model outperforms previous models on six Persian summarization tasks . it also outperformed previous models in textual entailment, question paraphrasing, and question answering . |
Copied to clipboard
| Challenge: | Recent years have witnessed increased interest in abstractive summarisation thanks to the popularity of neural network models and the availability of datasets containing hundreds of thousands of document-summary pairs. |
| Approach: | They propose to create a cross-lingual summarisation corpus with long documents in a source language associated with multi-sentence summaries in . target language. |
| Outcome: | The proposed task can be applied to several other languages and covers twelve languages and directions. |
Copied to clipboard
| Challenge: | supervised summarization has been traditionally approached with unsupervised, weakly-supervised and few-shot learning techniques. |
| Approach: | They propose to combine a large dataset of opinion summaries with user reviews to form a supervised summarizer. |
| Outcome: | The proposed method improves the quality of summarization and reduces hallucinations in the summarizer. |
Copied to clipboard
| Challenge: | Abstractive summarization models have been proven effective in creating fluent and informative summaries, but they suffer from the short-range dependency problem, causing them to produce summary that miss the key points of document. |
| Approach: | They propose a neural topic model empowered with normalizing flow to capture global semantics of the document and integrate them into the summarization model. |
| Outcome: | The proposed model outperforms state-of-the-art summarization models on five common text summarizing datasets, namely CNN/DailyMail, XSum, Reddit TIFU, arXiv, and PubMed. |
Copied to clipboard
| Challenge: | Pre-trained word embeddings and self-training have been used in dependency parsing tasks for years. |
| Approach: | They compare tri-training and pretrained word embeddings in dependency parsing . they use language-specific FastText and ELMo embedds and multilingual BERT embedders . |
| Outcome: | The proposed methods are tri-training and pretrained word embeddings. |
Copied to clipboard
| Challenge: | Existing methods for zero-shot cross-domain slot filling do not achieve effective knowledge transfer to the target domain. |
| Approach: | They propose a novel approach based on prototypical contrastive learning and a dynamic label confusion strategy for zero-shot slot filling. |
| Outcome: | The proposed model improves on unseen slots while setting new state-of-the-arts on slot filling task. |
Copied to clipboard
| Challenge: | Existing methods to integrate neural networks and symbolic rules have their merits and weaknesses. |
| Approach: | They propose to integrate regular expressions into neural networks for a slot filling task . they use finite-state transducers to convert regular expression into a neural network . their model has superior zero-shot and few-shot performance . |
| Outcome: | The proposed model outperforms rules in zero-shot and few-shot scenarios and is competitive when training data is available. |
Copied to clipboard
| Challenge: | a meta-analysis of published studies shows that the causal direction of data collection can explain some trends in NLP . semi-supervised learning and domain adaptation performance differ on a number of tasks . |
| Approach: | They argue that the causal direction of the data collection process has nontrivial implications . authors categorize common NLP tasks according to their causal direction . they also empirically assay the validity of the ICM principle for text data . |
| Outcome: | The proposed model can explain differences in semi-supervised learning and domain adaptation performance across settings. |
Copied to clipboard
| Challenge: | Recent pretrained language models extend from millions to billions of parameters. |
| Approach: | They propose a technique which forwards on a whole network while backwarding on resetting the gradients of the non-child network during the backward process. |
| Outcome: | The proposed technique outperforms the vanilla fine-tuning technique on various downstream tasks and can achieve better generalization performance by large margins. |
Copied to clipboard
| Challenge: | Knowledge Graph Embeddings (KGEs) map entities and relations from knowledge graphs into a geometric space. |
| Approach: | They propose a neuro differential KGE that embeds nodes of a KG on the trajectories of Ordinary Differential Equations (ODEs) they represent each relation (edge) in a knowledge graph as a vector field on several manifolds. |
| Outcome: | The proposed model can preserve graph characteristics including structural aspects and semantics and avoid wrong inferences. |
Copied to clipboard
| Challenge: | Existing approaches to weakly supervised training lack labeled data . weakly-supervised training can result in heuristic but noisy labels . |
| Approach: | They propose a scheme that allows to control influence of signals associated with specific labeling functions. |
| Outcome: | The proposed scheme improves results compared to weakly supervised learning with a pre-trained transformer language model and a feature-based baseline. |
Copied to clipboard
| Challenge: | Backdoor attacks can manipulate the output of deep neural networks and possess high insidiousness. |
| Approach: | They propose a textual backdoor defense based on outlier word detection that can handle all the textual attacks. |
| Outcome: | The proposed method can handle all the textual backdoor attack situations. |
Copied to clipboard
| Challenge: | Existing approximations of dot-product attention ignore the value vectors . a value-aware objective outperforms an optimal approximate that ignores values . |
| Approach: | They propose an approximation of a value-aware objective that substantially outperforms an optimal approximate that ignores values. |
| Outcome: | The proposed value-aware objective outperforms an optimal approximation that ignores values in the context of language modeling. |
Copied to clipboard
| Challenge: | Existing question generation methods rely on large amounts of synthetically generated datasets and costly computational resources. |
| Approach: | They propose a framework for domain adaptation that combines question generation and domain-invariant learning to answer out-of-domain questions in settings with limited text corpora. |
| Outcome: | The proposed framework improves on state-of-the-art questions in a domain with limited text corpora. |
Copied to clipboard
| Challenge: | Using human-labeled examples, case-based reasoning can solve complex problems from scratch . case-Based reasoning is a paradigm that is used to solve complex problem . |
| Approach: | They propose a neuro-symbolic CBR approach for question answering over large knowledge bases. |
| Outcome: | The proposed approach outperforms the current state of the art on a CWQ dataset by 11% on accuracy. |
Copied to clipboard
| Challenge: | Open-domain question answering uses evidence retrieved from large corpus to answer questions . state-of-the-art approaches require intermediate evidence annotations for training . however, such intermediate annotations are expensive and methods that rely on them cannot transfer to the more common setting . |
| Approach: | They propose an open-domain question answering approach that alternately finds evidence from an up-to-date model and encourages the model to learn the most likely evidence. |
| Outcome: | The proposed approach improves over weak retrievers on multi-hop and single-hop benchmarks without using evidence labels. |
Copied to clipboard
| Challenge: | A flaw in QA evaluation is that annotations often only provide one answer . therefore, model predictions semantically equivalent to the answer but superficially different are considered incorrect. |
| Approach: | They explore using alias entities from knowledge bases to extract additional answers . they incorporate additional answers for evaluation and model training with equivalent answers based on the results . |
| Outcome: | The proposed solution improves the accuracy of evaluation with additional answers and improves model training with equivalent answers. |
Copied to clipboard
| Challenge: | Despite substantial overlap, subtle but significant distinctions exert an outsize influence on research . one paradigm values creating more intelligent QA systems, the other paradigm values building QA system that appeals to users. |
| Approach: | They propose to use the Cranfield and Manchester paradigms to describe research working towards building human-like, intelligent QA systems. |
| Outcome: | The proposed paradigms are based on the findings of two recent studies on question answering (QA) the Cranfield paradigm is not new, but the Manchester paradigm is christened as the most eclectic in QA . |
Copied to clipboard
| Challenge: | Numerical reasoning based machine reading comprehension models have achieved near-human performance on a variety of benchmarks, but are they capable of learning to reason? |
| Approach: | They propose to use a DROP benchmark to measure machine reading comprehension and investigate models that have achieved near-human performance over standard metrics. |
| Outcome: | The DROP benchmark has inspired the design of specialized BERT and embedding the results into a specialized model. |
Copied to clipboard
| Challenge: | Existing knowledge base population systems require a machine translation task to generate multiple facts, but the fact order is not considered. |
| Approach: | They propose a knowledge base population task that aims to discover facts about entities from texts and expand a KB with these facts. |
| Outcome: | The proposed networks achieve state-of-the-art (SoTA) performance on two benchmark datasets. |
Copied to clipboard
| Challenge: | Existing methods for relation extraction ignore the incompleteness of existing knowledge bases . current methods are too weak and cause noises when training and testing are not based on training data. |
| Approach: | They propose a method to automatically align unstructured text with relation instances in a knowledge base . they use heuristics to leverage the memory mechanism of deep neural networks to find out possible FN samples . |
| Outcome: | Experiments on two wildly-used benchmark datasets show the effectiveness of the proposed method. |
Copied to clipboard
| Challenge: | Existing methods for entity set expansion define the expansion boundary using seed-based distance metrics, which are hard to adjust due to the extremely sparse supervision. |
| Approach: | They propose a new learning method for bootstrapping which jointly models the bootstraping process and boundary learning process in a GAN framework. |
| Outcome: | The proposed method achieves the new state-of-the-art performance for entity set expansion. |
Copied to clipboard
| Challenge: | Information Extraction (IE) aims to extract structural information from unstructured texts. |
| Approach: | They propose a framework that aims to uncover the main causalities behind data in the view of causal inference. |
| Outcome: | The proposed framework can detect the main causalities behind data in the view of causal inference. |
Copied to clipboard
| Challenge: | Open Information Extraction (OpenIE) aims to discover textual facts from a given sentence. |
| Approach: | They propose a non-autoregressive framework that generates a fact graph and a graph with an edge linking two nodes that belong to the same fact. |
| Outcome: | The proposed framework outperforms current state-of-the-art methods on two benchmark datasets and significantly outperformed the existing ones. |
Copied to clipboard
| Challenge: | Existing methods for open relation extraction (OpenRE) are designed for predefined relations, which cannot deal with new emerging relations in the real world. |
| Approach: | They propose a relation-oriented clustering model that leverages readily available labeled data to learn a relationship-oriented representation. |
| Outcome: | The proposed model reduces error rate by 29.2% and 15.7% on two datasets compared with current SOTA methods. |
Copied to clipboard
| Challenge: | Existing methods for generating explanatory notes for language learners are inadequate . nagata et al. demonstrates that neural-retrieval-based methods can generate feedback comments for preposition use . |
| Approach: | They investigate three different methods for generating feedback comments for preposition use . grammatical and writing items can also be used to generate feedback comments . |
| Outcome: | The proposed methods outperform neural-retrieval-based methods in generating feedback comments for preposition use. |
Copied to clipboard
| Challenge: | Recent efforts to predict chatbot failure hatches vital apprehensions due to complexity of human conversation. |
| Approach: | They propose a model that integrates dialogue satisfaction estimation and handoff prediction in one multi-task learning framework. |
| Outcome: | The proposed model integrates dialogue satisfaction estimation and handoff prediction in one multi-task learning framework. |
Copied to clipboard
| Challenge: | Notable PLMs are available for text classification tasks, but performance of PLM on downstream tasks may be limited by the availability of training set. |
| Approach: | They propose a meta-learning framework to learn the transferable knowledge across tasks using PLMs. |
| Outcome: | The proposed framework outperforms baselines on seven datasets and is task-agnostic and unbiased. |
Copied to clipboard
| Challenge: | Knowledge graph inference has been studied extensively due to its wide applications. |
| Approach: | They propose a framework that restricts logical rules to be definite Horn rules and can exploit the knowledge in logical rule-based reasoning and KGE in an extremely efficient way. |
| Outcome: | The proposed framework can exploit the knowledge in logical rules and improve KGE in an extremely efficient way. |
Copied to clipboard
| Challenge: | Existing methods to improve the learning of data-scarce target domains have negative transfer due to the data distributions between source and target domain. |
| Approach: | They propose a method that uses a reinforced selector to select helpful data for transfer learning and a Wasserstein-based discriminator to maximize the distance between the selected data and target data. |
| Outcome: | The proposed method performs better on three real-world text mining tasks. |
Copied to clipboard
| Challenge: | Existing work performs code repair and commit message generation independently. |
| Approach: | They propose a cascaded method to repair program codes and generate commit messages in a unified framework. |
| Outcome: | The proposed model significantly outperforms baselines on a buggy-fixed-commit dataset. |
Copied to clipboard
| Challenge: | a recent study shows that late-interaction methods trade off retrieval accuracy and efficiency by exploiting cross-modal interactions only in the late stage. |
| Approach: | They propose an inflating and shrinking approach to exploit cross-modal interactions . they inflate code inputs and shrink code outputs to exploit interactions progressively . |
| Outcome: | The proposed method exploits cross-modal interactions in the late stage to achieve retrieval speed. |
Copied to clipboard
| Challenge: | Existing methods for video grounding are not end-to-end, i.e., they rely on time-consuming post-processing steps to refine predictions. |
| Approach: | They propose an end-to-end multi-modal Transformer model that uses two encoders and a cross-modal decoder for grounding prediction. |
| Outcome: | The proposed model is 4.9% faster than existing models and is based on a set of encodings and decoders. |
Copied to clipboard
| Challenge: | Existing benchmarks for visual-grounded models have focused on synthetic images . et al., 2018: compositional generalization is crucial for building models that generalize to new settings. |
| Approach: | They propose a test-bed for visually-grounded compositional generalization with real images. |
| Outcome: | The proposed test-bed enables compositional splits where models need to generalize to new concepts and compositions in a zero- or few-shot setting. |
Copied to clipboard
| Challenge: | Pretrained vision-and-language BERTs aim to learn representations that combine information from both modalities. |
| Approach: | They propose a diagnostic method based on cross-modal input ablation to assess the extent to which pretrained models integrate cross-module information. |
| Outcome: | The proposed method evaluates the model's performance on the other modality based on inputs from one or both modality. |
Copied to clipboard
| Challenge: | Existing methods for data augmentation involve performing mathematical operations over the raw input samples or their latent states representations, but these operations are performed in the Euclidean space, simplifying these representations and resulting in noisy interpolations. |
| Approach: | They propose a model-, data-, and modality-agnostic interpolative data augmentation technique operating in the hyperbolic space that captures the complex geometry of input and hidden state hierarchies better than its contemporaries. |
| Outcome: | The proposed technique outperforms state-of-the-art methods on benchmark and low resource datasets across speech, text, and vision modalities. |
Copied to clipboard
| Challenge: | Existing studies only consider a single event sequence corresponding to one common protagonist. |
| Approach: | They propose a Transformer-based model which integrates deep event-level and script-level information for script event prediction. |
| Outcome: | The proposed model is superior to existing models on the New York Times corpus . it utilizes rich information in the text to obtain more comprehensive representations . |
Copied to clipboard
| Challenge: | Existing approaches to consolidate textual inputs are difficult to implement . a recent study aims to capture content overlap by combining multiple textual elements . |
| Approach: | They propose to align predicate-argument relations across texts to represent content overlap . their setting exploits QA-SRL, utilizing question-answer pairs to capture predicates . |
| Outcome: | The proposed task captures content overlap beyond lexical similarity and complements cross-document coreference with proposition-level links, offering potential use for downstream tasks. |
Copied to clipboard
| Challenge: | Large pre-trained language models for textual data have an unconstrained output space . when fine-tuned to target constrained formal languages like SQL, these models often generate invalid code, rendering it unusable. |
| Approach: | They propose a method for constraining auto-regressive decoders of language models through incremental parsing. |
| Outcome: | The proposed method can find valid output sequences by rejecting inadmissible tokens . it can be used on Spider and CoSQL text-to-SQl translation tasks . |
Copied to clipboard
| Challenge: | Semantic sentence embeddings are usually supervisedly built minimizing distances between pairs of embeddable sentences labelled as semantically similar by annotators. |
| Approach: | They propose a language-independent approach to build large datasets of pairs of informal texts weakly similar, without manual human effort, exploiting Twitter’s powerful signals of relatedness: replies and quotes of tweets. |
| Outcome: | The proposed model learns classical Semantic Textual Similarity, and excels on tasks where pairs of sentences are not exact paraphrases. |
Copied to clipboard
| Challenge: | linguistic models have a higher correlation with human ground truth ratings than labeled data . word vectors have often been evaluated on standard word relatedness benchmarks . |
| Approach: | They propose to use unsupervised, supervised, and finally supervised methods to extract emotional associations from pretrained vectors and models. |
| Outcome: | The proposed method shows higher correlation with ground truth ratings than state-of-the-art lexicons based on labeled data. |
Copied to clipboard
| Challenge: | Existing studies show that word choice is driven by demographics within the United States. |
| Approach: | They develop computational methods to study word choice within a sociolinguistic lexical variable . they use two variables to test for attitudes towards sexuality and gender in the u.s. |
| Outcome: | The proposed methods allow us to examine attitudes towards sexuality and gender in the United States through two lexical variables. |
Copied to clipboard
| Challenge: | Moral sentiment is often motivated by its targets, which can correspond to individuals or collective entities. |
| Approach: | They propose a model to predict moral attitudes towards entities and moral foundations jointly using tweets written by US politicians. |
| Outcome: | The proposed model predicts moral attitudes towards entities and moral foundations jointly from tweets written by US politicians. |
Copied to clipboard
| Challenge: | Existing studies have found that presenting uncertainty in science communications influences people's perception of scientific findings and trust in science. |
| Approach: | They propose a model that models both the level and the aspects of certainty in scientific findings using an annotated dataset. |
| Outcome: | The proposed model can predict overall certainty and individual aspects of scientific findings with pre-trained language models, providing a more complete picture of the author’s intended communication. |
Copied to clipboard
| Challenge: | Various measures have been proposed to quantify human-like social biases in word embeddings, but they can suffer from measurement error. |
| Approach: | They propose to assess the reliability of word embedding gender bias measures by examining their reliability across different choices of random seeds, scoring rules and words. |
| Outcome: | The proposed measures can suffer from measurement error, and the results inform better design of word embedding gender bias measures. |
Copied to clipboard
| Challenge: | Existing methods for rumor detection are limited to the strict relation of user responses or oversimplify the conversation structure. |
| Approach: | They propose a method that reinforces interaction of user opinions while reducing negative impact imposed by irrelevant posts. |
| Outcome: | The proposed method improves performance on three Twitter datasets and can detect rumors at early stages. |
Copied to clipboard
| Challenge: | despite the importance of bill-to-bill linkages, existing approaches fail to address semantic similarities across bills. |
| Approach: | They propose a 5-class classification task that closely reflects the nature of the bill generation process. |
| Outcome: | The proposed method captures similarities across legal documents at various levels of aggregation. |
Copied to clipboard
| Challenge: | Using two additional wordsets, we compute the relative polarization of a topical wordsetting across two distributional representations. |
| Approach: | They propose a new measure to compute the relative polarization of a topical wordset across two distributional representations using two additional wordsetes deemed to have opposite valence to represent two different poles. |
| Outcome: | The proposed measure is validated by a case study and validated in a randomized controlled trial. |
Copied to clipboard
| Challenge: | Existing datasets for humour classification are limited due to the subjectivity of the content and the multiple interpretations of the data. |
| Approach: | They propose to annotate a multi-modal humour-annotated dataset using stand-up comedy clips and compute a humor quotient using the audience's laughter. |
| Outcome: | The proposed scoring mechanism is validated by comparing with manual scoring methods and achieves an accuracy of 0.813 in terms of QWK. |
Copied to clipboard
| Challenge: | Several studies have identified such linguistic classes of words that occur frequently in natural language text and are bias-inducing by virtue of their framing effects. |
| Approach: | They propose to use linguistic cues to induce subtle biases through implied sentiment and presupposed facts to influence the distribution of the generated text. |
| Outcome: | The proposed models are sensitive to these framing effects, but show that they lead to measurable style and topic differences in the generated text, leading to language that is, on average, more polarised and more skewed towards controversial entities and events. |
Copied to clipboard
| Challenge: | Sentence embedding is a set of effective and versatile techniques for converting raw text into numerical vector representations. |
| Approach: | They propose a generic and end-to-end approach to embed sentences from a partially labeled dataset using supervised methods. |
| Outcome: | The proposed approach achieves state-of-the-art results using only a small fraction of labeled sentence pairs on various benchmark tasks. |
Copied to clipboard
| Challenge: | Existing studies have used probing tasks to verify the presence of linguistic properties in vector representations, but it is unclear whether they can be manipulated to indirectly steer them. |
| Approach: | They validate a geometric mapping technique to transform linguistic properties without tuning . they use a pre-trained multilingual autoencoder to transform three linguistic property . |
| Outcome: | The proposed method can be used without tuning of the pre-trained autoencoder . the results are validated in monolingual and cross-lingual settings . |
Copied to clipboard
| Challenge: | Linguistic typology generally divides synthetic languages into groups based on their morphological fusion. |
| Approach: | They propose to quantify the degree of fusion of morphological features in a surface form . they recapitulate the usual linguistic classifications for concatenative systems . |
| Outcome: | The proposed measure recapitulates the usual classifications for concatenative systems and provides new measures for nonconcatenating ones. |
Copied to clipboard
| Challenge: | Existing studies have focused on agent-based simulations of language emergence. |
| Approach: | They propose to model the trade-off between word order and inflection in natural languages by using neural network agents. |
| Outcome: | The results show that neural agents strive to maintain the utterance type distribution observed during learning, rather than developing a more efficient or systematic language. |
Copied to clipboard
| Challenge: | Existing approaches to same side stance classification (S3C) require domain knowledge and semantic inference to solve the task. |
| Approach: | They propose to use same side stance classification to predict whether two arguments argue for the same stance for a given pair of arguments. |
| Outcome: | The proposed model fails to generalize both within and across topics and domains when adjusting the sampling strategy to a more adversarial scenario. |
Copied to clipboard
| Challenge: | Unlike most of the previous work focusing on the English language, this paper focuses on the Chinese ORL task. |
| Approach: | They propose to use a standard English MPQA dataset to construct a Chinese ORL dataset and investigate the effectiveness of cross-lingual transfer methods. |
| Outcome: | The proposed method is able to detect and improve the performance of the proposed method in Chinese. |
Copied to clipboard
| Challenge: | Current research in automatic summarisation is expensive to create, posing a challenge for any language. |
| Approach: | They propose to use a large-scale multilingual summarisation dataset with articles in 92 languages and more than 35 writing scripts to generate a multilingual dataset. |
| Outcome: | The proposed method is the largest, most inclusive, existing dataset and one of the largest and most inclusive datasets for any NLP task. |
Copied to clipboard
| Challenge: | Recent efforts to develop deep learning models for text generation tasks are challenging for non-experts. |
| Approach: | They propose methods to automatically create deep learning models for extractive and abstractive summarization tasks using large language models. |
| Outcome: | The proposed methods achieve near state-of-the-art performance on a range of datasets. |
Copied to clipboard
| Challenge: | Post-editing (PE) machine translation (MT) output can save time and reduce errors. |
| Approach: | They propose to use automatic word-level quality estimation to predict correctness of MT output to flag problematic output. |
| Outcome: | The proposed model is not good enough to support human translations, but is based on a visualization reflecting uncertainty of the model. |
Copied to clipboard
| Challenge: | Massively multilingual language models offer state-of-the-art cross-lingual transfer performance on a range of NLP tasks, but there is a profound performance gap between resource-rich and resource-poor target languages. |
| Approach: | They propose a series of data-efficient methods that enable quick and effective adaptation of pretrained multilingual models to low-resource languages and unseen scripts. |
| Outcome: | The proposed methods improve learning of the new dedicated embedding matrix in the target language and for low-resource languages written in unseen scripts. |
Copied to clipboard
| Challenge: | a recent study has shown that MT post-editing can reduce translation quality and speed . a large-scale study involving 30 professional translators examined the relationship between MT performance and post-edited outputs. |
| Approach: | They examine the relationship between MT performance and post-editing time and quality . they use neural MT of high quality to improve translation quality based on phrase-based MT . |
| Outcome: | The proposed model is not stable predictor of time or quality, the authors say . they find that better MT systems lead to fewer changes in the sentences . |
Copied to clipboard
| Challenge: | Recent advances in multilingual natural language processing have improved performance on benchmarks such as XTREME and XGLUE by 13 points . however, improvements have been easier to achieve in some tasks than others . |
| Approach: | They extend XTREME to XTRAME-R, which includes ten natural language understanding tasks and covers 50 typologically diverse languages. |
| Outcome: | The proposed framework improves the performance on the XTREME multilingual benchmark by 13 points compared to human-level performance on English transfer learning. |
Copied to clipboard
| Challenge: | Lexical disambiguation is a major challenge for machine translation systems . previous work focused on automatic post-hoc analysis of translations, but rules of what makes a disambiguations correct or incorrect tend to be imprecise. |
| Approach: | They propose a black-box method that uses contrastive conditioning to detect disambiguation errors. |
| Outcome: | The proposed method is scalable and reliable for disambiguation evaluations. |
Copied to clipboard
| Challenge: | Existing models for extractive rationales do not work as well on reasoning tasks requiring free-text rationale. |
| Approach: | They propose to use pipelines to extract rationales from input words and to use them to explain reasoning tasks. |
| Outcome: | The proposed models exhibit desirable properties for explaining commonsense question-answering and natural language inference, indicating their potential for producing faithful free-text rationales. |
Copied to clipboard
| Challenge: | Integrated Gradients (IG) is widely adopted due to its desirable explanation axioms and the ease of gradient computation. |
| Approach: | They propose an attribution-based explanation algorithm that uses averaging the model's output gradient interpolated along a straight-line path in the input data space. |
| Outcome: | The proposed method is compared with IG on multiple sentiment classification datasets. |
Copied to clipboard
| Challenge: | a new technique for exploring contextualized vector space is proposed . masked prediction of a word in a sentence allows controlled exploration of the space . |
| Approach: | They propose a method for exploring regions around individual points in a contextualized vector space . they use a static embedding to induce a "pseudoword" vector and masked prediction of a word . |
| Outcome: | The proposed method investigates the geometry of the contextualized space around individual instances of a word . it uses a static embedding to induce a contextualized "pseudoword" vector . |
Copied to clipboard
| Challenge: | Sequence models produce accurate predictions, but their decision making processes are hard to explain. |
| Approach: | They propose an efficient algorithm to approximate sequential objective by identifying the most faithful rationales. |
| Outcome: | The proposed algorithm is best at optimizing the sequential objective and provides the most faithful rationales. |
Copied to clipboard
| Challenge: | despite popularity of influence functions, their computational cost does not scale well with model and training data size. |
| Approach: | They propose a fast parallel variant that approximates the “influences” of training data-points for test predictions. |
| Outcome: | The proposed method achieves about 80X speedup while being highly correlated with the original influence values. |
Copied to clipboard
| Challenge: | Recent work on large language models has made this hypothesis popular . but, word order is not important enough to make sentence structure relevant . |
| Approach: | They propose an efficient procedure that finds word order having highest likelihood under a fixed language model. |
| Outcome: | The proposed procedure can be used to find the ordering of a bag of words having the highest likelihood under a fixed language model. |
Copied to clipboard
| Challenge: | Named entity recognition models require abundant high-quality annotations to train . distant supervision may induce incomplete and noisy labels, making supervised learning ineffective. |
| Approach: | They propose a noise-robust learning scheme for training named entity recognition models using only distantly-labeled data and a self-training method that uses contextualized augmentations created by pre-trained language models. |
| Outcome: | The proposed method outperforms existing supervised NER models on three datasets by significant margins. |
Copied to clipboard
| Challenge: | Existing approaches to solve this problem generate embeddings for noun and relation phrases . ambiguous subject-relation-object triples are created by open knowledge graphs . |
| Approach: | They propose a model to learn both embeddings and cluster assignments in an end-to-end approach . they propose CUVA to be able to group noun and relation phrases using embeddable features . |
| Outcome: | The proposed model outperforms state-of-the-art methods over multiple benchmarks. |
Copied to clipboard
| Challenge: | Existing knowledge graph embedding methods to learn representations of knowledge graphs are conceptually simple and can be applied to tasks like factoid question answering (Saxena et al., 2020) and reasoning. |
| Approach: | They propose a Hierarchical Transformer model to jointly learn Entity-relation composition and Relational contextualization based on a source entity’s neighborhood. |
| Outcome: | The proposed model achieves state-of-the-art on multiple link prediction datasets and can be integrated into BERT and demonstrate its effectiveness on two Freebase factoid question answering datasets. |
Copied to clipboard
| Challenge: | Existing methods to build named entity recognition systems with limited labeled data are lacking. |
| Approach: | They propose three orthogonal schemes to build named entity recognition systems when labeled data is limited. |
| Outcome: | The proposed NER systems outperform existing methods on few-shot and training-free settings. |
Copied to clipboard
| Challenge: | Existing approaches to generate named entity lexica for lower-resource languages are under performing. |
| Approach: | They propose a technique to automatically mine cross-lingual named-entity lexica from mined web data. |
| Outcome: | The proposed technique outperforms baselines at extracting cross-lingual entity pairs and mines 164 million entity pairs from 120 different languages aligned with English. |
Copied to clipboard
| Challenge: | Existing methods for event-event temporal relation extraction are sparse on event-time information. |
| Approach: | They propose a model for event-event temporal relation classification and an auxiliary task, relative event time prediction, which predicts the event time as real numbers. |
| Outcome: | The proposed model significantly improves the RoBERTa-based baseline and achieves state-of-the-art performance on MATRES dataset. |
Copied to clipboard
| Challenge: | State-of-the-art NLP models adopt shallow heuristics that limit their generalization capability. |
| Approach: | They propose to use heuristics that limit their generalization capability to model lexical overlap with the training set in Named-Entity Recognition and Event or Type heuristic in Relation Extraction to test their models. |
| Outcome: | The proposed model can perform better on the two key tasks, while the retention of training relation triples. |
Copied to clipboard
| Challenge: | metric BaryScore is used to evaluate text generation based on deep contextualized embeddings. |
| Approach: | They propose to model the layer output of deep contextualized embeddings as a probability distribution rather than a vector embeddable layer. |
| Outcome: | The proposed metric outperforms other BERT based metrics and exhibits more consistent behaviour in particular for text summarization. |
Copied to clipboard
| Challenge: | a new protocol allows for a multilingual hierarchy of concepts and images based on native speakers . the results suggest that the current models are not robust enough to handle multilingual data . |
| Approach: | They propose a protocol to construct an ImageNet-style hierarchy representative of more languages and cultures. |
| Outcome: | The proposed protocol lets the selection of concepts and images be entirely driven by native speakers, rather than scraping them automatically. |
Copied to clipboard
| Challenge: | Pre-trained language models have boosted performance on some WS benchmarks, but the source of improvement is not clear. |
| Approach: | They propose a method that uses twin sentences for evaluation and two new baselines that account for artifacts in WS benchmarks. |
| Outcome: | The proposed evaluation method is suboptimal for the Winograd Schema . it uses twin sentences to account for commonsense reasoning abilities . |
Copied to clipboard
| Challenge: | Entity disambiguation (ED) is the last step of entity linking when candidate entities are reranked according to the context they appear in. |
| Approach: | They propose a dataset that includes 16K short text snippets annotated with entity mentions to evaluate EL models. |
| Outcome: | The proposed dataset shows that the performance of EL systems is overestimated . the results show that the EL system performance is significantly better on the ShadowLink benchmark . |
Copied to clipboard
| Challenge: | XLM-R model outperforms other pre-trained models in annotated data. |
| Approach: | They adapt the data collection protocol for MNLI and collect 18K sentence pairs annotated by crowd workers and experts. |
| Outcome: | The proposed dataset outperforms other pre-trained models on the expert-annotated data. |
Copied to clipboard
| Challenge: | supervised learning is a key component of offensive language detection, but there is little attention given to the quality of annotated data. |
| Approach: | They propose to examine the level of agreement among annotators while selecting data to create offensive language datasets, a task involving a high level of subjectivity. |
| Outcome: | The proposed datasets show that annotators' agreement has a strong effect on classifiers performance and robustness. |
Copied to clipboard
| Challenge: | Graph-based dependency parsers can be improved without compromising on accuracy or accuracy. |
| Approach: | They propose two approaches to single-root dependency parsing that yield speed ups . they show that one approach is fully correct and finds the optimal dependency tree . |
| Outcome: | The proposed approach finds the optimal dependency tree without loss of accuracy or optimality. |
Copied to clipboard
| Challenge: | Spanning trees are a fundamental model of dependency structure in natural language processing, syntactic dependency trees. |
| Approach: | They propose to use a spanning tree sampling algorithm to faithfully sample dependency trees from a graph subject to a root constraint. |
| Outcome: | The proposed sampling algorithm can sample K trees without replacement in O(K N3 + K2 N) time. |
Copied to clipboard
| Challenge: | Existing discontinuous constituent parsers are slow and lack accuracy and speed . however, discontinuous parsing can be solved by any off-the-shelf continuous parser . |
| Approach: | They propose to reduce discontinuous constituent parsing to a continuous problem by reordering tokens. |
| Outcome: | The proposed method is on par with state-of-the-art methods but considerably faster. |
Copied to clipboard
| Challenge: | Existing neural CCG parsing methods do not support CCG derivations that violate the rule schemata. |
| Approach: | They propose a new representation for CCG derivations that decomposes CCG into several independent pieces. |
| Outcome: | The proposed representation decomposes CCG derivations into independent pieces . it prevents span-based models from violating the schemata . |
Copied to clipboard
| Challenge: | Existing methods for fine-tuning intermediate tasks are inefficient and expensive. |
| Approach: | They propose to use a set of 42 intermediate and 11 target English classification, multiple choice, question answering, and sequence tagging tasks to identify the best settings for intermediate transfer learning. |
| Outcome: | The proposed methods achieve an average Regret@3 of 1% across all target tasks. |
Copied to clipboard
| Challenge: | Existing Transformers that scale to long sequences are not compatible with relative position encoding. |
| Approach: | They propose a Performer-based model with relative position encoding that scales linearly on long sequences. |
| Outcome: | The proposed model outperforms performer on long sequences with no computational overhead and outperformed vanilla Transformer on most of the tasks. |
Copied to clipboard
| Challenge: | Pruning methods have proven to be effective at reducing model size, while distillation methods are proven for speeding up inference. |
| Approach: | They propose a block pruning approach that integrates structured pruning methods with the movement pruning paradigm for fine-tuning. |
| Outcome: | The proposed model is 2.4x faster, 74% smaller and faster than distilled models on classification and generation tasks. |
Copied to clipboard
| Challenge: | Efficient transformers outperform recurrent neural networks in natural language generation, but this comes with significant computational cost and memory footprint during generation. |
| Approach: | They propose to convert a pretrained transformer into its efficient recurrent counterpart, improving efficiency while maintaining accuracy. |
| Outcome: | The proposed transformers outperform recurrent neural networks in natural language generation but come with significant computational and memory footprint during generation. |
Copied to clipboard
| Challenge: | Large language models such as BERT are used in many NLP tasks, but their pretraining phase can be prohibitively expensive for startups and academic research groups. |
| Approach: | They propose a recipe for pretraining a large language model in 24 hours using a low-end deep learning server. |
| Outcome: | The proposed model can be trained on GLUE tasks at fraction of the cost of pretraining. |
Copied to clipboard
| Challenge: | Recent studies on compression of pretrained language models usually use preserved accuracy as the metric for evaluation. |
| Approach: | They propose two new metrics that measure how closely a compressed model mimics the original model. |
| Outcome: | The proposed metrics measure how closely a compressed model (i.e., student) mimics the original model (e.g., teacher). |
Copied to clipboard
| Challenge: | In IndoBERTweet, a pretraining model for Indonesian Twitter is extended with domain-specific vocabulary. |
| Approach: | They propose a pretraining model that extends a monolingual Indonesian BERT model with domain-specific vocabulary. |
| Outcome: | The proposed model can be initialized with the average BERT subword embedding five times faster than existing methods for vocabulary adaptation. |
Copied to clipboard
| Challenge: | ML models with handcrafted features are linguistically explainable, expandable, and competent against the modern neural models. |
| Approach: | They propose to combine traditional ML models with ML transformers to improve readability assessment by 99% accuracy. |
| Outcome: | The proposed model achieves state-of-the-art (SOTA) accuracy on popular datasets. |
Copied to clipboard
| Challenge: | Current NLP models produce unreliable or catastrophic predictions when training and test distributions differ . current models tend to produce unreliability or even catastrophic predictions that hurt user trust. |
| Approach: | They categorize examples as exhibiting a background shift or semantic shift and use calibration and density estimation methods to detect OOD examples. |
| Outcome: | The proposed methods beat calibration methods in background shift settings and perform worse in semantic shift settings. |
Copied to clipboard
| Challenge: | Recent work focused on training largescale and complex neural network models, but they are opaque in terms of their decision-making process. |
| Approach: | They propose a multi-task teacher-student framework for self-training pre-trained language models with limited task-specific labels and annotated rationales. |
| Outcome: | The proposed model improves performance in low-resource settings by making it aware of its rationalized predictions. |
Copied to clipboard
| Challenge: | In supervised and unsupervised learning, adding loss terms often leads to improved performance. |
| Approach: | They propose an algorithm that balances the gradient magnitude of loss terms across all layers . they use Adam to add loss terms to neural models, but add more terms as they are added . |
| Outcome: | The proposed method improves performance and improves training outcomes. |
Copied to clipboard
| Challenge: | Classification problems with thousands or more classes occur in NLP, for example language models or document classification. |
| Approach: | a new algorithm uses a binary tree with sparse hyperplanes and small softmax classifiers at the leaves to predict the top class. |
| Outcome: | The proposed model is faster at inference because the input follows a single path to a leaf and the softmax classifier operates on a small subset of the classes. |
Copied to clipboard
| Challenge: | Named entity recognition (NER) is a method of detecting entity spans and classifying them into predefined categories. |
| Approach: | They propose a method to iteratively perform noisy label refinery by using self-collaborative denoising learning. |
| Outcome: | The proposed learning paradigm exploits reliable labels and communicates with unreliable annotations by collaborative denoising. |
Copied to clipboard
| Challenge: | a recent study shows that drawing inferences between open domain predicates is a necessity for true language understanding. |
| Approach: | They propose to reinterpret the Distributional Inclusion Hypothesis to model entailment between predicates of different valencies. |
| Outcome: | The proposed graphs are more useful than using the same valency evidence, the authors show . they show that drawing on evidence across valencies answers more questions than using only the same evidence. |
Copied to clipboard
| Challenge: | Existing work on sentence ordering has focused on exploiting different categories of features like coreference clues. |
| Approach: | They propose a sentence ordering task as a conditional text-to-marker generation problem that leverages a pre-trained Transformer-based model to identify a coherent order for a given set of shuffled sentences. |
| Outcome: | The proposed model performs well across 7 datasets in Perfect Match Ratio and Kendall’s tau. |
Copied to clipboard
| Challenge: | State-of-the-art Ontology Alignment systems are based on domain-dependent approaches with handcrafted rules or domain-specific architectures, making them unscalable and inefficient. |
| Approach: | They propose a Deep Learning based model that exploits syntactic and semantic information encoded in ontologies by using a dual-attention mechanism. |
| Outcome: | The proposed model exploits syntactic and semantic information encoded in ontologies and is flexible and scalable to different domains with minimal effort. |
Copied to clipboard
| Challenge: | Recent research shows that automatic generation of synthetic utterance-program pairs can alleviate the first problem, but its potential for the second has thus far been under-explored. |
| Approach: | They propose to generate synthetic utterance-program pairs for improving compositional generalization in semantic parsing by using structurally-diverse examples. |
| Outcome: | The proposed approach leads to dramatic improvements in compositional generalization and moderate improvements in the traditional i.i.d setup. |
Copied to clipboard
| Challenge: | lexical substitution tasks require a system to provide adequate replacements for a word in a given context. |
| Approach: | They propose a generative approach to lexical substitution using a seq2seq model to generate suitable replacements for a word in context. |
| Outcome: | The proposed approach achieves state-of-the-art on different benchmarks and human evaluation of the generated substitutes. |
Copied to clipboard
| Challenge: | Recent studies have shown that news media exaggerate scientific papers by exagging their findings. |
| Approach: | They propose a method to detect when a news article has exaggerated a scientific finding . they use annotated press release/abstract pairs to compare machine learning models . |
| Outcome: | The proposed method outperforms PET and supervised learning on a multi-task version of Pattern Exploiting Training. |
Copied to clipboard
| Challenge: | Phrase representations derived from pretrained language models often lack lexical similarity to determine semantic relatedness. |
| Approach: | They propose a contrastive fine-tuning objective that enables BERT to produce more powerful phrase embeddings by fine- tuning a dataset of diverse phrasal paraphrases and a large-scale dataset of phrases in context. |
| Outcome: | The proposed model outperforms baseline models across phrase-level similarity tasks while also showing increased lexical diversity between nearest neighbors in the vector space. |
Copied to clipboard
| Challenge: | Existing work on semantic change detection methods has focused on generic research questions and datasets, using them as a training ground for proof-of-concept studies. |
| Approach: | They propose to use type-level embeddings to detect new semantic shifts and token-level embeddeds to isolate regionally specific occurrences. |
| Outcome: | The proposed method is comparable to state-of-the-art on diachrony tasks, but it does not translate to practical value in detecting new semantic shifts. |