Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
Copied to clipboard
| Challenge: | Existing tools for natural language resolution fail to handle ambiguous referents . ambiguity arises when the language is underspecified or there are multiple candidate referent. |
| Approach: | They investigate how pragmatic modulators outside of the linguistic content are critical for correct interpretation of referents in underspecified contexts. |
| Outcome: | The proposed method can be used to resolve referents in human environments. |
Copied to clipboard
| Challenge: | Existing approaches to train conditional languagemodels without supervised learning fail to scale to large action spaces, thus allowing to train a language agent by only interacting with its environment without any task-specific prior knowledge. |
| Approach: | They propose an original approach to train conditional languagemodels without supervised learning by only using reinforcement learning. |
| Outcome: | The proposed approach avoids the dependency to labelled datasets and reduces pretrained policy flaws such as language or exposure biases. |
Copied to clipboard
| Challenge: | Existing adaptive policies for simultaneous neural machine translation use monotonic attention to perform read/write decisions based on the partial source and target sequences. |
| Approach: | They propose a framework to aid monotonic attention with an external language model to improve its decisions. |
| Outcome: | The proposed approach improves on English-German and English-French translation tasks by using a language model. |
Copied to clipboard
| Challenge: | Existing research on automatic text summarization does not fully align with students’ needs. |
| Approach: | They propose a survey methodology that can be used to investigate the needs of users of automatically generated summaries. |
| Outcome: | The proposed method can be easily adjusted to investigate different user groups. |
Copied to clipboard
| Challenge: | Currently available grammatical error correction datasets focus on written essays . a novel dataset is presented to improve the accuracy of existing educational chatbots . |
| Approach: | They propose a novel grammatical error correction dataset using essays and other long-form text written by language learners. |
| Outcome: | The proposed dataset improves the performance of a conversational chatbot in a human-machine conversational setting. |
Copied to clipboard
| Challenge: | Existing methods to measure diversity of chitchat model responses have been proposed to measure iteratively. |
| Approach: | They propose a metric which uses Natural Language Inference to measure the semantic diversity of a set of model responses for a conversation. |
| Outcome: | The proposed metric improves the diversity of a sampled set of responses using a new generation procedure called Diversity Threshold Generation. |
Copied to clipboard
| Challenge: | Existing few-shot text classification methods often lack labeled data in real-world tasks. |
| Approach: | They propose a meta-learning method that encodes how to attend for given tasks . they evaluate the method on five benchmark datasets and show it is competitive . |
| Outcome: | The proposed method performs better on five benchmark datasets than previous methods on labeled data. |
Copied to clipboard
| Challenge: | Existing works of knowledge infusion depend on multi-task learning frameworks, which are inefficient and require large-scale retraining when new knowledge is considered. |
| Approach: | They propose a method which integrates knowledge-generated attention maps into the self-attention mechanism and integrates it into the model. |
| Outcome: | The proposed model outperforms existing methods on academic datasets and industry-scale ad relevance applications. |
Copied to clipboard
| Challenge: | Recent advances in machine learning have led to the use of contrastive loss for representation learning. |
| Approach: | They propose to use batch-softmax contrastive loss to train pairwise sentence embeddings . they propose to take a batch-softermax contrastitive loss and train it with different loss functions . |
| Outcome: | The proposed model improves on a number of datasets and pairwise sentence scoring tasks. |
Copied to clipboard
| Challenge: | a large dataset of news article revision histories provides clues to narrative and factual evolution in news articles. |
| Approach: | They propose tasks to predict edit-actions performed during version updates . they define article-level edit actions: Addition, Deletion, Edit and Refactor . |
| Outcome: | The proposed dataset is large-scale and multilingual and spans 15 years . it shows that edit-actions are predictable and are likely to be based on factual evolution . |
Copied to clipboard
| Challenge: | Using neural networks, we can model the impact of speaker role on language use through the game of Mafia. |
| Approach: | They analyze the effect of speaker role on language use through the game of Mafia, in which players are assigned either an honest or a deceptive role. |
| Outcome: | The proposed model outperforms a standard BERT-based text classification approach on two auxiliary tasks and identifies features that distinguish between player roles. |
Copied to clipboard
| Challenge: | Semantic parsing models fail at compositional generalization due to lack of reasoning ability. |
| Approach: | They propose to use subtree substitution for compositional data augmentation to increase the number of subtreas with similar semantic functions as exchangeable. |
| Outcome: | The proposed method improves performance on Scan and GeoQuery, and new SOTA on compositional split of GeoQuery. |
Copied to clipboard
| Challenge: | Labelled data is the foundation of most natural language processing tasks, but there are valid beliefs about what the correct data labels should be. |
| Approach: | They propose two contrasting paradigms for data annotation that encourage annotator subjectivity . they propose a descriptive paradigm that allows for the surveying and modelling of different beliefs . |
| Outcome: | The proposed paradigms encourage annotator subjectivity, while the prescriptive paradigm discourages it. |
Copied to clipboard
| Challenge: | DL-based graders often lack the ability to explain and justify how a prediction is made, which decreases their trustworthiness and hinders educators from embracing them in practice. |
| Approach: | They conducted a user study to determine whether DL-based graders align with human grader . they also ran a randomized controlled experiment to explore the impact of highlighting important words detected by DL grader. |
| Outcome: | The proposed method enables human graders to identify important words when marking short answer questions. |
Copied to clipboard
| Challenge: | Existing knowledge-grounded dialogue systems perform poorly on unseen topics due to limited topics covered in training data. |
| Approach: | They propose a language model that homogenizes different knowledge sources to a unified knowledge representation for knowledge-grounded dialogue generation tasks. |
| Outcome: | The proposed language model generalizes well across knowledge-grounded dialogue tasks. |
Copied to clipboard
| Challenge: | Despite advances in self-supervised learning, there is a lack of models that can effectively capture both intra- and intra-item semantics for semi-structured session data. |
| Approach: | They propose a graph-based transformer model for semi-structured session data that captures both intra- and intra-item semantics. |
| Outcome: | The proposed model outperforms baselines in three session search and entity linking tasks by up to 9%. |
Copied to clipboard
| Challenge: | Recent research has made great strides towards understanding the ideological bias (i.e., stance) of news media along the left-right spectrum. |
| Approach: | They propose a novel approach for the study of ideology based on its left or right positions on the issue being discussed. |
| Outcome: | The proposed method allows for the quantitative and temporal measurement and analysis of polarization as a multidimensional ideological distance. |
Copied to clipboard
| Challenge: | Pretrained language models provide high-quality contextualized word embeddings, but training question answering models requires large amounts of annotated data for specific domains. |
| Approach: | They propose a framework for automatically generating more non-trivial question-answer pairs to improve model performance. |
| Outcome: | The proposed framework outperforms state-of-the-art (SOTA) pretrained language models and transfer learning approaches on standard question-answering benchmarks. |
Copied to clipboard
| Challenge: | Existing methods for interpreting the underlying dynamics of Transformers have been criticized for their lack of reliability. |
| Approach: | They propose a token attribution analysis method that incorporates all components in the encoder block and aggregates this across layers. |
| Outcome: | The proposed method significantly outperforms existing methods on saliency scores and correlation with gradient-based salience scores. |
Copied to clipboard
| Challenge: | Aspect sentiment triplet extraction (ASTE) is a challenging subtask in aspect-based sentiment analysis. |
| Approach: | They propose a bidirectional machine reading comprehension method to extract triplets of aspects, opinions and sentiments with complex correspondence from the context. |
| Outcome: | The proposed method achieves state-of-the-art on multiple benchmark datasets. |
Copied to clipboard
| Challenge: | Existing topic models adopt a fully unsupervised setting and their discovered topics may not reflect user preferences well due to their unsupervised nature. |
| Approach: | They propose a framework that allows out-of-vocabulary seeds to be used to find latent topics from text corpora. |
| Outcome: | The proposed framework can find topics that are never seen in the corpus and can benefit from the general knowledge of pre-trained language models. |
Copied to clipboard
| Challenge: | NLP-powered automatic question generation (QG) techniques have not been widely adopted in classrooms to date. |
| Approach: | They propose to identify key impediments and improve the usability of NLP-powered automatic question generation techniques by understanding how instructors construct questions and identifying touch points to enhance the underlying NLP models. |
| Outcome: | The proposed methods can be used by 11 instructors across 7 universities and highlight their needs and needs when creating questions. |
Copied to clipboard
| Challenge: | Social media and Internet forums are valuable sources of citizens’ opinions, which can be analyzed for community development and user behavior analysis. |
| Approach: | They present a pre-training and annotated datasets of Swahili and an emotion classification datasets that are manually annotating by two native Swahils. |
| Outcome: | The proposed model outperforms existing monolingual language model in almost all downstream tasks. |
Copied to clipboard
| Challenge: | Evaluating natural language generation systems is difficult, as there are many ways to express similar things in text. |
| Approach: | They combine interviews with NLG practitioners to examine ethical considerations and their implications for NLG evaluation. |
| Outcome: | The findings of the study surface goals, community practices, assumptions, and constraints that shape NLG evaluations, and examine their implications and how they embody ethical considerations. |
Copied to clipboard
| Challenge: | Existing methods for extractive and abstract summarization are limited to short abstracts . however, extended summaries provide detailed information beyond coarse information . |
| Approach: | They propose an extractive summarization tool that utilizes the introductory information of documents as pointers to their salient information. |
| Outcome: | The proposed extractive summarization improves on existing datasets with human-written summaries . the proposed summarizing improves in terms of cohesion and completeness compared to baselines and state-of-the-art . |
Copied to clipboard
| Challenge: | a method to control affective prosody of text-to-speech systems is proposed to use phoneme-level intermediate features as levers . DS is used to disentangle features relating to affective proody from those due to acoustics conditions and speaker identity . |
| Approach: | They propose a method to control the emotional prosody of Text to Speech systems by using phoneme-level intermediate features as levers. |
| Outcome: | The proposed method improves over the prior art in emulating emotion in speech . it adds the much-coveted "human touch" in machine dialogue, the authors say . |
Copied to clipboard
| Challenge: | Recent research shows that different forms of natural language-based interaction prove suitable to support users in accomplishing various visualization tasks. |
| Approach: | They propose a taxonomy of visualization tasks and a classification system to illustrate the state-of-the-art of natural language-based interaction in visualization. |
| Outcome: | The proposed model can support annotations, recommendations, explanations, and documentation tasks. |
Copied to clipboard
| Challenge: | Temporal reading comprehension (TRC) is a natural way to study temporal relations since natural language questions are flexible to capture divergent temporal relationships. |
| Approach: | They propose a reading comprehension approach that uses precise question understanding . they embed a temporal ordering question into two vectors and evaluate the temporal relation based on that . |
| Outcome: | The proposed approach outperforms strong baselines and achieves state-of-the-art performance on the TORQUE dataset. |
Copied to clipboard
| Challenge: | Existing studies on how NLP systems could be used in clinical practice focus on technical difficulties and usability challenges involved in implementing them. |
| Approach: | They propose to use Speech Recognition to transcribe the audio of a medical consultation and then to train sequence-to-sequence models to summarise the transcript into a consultation note. |
| Outcome: | The proposed system generates notes in real time during a doctor-patient consultation and is able to capture the salient points of a consultation . the proposed system is based on three rounds of user studies in a live telehealth clinic and identifies a number of clinical use cases that could prove challenging for the system. |
Copied to clipboard
| Challenge: | Cross-lingual question answering systems are becoming more and more important . a new approach can be generalized to more than 20 languages and outperforms previous models by 12% . |
| Approach: | They propose a cross-lingual question answering system that can be generalized to more than 20 languages . their approach can outperform previous models by 12% on multiple languages based on a dataset . |
| Outcome: | The proposed approach outperforms the previous models on multiple languages by 12% . it can be generalized to more than 20 languages and outperformed all previous models by 2% . |
Copied to clipboard
| Challenge: | Existing approaches to generate generic responses are ignoring low-frequency but generic responses and bringing low- frequency but meaningless responses. |
| Approach: | They propose a negative training paradigm that reminds dialogue models not to generate high-frequency responses during training. |
| Outcome: | The proposed method outperforms previous methods in the generic response problem while minimizing low-frequency but meaningless responses. |
Copied to clipboard
| Challenge: | Existing studies on back translation (BT) focus on beam search or random sampling . a new method to generate synthetic data with a backward model is proposed to improve BT performance. |
| Approach: | They propose a method to generate synthetic data to trade off quality and importance factors . back translation (BT) is one of the most significant technologies in NMT research fields . |
| Outcome: | The proposed method outperforms the baseline methods on WMT14 DE-EN, EN-DE, and RU-EN benchmark tasks. |
Copied to clipboard
| Challenge: | Automated text summarization systems involve humans for preparing data or evaluating model performance, yet, there is no systematic understanding of human-AI interactions and how to design for them. |
| Approach: | They conducted a systematic literature review of 70 papers and designed prototypes for each interaction. |
| Outcome: | The proposed design considerations were based on the results of a systematic literature review of 70 papers and interviews with 16 users. |
Copied to clipboard
| Challenge: | Recent studies show that auto-encoders perform language generation, smooth sentence interpolation, and style transfer over unseen attributes using unlabelled datasets in a zero-shot manner. |
| Approach: | They propose a discrete token-based perturbation approach to map "similar" sentences close by in latent space. |
| Outcome: | The proposed model can generate and perform language generation, style transfer and sentence interpolation tasks on unlabelled datasets in a zero-shot manner. |
Copied to clipboard
| Challenge: | Automated summarization methods are efficient but can suffer from low quality. |
| Approach: | They conducted an experiment with 72 participants to compare post-editing provided summaries with manual summarization for summary quality, human efficiency, and user experience. |
| Outcome: | The results show that post-editing improves summary quality, human efficiency, and user experience on formal (XSum news) and informal (Reddit posts) text. |
Copied to clipboard
| Challenge: | Despite recent advances in machine translation, a tremendous amount of translated content in the world is still written by humans. |
| Approach: | They propose a task of translation error correction (TEC) that corrects human-generated translations by correcting all errors in a source sentence and a human-created translation. |
| Outcome: | The proposed system improves translation accuracy by 5.1 points compared to MT systems with human errors . |
Copied to clipboard
| Challenge: | SpanBERT model is more robust than RoBERTa, despite having similar accuracy on unperturbed test data. |
| Approach: | They propose a pipeline to replace entity names with names from a variety of sources. |
| Outcome: | The proposed model performs worse when entities are renamed, the authors show . SpanBERT, which is pretrained with span-level masking, is more robust than RoBERTa . |
Copied to clipboard
| Challenge: | In the context of data labeling, researchers are interested in having humans select rationales . |
| Approach: | They conducted an online user study to understand how humans select rationales . they found that participants were near unanimous in their data labels . |
| Outcome: | The results show that participants selected 12% of input tokens as rationales, but fewer if unable to drag over multiple tokens at once. |
Copied to clipboard
| Challenge: | Recent studies show that fine-tuning pre-trained language models with a small set of labeled utterances in a supervised manner is helpful, but it yields an anisotropic feature space, which may suppress the expressive power of the semantic representations. |
| Approach: | They propose to regularize supervised pre-training towards isotropy by contrastive learning and correlation matrix regularizers. |
| Outcome: | The proposed methods improve supervised pre-training by regularizing the feature space towards isotropy. |
Copied to clipboard
| Challenge: | Existing methods for misinformation detection are limited to judging each document in isolation. |
| Approach: | They propose a task of cross-document misinformation detection that detects fake news from a cluster of topically related news documents. |
| Outcome: | The proposed method outperforms existing methods by up to 7 F1 points on this task. |
Copied to clipboard
| Challenge: | a new method for compositional action recognition is proposed to address the problem of zero-shot learning. |
| Approach: | They propose a method to generalize compositional action recognition models to new verbs and nouns . they use knowledge graphs to extract disentangled feature representations for verbs, noun and type constraint . |
| Outcome: | The proposed approach improves generalization ability of the compositional action recognition model to novel verbs and nouns that are unseen during training time. |
Copied to clipboard
| Challenge: | Prior work has shown that providing users with a machine-written draft or sentence-level continuations has limited success since the generated text tends to deviate from users’ intention. |
| Approach: | They propose to train a rewriting model that modifies specified spans of text within the user’s original draft to introduce descriptive and figurative elements in the text. |
| Outcome: | The proposed model is rated more helpful by users than a baseline infilling language model on a user study through Amazon Mechanical Turk. |
Copied to clipboard
| Challenge: | Existing models are vulnerable to adversarial attacks, but their vulnerability is underexplored. |
| Approach: | They propose to concatenate a perturbed but semantically similar tweet into a model that fools stock prediction models. |
| Outcome: | The proposed method achieves consistent success rates and causes significant monetary loss in trading simulation by simply concatenating a perturbed but semantically similar tweet. |
Copied to clipboard
| Challenge: | Multilingual Neural Machine Translation (MNMT) systems are often limited to many-to-one directions and suffer from poor performance in one-to one directions. |
| Approach: | They propose to build multilingual machine translation systems that serve arbitrary X-Y directions while leveraging multilinguality with a two-stage training strategy of pretraining and finetuning. |
| Outcome: | The proposed system outperforms baseline bilingual models and pivot translation models in most directions without the need for architecture change or extra data collection. |
Copied to clipboard
| Challenge: | Variational Autoencoder (VAE) is an effective framework to model the interdependency for non-autoregressive neural machine translation (NAT). |
| Approach: | They propose to use Variational Autoencoder to model interdependency for non-autoregressive neural machine translation (NAT) a posterior consistency regularization approach is proposed to improve translation quality . |
| Outcome: | The proposed model is 1.5/0.7 and 0.8/0.3 BLEU points faster than the baseline model. |
Copied to clipboard
| Challenge: | Existing systems that embed and amplify gender bias can still exhibit and exacerbate this problem. |
| Approach: | They propose a multi-step system that combines the positive aspects of rule-based and neural rewriting models to provide personalized outputs based on the users’ grammatical gender preferences. |
| Outcome: | The proposed system achieves 88.42 M2 F0.5 on a blind test set and improves over previous work on the first-person-only version of this task by 3.05 absolute increase in M2F0.5. |
Copied to clipboard
| Challenge: | Large language models are capable of generating fluent-appearing text with little task-specific supervision. |
| Approach: | They propose a pipeline that combines GPT-3 with a supervised filter that incorporates binary acceptability judgments from humans in the loop. |
| Outcome: | The proposed model can generate freetext explanations in a fewshot setting with human-written examples. |
Copied to clipboard
| Challenge: | Existing studies only explore entity representations, but propose a novel triple perspective for relation extraction. |
| Approach: | They propose to explicitly introduce relation representation and jointly represent it with entities to identify valid triples. |
| Outcome: | The proposed method is based on ablations and document-level relation extraction and joint entity and relation extraction. |
Copied to clipboard
| Challenge: | Meta-learning is an emerging field in machine learning, but there is no systematic survey of these approaches in NLP. |
| Approach: | They propose to introduce meta-learning and the common approaches and summarize their work and review their work in the NLP community. |
| Outcome: | The proposed methods improve performance in many NLP tasks but are limited to domains, languages, countries, or styles. |
Copied to clipboard
| Challenge: | despite its importance, little attention has been paid to improving the robustness of multimodal models. |
| Approach: | They propose simple diagnostic checks for modality robustness in a trained multimodal model . they find MSA models highly sensitive to a single modality, which creates issues . |
| Outcome: | The proposed checks show that models are highly sensitive to a single modality, which creates issues in their robustness. |
Copied to clipboard
| Challenge: | Variational Auto-Encoders are often used for text generation tasks due to the sequential nature of the text. |
| Approach: | They propose a variational Transformer framework that learns a series of layer-wise latent variables with each inferred from those of lower layers and tightly coupled with the hidden states by low-rank tensor product. |
| Outcome: | The proposed framework can learn latent variables from lower layers and incorporate more information. |
Copied to clipboard
| Challenge: | Existing approaches to mitigate demographic biases evaluate on monolingual data, however, multilingual data has not been examined. |
| Approach: | They propose a standard domain adaptation model to reduce gender bias in multilingual contexts. |
| Outcome: | The proposed model reduces gender bias and improves on two text classification tasks with three fair-aware baselines. |
Copied to clipboard
| Challenge: | Named entity recognition (NER) tasks require large labeled datasets to perform . compared to prior work, relative improvements in F1 of up to 16% are found . |
| Approach: | They propose to use self-training, knowledge distillation, and transfer learning to learn SLU models . they compare pipeline and pipeline approaches to find out how to use external data . |
| Outcome: | The proposed models improve performance beyond pre-trained models in resource-constrained settings . the best baseline model is a pipeline approach, while the best performance is achieved by an E2E model. |
Copied to clipboard
| Challenge: | Current approaches for controlling dialogue response generation focus on high-level attributes like style, sentiment, or topic. |
| Approach: | They propose a method that allows for more fine-grained control of dialogue response generation . they propose utterances that encourage the generation of control words in the future . |
| Outcome: | The proposed method outperforms state-of-the-art constrained generation baselines on task-oriented dialogue datasets and shows that it is more fine-grained than previous methods. |
Copied to clipboard
| Challenge: | Dialogue Sentence Embedding (DSE) is a self-supervised contrastive learning method that learns effective dialogue representations suitable for a wide range of dialogue-oriented tasks. |
| Approach: | They propose a self-supervised contrastive learning method that learns dialogue representations suitable for a wide range of dialogue tasks. |
| Outcome: | The proposed method outperforms baselines on five dialogue tasks on a few-shot and zero-shot datasets. |
Copied to clipboard
| Challenge: | a recent study examines the morality of NLP models that can take in arbitrary text and output a moral judgment . a Delphi project is a popular system for moral prediction, but it has received criticism . |
| Approach: | They propose to critique NLP methods for automating ethical decision-making . they examine a nascent task of predicting moral and ethical decisions from text . |
| Outcome: | The proposed model is unsafe at any accuracy, the authors argue . they argue that the proposed model could be useful in NLP, but not in AI. |
Copied to clipboard
| Challenge: | Existing paradigms for text generation are left-to-right decoding from autoregressive language models. |
| Approach: | They propose a decoding algorithm that incorporates heuristic estimates of future cost that are efficient for large-scale language models. |
| Outcome: | The proposed method outperforms baselines on five generation tasks and achieves new state-of-the-art performance on table-to-text generation, constrained machine translation, and keyword-constrained generation. |
Copied to clipboard
| Challenge: | Existing methods for multilingual sequence-to-sequence pretraining rely on monolingual corpora and do not use strong cross-lingual signal contained in parallel data. |
| Approach: | They propose a method that replaces monolingual words with a bilingual dictionary and predicts the reference translation according to a parallel corpus instead of recovering the original sequence. |
| Outcome: | The proposed method improves machine translation and cross-lingual natural language inference by 2.0 BLEU points and 6.7 accuracy points over existing methods at a fraction of their computational cost. |
Copied to clipboard
| Challenge: | Existing work on toxic speech classification relies on generic and repetitive explanations . elucidating toxic speech can help with downstream tasks such as debiasing . |
| Approach: | They propose a knowledge-informed encoder-decoder framework to generate toxic text explanations . they use multiple knowledge sources to generate detailed explanations of toxic text . |
| Outcome: | The proposed model outperforms state-of-the-art models significantly in generating toxic explanations . the proposed model can generate detailed explanations of toxic speech compared to baselines compared with baseline models . |
Copied to clipboard
| Challenge: | a recent study shows that current NLP models operate non-incrementally, causing unacceptable delays for the user. |
| Approach: | They propose a streaming BERT-based sequence tagging model that detects disfluencies in real-time . they train the model to decide whether to immediately output a prediction or wait for further context . |
| Outcome: | The proposed model produces accurate predictions sooner than baselines, with lower flicker . disfluencies hurt readability of ASR transcripts, erode model performance on downstream tasks . |
Copied to clipboard
| Challenge: | Content-based collaborative filtering (CF) predicts user-item interactions based on both items’ interaction history and item content information. |
| Approach: | They propose to combine item encodings with a multi-modality approach to improve training efficiency by 146x . |
| Outcome: | The proposed model improves training efficiency (up to 146x) on five datasets from two task domains of Knowledge Tracing and News Recommendation. |
Copied to clipboard
| Challenge: | Existing studies have focused on general response generation with neural network-based approaches, but none have addressed specific types of repetitions. |
| Approach: | They propose a weighted label smoothing method for explicitly learning which words to repeat during fine-tuning and a repetition scoring method that can output more appropriate repetitions during decoding. |
| Outcome: | The proposed method outperforms baselines in automatic and human evaluations on a pre-trained language model for generating repetitions. |
Copied to clipboard
| Challenge: | Existing text-based speech-to-speech translation systems rely on cascaded approach . text-to text translation systems require text generation and a single input to generate output . |
| Approach: | They propose a textless speech-to-speech translation system that can translate speech from one language into another without the need of text data. |
| Outcome: | The proposed system can translate speech from one language into another without text data. |
Copied to clipboard
| Challenge: | Existing studies on weak supervision for NLU focus on a specific task or simulate weak supervision signals from ground-truth labels. |
| Approach: | They propose a benchmark to advocate and facilitate research on weak supervision for NLU . they use document-level and token-level prediction tasks as examples . |
| Outcome: | The proposed benchmark advocates and facilitates research on weak supervision for NLU tasks. |
Copied to clipboard
| Challenge: | Despite advances in open information extraction, many systems focus on covering more information over compactness of constituents. |
| Approach: | They propose a neural OpenIE system that produces compact extractions with overlapping constituents by using a pipelined approach. |
| Outcome: | The proposed system produces 1.5x-2x more compact extractions than previous systems, with high precision, establishing a new state-of-the-art in OpenIE. |
Copied to clipboard
| Challenge: | a new dataset evaluates the ability of AI systems to reason about scene change imagination . a large human-model performance gap exists in the dataset . |
| Approach: | They propose a dataset to evaluate AI's ability to reason about scene change imagination . they use an image and a commonsense question to imagine a counterfactual scene change . |
| Outcome: | The proposed dataset evaluates the ability of AI systems to reason about scene change imagination. |
Copied to clipboard
| Challenge: | Pre-trained models are the state of the art in linguistics. |
| Approach: | They compare the performance of pre-trained and native English language models on the task of article prediction set up as a three way choice (a/an, the, zero) they argue that BERT captures a high level generalisation of article use akin to human intuition. |
| Outcome: | The proposed model outperforms humans on the linguistically interesting task of article prediction. |
Copied to clipboard
| Challenge: | a table-based question answering system requires complex reasoning and alignment between questions and tables. |
| Approach: | They propose a table-based QA model that consumes both natural and synthetic data . they combine retrieval with masking to pair natural sentences with QA . |
| Outcome: | The proposed model outperforms existing models in few-shot and full settings and on WikiTableQuestions. |
Copied to clipboard
| Challenge: | Existing methods to train language models without memorizing sensitive data are mismatched and can be difficult to screen and filter. |
| Approach: | They propose a method to train language generation models while protecting the confidential segments of training data. |
| Outcome: | The proposed method prevents unintended memorization by randomizing parts of the training process while protecting strong confidentiality. |
Copied to clipboard
| Challenge: | Existing methods for knowledge retrieval and answer prediction have left open questions about the quality and relevance of the retrieved knowledge and how the reasoning processes over implicit and explicit knowledge should be integrated. |
| Approach: | They propose a Knowledge Augmented Transformer which integrates both implicit and explicit knowledge in an encoder-decoder architecture while simultaneously reasoning over both knowledge sources during answer generation. |
| Outcome: | The proposed model achieves a strong state-of-the-art (+6% absolute) on the open-domain multimodal task of OK-VQA. |
Copied to clipboard
| Challenge: | Existing theories on how humans track discourse entities are based on the idea that humans maintain explicit memory representations for each entity that encode all properties of an entity and its relation to other entities. |
| Approach: | They adapt the psycholinguistic assessment of language models paradigm to higher-level linguistic phenomena and introduce an English evaluation suite that targets the knowledge of the interactions between sentential operators and indefinite NPs. |
| Outcome: | The evaluation suite targets the knowledge of the interactions between sentential operators and indefinite NPs and the models are challenged by multiple NP's and their behavior is not systematic. |
Copied to clipboard
| Challenge: | Recent research suggests that data order can have a significant impact on the performance of finetuned models for natural language understanding. |
| Approach: | They use paced curriculum learning to rank data and sample training mini-batches with increasing levels of difficulty during finetuning. |
| Outcome: | The proposed model improves performance for socialIQA, CosmosQA, CODAH, HellaSwag, WinoGrande in both tuning settings. |
Copied to clipboard
| Challenge: | Document dependency graphs (TDGs) are used to understand the temporal relations between events mentioned in a document and to improve downstream tasks such as timeline creation and time-aware summarization. |
| Approach: | They propose a temporal dependency graph parser that takes input from a text document and produces a graph that incorporates longer range dependencies. |
| Outcome: | The proposed framework outperforms existing models on three datasets and improves tasks such as timeline creation, time-aware summarization, and temporal information extraction. |
Copied to clipboard
| Challenge: | Abstractive summarization models suffer from the problem of hallucinations, where a summary contains facts or entities not present in the original document. |
| Approach: | They propose an abstractive summarization model that addresses the problem of factuality during pre-training and fine-tuning. |
| Outcome: | Experiments on three downstream tasks show that FactPEGASUS significantly improves factuality compared to the original pre-training objective in zero-shot and few-shot settings. |
Copied to clipboard
| Challenge: | Suicidal behaviors, including suicide attempts (SA) and suicide ideations (SI), are leading risk factors for death by suicide. |
| Approach: | They first built a Suicide Attempt and Ideation Events (ScAN) dataset, a subset of the publicly available MIMIC III dataset spanning over 12k+ EHR notes with 19k+ annotated SA and SI events information. |
| Outcome: | The proposed model is based on the publicly available MIMIC III Suicide Attempt and Ideation Events Retriever (ScANER) dataset and achieves a macro-weighted F1 score of 0.83 for identifying suicidal behavioral evidences and a micro-weighting score of 0.8 and 0.60 for classification of SA and SI for the patient’s hospital-stay. |
Copied to clipboard
| Challenge: | Language representations are an efficient tool used across NLP, but they are strife with encoded societal biases. |
| Approach: | They investigate the encoded biases in Hindi language representations based on cultural and historical contexts . they emphasize the necessity of social-awareness along with linguistic and grammatical artefacts when modeling language representation . |
| Outcome: | The proposed model reflects the cultural and cultural diversity of the region in which it is used . the model is based on the language and culture of the language being used based upon the study . |
Copied to clipboard
| Challenge: | Existing methods for generating homographic puns are heavy-weighted due to the lack of training data. |
| Approach: | They propose a way to generate pun sentences that does not require training on existing puns. |
| Outcome: | The proposed method outperforms baseline models and state-of-the-art models by a large margin. |
Copied to clipboard
| Challenge: | Existing empathetic dialogue models lack emotion-dependent response generation . elaine mccartney: "i'm sorry to hear that! " |
| Approach: | They propose a model to generate empathetic responses with human-consistent intents . they aim to address the bias of the empathic intent distribution between epd models and humans . |
| Outcome: | The proposed model outperforms state-of-the-art models in terms of empathy, relevance, and diversity on automatic and human evaluation. |
Copied to clipboard
| Challenge: | Existing datasets for Yes/No QA are lacking information needed to answer a Yes/Non question. |
| Approach: | They extend the Yes/No QA task by adding questions with an IDK answer to a BoolQ dataset and create out-of-domain test sets for the task. |
| Outcome: | The proposed dataset includes paragraphs together with naturally occurring questions whose answer is either "Yes" or "No". |
Copied to clipboard
| Challenge: | Abstract Meaning Representation parsers rely on node-to-word alignments, but lack the complexity of the pipeline. |
| Approach: | They propose a neural aligner for abstract meaning representation that learns node-to-word alignments without relying on pipelines. |
| Outcome: | The proposed approach improves accuracy and generalization from AMR2.0 to AMR3.0 corpora. |
Copied to clipboard
| Challenge: | Recent Part-Of-Speech (POS) induction models assume certain independence assumptions that do not hold in real languages. |
| Approach: | They propose a Masked Part-of-Speech Model (MPoSM) that can model arbitrary tag dependency and perform POS induction through the objective of masked POS reconstruction. |
| Outcome: | The proposed model can model arbitrary tag dependency and perform POS induction through the objective of masked POS reconstruction. |
Copied to clipboard
| Challenge: | Cognitive science has long promoted the formation of mental models as central to understanding and question-answering. |
| Approach: | They train a new model, DREAM, to answer questions that elaborate the scenes that situated questions are about and then provide those elaborations as additional context to a question-answering (QA) model. |
| Outcome: | The proposed model is able to create better scene elaborations than a representative state-of-the-art, zero-shot model. |
Copied to clipboard
| Challenge: | Pre-trained language models (PTLMs) have been shown to perform well on natural language tasks. |
| Approach: | They propose a commonsense contextualizer conditioned on sentences as input to make it generically usable in tasks involving natural language text. |
| Outcome: | The proposed model improves on existing methods on CSQA, ARC, QASC and OBQA datasets. |
Copied to clipboard
| Challenge: | Pre-trained language models have increased the performance of data-driven natural language processing (NLP) models on a wide variety of tasks. |
| Approach: | They propose a model-free approach to probing via prompting which formulates probing as a prompting task and combine pruning to analyze where the model stores the linguistic information in its architecture. |
| Outcome: | The proposed approach extracts information from pre-trained models while learning much less on its own. |
Copied to clipboard
| Challenge: | Task-oriented dialog systems can't handle multiplesearch results when querying a database due to the lack of such scenarios in existing datasets. |
| Approach: | They propose a task that focuses on disambiguating database search results by synthetically generating turns through a pre-defined grammar and collecting human paraphrases for a subset. |
| Outcome: | The proposed task improves performance on DSR-disambiguation even in the absence of in-domain data, suggesting it can be learned as a universal dialog skill. |
Copied to clipboard
| Challenge: | Defining task-specific schemas is the first step of building a task-oriented dialog system. |
| Approach: | They propose an unsupervised approach for slot schema induction from unlabeled dialog corpora using in-domain language models and unsupervised parsing structures. |
| Outcome: | The proposed method shows significant performance improvement on multi-domain and SGD datasets. |
Copied to clipboard
| Challenge: | Recent advances in large-scale language modeling and generation have enabled the creation of dialogue agents that exhibit human-like responses in a wide range of conversational scenarios. |
| Approach: | They propose a framework in which dialogue agents can evaluate the progression of a conversation toward or away from desired outcomes and use this signal to inform planning for subsequent responses. |
| Outcome: | The proposed framework evaluates the progression of a conversation toward or away from desired outcomes and uses this signal to inform planning for subsequent responses. |
Copied to clipboard
| Challenge: | Recent advances in techniques for generating realistic synthetic content pose a diverse set of problems with significant societal consequences. |
| Approach: | They propose to use paragraph-level detectors to detect tampering of full-length documents under a variety of threat models to detect machine-generated text. |
| Outcome: | The proposed detectors can detect the tampering of full-length documents under a variety of threat models. |
Copied to clipboard
| Challenge: | Prior work on labeling arguments extracted from peer review text has focused qualified labor force on labelling arguments extracted by the text. |
| Approach: | They synthesize label sets from prior work and extend them to include fine-grained annotations of review and rebuttal sentences. |
| Outcome: | The proposed dataset synthesizes label sets from prior work and extends them to include fine-grained annotation of review and rebuttal sentences. |
Copied to clipboard
| Challenge: | Existing reading comprehension datasets focus on single-span answers, but multi-spread questions are less studied. |
| Approach: | They propose a new reading comprehension dataset that focuses on multi-span questions . they introduce new metrics for the purposes of multi--spontaneous question answering evaluation . |
| Outcome: | The proposed model beats baselines and achieves state-of-the-art on the existing dataset. |
Copied to clipboard
| Challenge: | Existing paradigms for text entry in augmentative and alternative communication (AAC) for people with severe motor impairments require 3-5 predictions to save keystrokes. |
| Approach: | They propose a paradigm in which phrases are abbreviated aggressively as word-initial letters. |
| Outcome: | The proposed paradigm can save up to 77% on expansions on conversation turn . the proposed paradigm could be used in augmentative and alternative communication (AAC) |
Copied to clipboard
| Challenge: | Pre-trained language models encode correlations between social groups and traits, like associating the group with the group. |
| Approach: | They adapt the Agency-Belief-Communion (ABC) stereotype model to a language model and introduce the sensitivity test (SeT) to measure stereotypical associations. |
| Outcome: | The proposed framework is used to measure stereotyping of intersectional identities in language models. |
Copied to clipboard
| Challenge: | Existing algorithms for pre-trained language models lack performance indicators for linguistic tasks such as structured prediction. |
| Approach: | They propose to measure the degree to which labeled trees are recoverable from an LM’s contextualized embeddings by probing to rank LMs for parsing dependencies in a given language. |
| Outcome: | The proposed approach predicts the best LM choice 79% of the time using less compute than training a full parser. |
Copied to clipboard
| Challenge: | Literature in Natural Language Processing (NLP) typically labels whole language with strict type of morphology, e.g. fusional or agglutinative. |
| Approach: | They propose to quantify morphological typology at the word and segment level by using two indices: synthesis (e.g. analytic to polysynthetic) and fusion (agglutinative to fusional). |
| Outcome: | The proposed method reduces the rigidity of NLP classification claims by measuring morphological diversity at the word and segment level. |
Copied to clipboard
| Challenge: | Empirical results show that our proposed model outperforms the state-of-the-art methods in terms of both automatic evaluation metrics and human judgment. |
| Approach: | They propose a model which uses large-scale commonsense and named entity based knowledge to ground dialogue on external knowledge and topic-specific knowledge associated with each utterance. |
| Outcome: | The proposed model outperforms the state-of-the-art methods on two benchmark datasets. |
Copied to clipboard
| Challenge: | Existing methods to allow domain adaptation to diverse domains are expensive and require continuing training in-domain. |
| Approach: | They propose a method to permit domain adaptation to many diverse domains using a computationally efficient adapter approach. |
| Outcome: | The proposed method allows domain adaptation to many diverse domains while avoiding negative interference between unrelated domains. |
Copied to clipboard
| Challenge: | Existing models for detecting hate expressed with emojis have weaknesses when used for sensitive applications such as content moderation. |
| Approach: | They propose a test suite of 3,930 short-form statements that evaluates hateful language expressed with emoji. |
| Outcome: | The proposed model performs better on emoji-based hate while maintaining strong performance on text-only hate. |
Copied to clipboard
| Challenge: | a framework to evaluate the performance and cost trade-offs between machine-translated and manually-created labelled data is presented. |
| Approach: | They propose a framework to evaluate the performance and cost trade-offs between machine-translated and manually-created labelled data for task-specific fine-tuning of massively multilingual language models. |
| Outcome: | The proposed framework can be used to evaluate cost trade-offs between machine-translated and manually-created labelled data for task-specific fine-tuning of massively multilingual models. |
Copied to clipboard
| Challenge: | Paraphrase generation is an important natural language generation task . however, the effectiveness of paraphrase generation can be limited due to the limited data available. |
| Approach: | They propose a weakly supervised approach to paraphrase generation that leverages reinforcement learning for effective model training with data selection. |
| Outcome: | The proposed model improves the state-of-the-art performance on four weakly supervised paraphrase generation tasks. |
Copied to clipboard
| Challenge: | Despite advances in machine translation quality estimation and evaluation, decoding is mostly oblivious to this. |
| Approach: | They propose to use a decoding framework that is quality-aware for neural machine translation . they compare various methods like N-best reranking and minimum Bayes risk decoding . |
| Outcome: | The proposed quality-aware decoding outperforms MAP-based decoding on four datasets and two model classes. |
Copied to clipboard
| Challenge: | Federated Learning (FL) is a machine learning technique that trains a model across multiple distributed clients holding local data samples, without ever storing client data in a central location. |
| Approach: | They propose to use pretrained models to study three multilingual language tasks . they also examine impact of non-IID text on FL in naturally occurring data . |
| Outcome: | The proposed methods perform better than centralized learning even when using non-IID partitioning. |
Copied to clipboard
| Challenge: | Existing methods for fine-tuning pre-trained language models ignore the potential of unlabeled data. |
| Approach: | They propose a framework that allows users to unleash the power of unlabeled data via self-training. |
| Outcome: | The proposed framework outperforms active learning and self-training baselines and improves the label efficiency of PLM fine-tuning by 56.2% on average. |
Copied to clipboard
| Challenge: | a novel approach to contrastive learning for language understanding is not fully explored . contrastive training has been widely applied to self-supervised representation learning . |
| Approach: | They propose a label anchored contrastive learning approach for language understanding using a class label. |
| Outcome: | The proposed approach improves on GLUE and CLUE benchmarks by 4.1% compared to the state-of-the-art approaches . the proposed approach also improves under the few-shot and data imbalance settings . |
Copied to clipboard
| Challenge: | Existing systems that generate *flashbacks* are monotonic and lack explicit guidance on how to insert them. |
| Approach: | They propose to use event temporal orders to encode events as temporal prompts . they leverage a Plan-and-Write framework enhanced by reinforcement learning to generate storylines . |
| Outcome: | The proposed method generates more interesting stories with *flashbacks* while maintaining textual diversity, fluency, and temporal coherence. |
Copied to clipboard
| Challenge: | Existing studies have shown that social media can help predict rises in infectious disease caseloads. |
| Approach: | They propose to use transformer-based language models to integrate infectious disease modelling into reddit embedding features in reddits in specific US states. |
| Outcome: | The proposed model outperforms other features at predicting upward trend signals in areas where epidemiological data is unreliable. |
Copied to clipboard
| Challenge: | In automatic essay grading, essay traits are important for scoring the essay holistically . a single-task learning system gives the best results for scoring essays holistically and scoring essay traits. |
| Approach: | They propose a way to score essays using a multi-task learning approach . they compare the MTL-based BiLSTM system to a single-task Learning approach based on LSTMs and BiLStms . |
| Outcome: | The proposed system gives better results for scoring essay holistically and scoring essay traits. |
Copied to clipboard
| Challenge: | Existing datasets focus on a single medium, information domain or specific application . authors propose novel methods for automated veracity assessment based on Natural Language Inference . |
| Approach: | They propose to build a PANACEA dataset that combines different data sources with different foci to ensure a unique set of claims. |
| Outcome: | The proposed methods are competitive with SOTA methods and provide a detailed discussion. |
Copied to clipboard
| Challenge: | Desire is a primitive instinct and a need for strongly expressing human desires to get or possess something. |
| Approach: | They propose to use MSED to model and understand human desire . they propose to provide a benchmark for human desire analysis . |
| Outcome: | The proposed dataset contains 9,190 text-image pairs with English text. |
Copied to clipboard
| Challenge: | Existing document-level relation extraction methods do not distinguish between mention-level features and entity-level feature . document-based methods are more challenging because of multiple mentions of entities. |
| Approach: | They propose a method which selectively attentions different entity mentions with respect to candidate relations and performs relation-specific representations of entities. |
| Outcome: | The proposed method improves relation-specific representations of entities on two benchmark datasets. |
Copied to clipboard
| Challenge: | Detecting out-of-context media is a problem in domains of public significance . a method that leverages automatically generated hard image-text mismatches is proposed . |
| Approach: | They propose a method that leverages automatically generated hard image-text mismatches to detect out-of-context media . they analyze tweets relevant to topics such as COVID-19, Climate Change and Military Vehicles . |
| Outcome: | The proposed method improves detection accuracy over a strong baseline on a set of fakes created by humans. |
Copied to clipboard
| Challenge: | Standard evaluation metrics, e.g., BLEU, TER and METEOR, focus on the quality of translations at the sentence level and do not consider discourse-level features. |
| Approach: | They propose to use a metric to take discourse coherence into consideration by categorizing discourse-related spans and calculating the similarity-based F1 measure of categorized spans. |
| Outcome: | The proposed metric possesses better selectivity and interpretability at the document-level, and is more sensitive to document- level nuances. |
Copied to clipboard
| Challenge: | Existing approaches to detect vaccine attitudes on social media require abundant annotations and pre-defined aspect categories. |
| Approach: | They propose a semi-supervised approach to detect vaccine attitudes on social media . they use an autoencoding architecture to learn from unlabelled data the topical information of the domain . |
| Outcome: | The proposed model outperforms existing aspect-based models on stance detection and tweet clustering. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated human-level performance on a vast spectrum of natural language tasks. |
| Approach: | They propose a method to infuse structured knowledge into large language models by directly training T5 models on factual triples of knowledge graphs (KGs). |
| Outcome: | The proposed method outperforms baseline models on FreebaseQA and WikiHop, as well as the Wikidata-answerable subset of TriviaQA and NaturalQuestions. |
Copied to clipboard
| Challenge: | Existing studies show that multilingual pre-trained models can learn to generalise across languages . however, it remains unclear how these models learn to learn multilingual representations . |
| Approach: | They propose a hypothesis that multilingual pre-trained models can derive language-universal abstractions about grammar by aligning morphosyntactic markers that fulfil a similar grammatical function across languages. |
| Outcome: | The proposed model can derive language-universal abstractions even without explicit supervision. |
Copied to clipboard
| Challenge: | Existing approaches to classify aspects with aspect sentiment bias are hard to find . |
| Approach: | They propose a no-aspect differential sentiment framework for the ABSA task that eliminates aspect sentiment bias and uses differential sentiment loss instead of cross-entropy loss to better classify the sentiments. |
| Outcome: | The proposed framework can be combined with almost all traditional ABSA methods. |
Copied to clipboard
| Challenge: | Existing methods for training pre-trained language models have limited practicality due to latency requirements. |
| Approach: | They propose a method that uses a Mixture-of-Experts structure to increase model capacity and inference speed. |
| Outcome: | The proposed method outperforms existing distillation methods on natural language understanding and question answering tasks. |
Copied to clipboard
| Challenge: | Recent studies show that self-attention based models have limitations on modeling sequential transformations. |
| Approach: | They propose to extract some explainable features from trained RNNs that are reminiscent of classical n-grams features. |
| Outcome: | The proposed models can model interesting linguistic phenomena such as negation and intensification. |
Copied to clipboard
| Challenge: | Existing approaches to Visual Question Generation (VQG) are trained to mimic an arbitrary choice of concept but only one or a few are captured by the human references. |
| Approach: | They propose a variant of Visual Question Generation which conditions the question generator on categorical information based on expectations on the type of question and the objects it should explore. |
| Outcome: | The proposed model improves on the current state of the art on an answer-category augmented VQA dataset and human evaluation validates that guidance helps the generation of questions that are grammatically coherent and relevant to the given image and objects. |
Copied to clipboard
| Challenge: | Existing methods to predict logical forms ignore the utilization of symbolic operations and lack reasoning ability and interpretability. |
| Approach: | They propose an operation-pivoted discrete reasoning framework that uses symbolic operations as neural modules to facilitate reasoning ability and interpretability. |
| Outcome: | Extensive experiments on DROP and RACENum datasets show the reasoning ability of OPERA. |
Copied to clipboard
| Challenge: | Existing approaches to Multi-document summarization are limited due to the extremely long input length. |
| Approach: | They propose an extract-then-abstract Transformer framework to overcome the problem . they leverage pre-trained language models to construct hierarchical extractors and abstractors . |
| Outcome: | The proposed framework outperforms baseline models with comparable model sizes and achieves the best results on the Multi-News, Multi-XScience, and WikiCatSum corpora. |
Copied to clipboard
| Challenge: | Existing methods of span representation are based on simple derivations from word representations and do not utilize compositional structures of natural language. |
| Approach: | They propose a hypertree neural network that is structured with constituency parse trees to improve representations of constituent spans. |
| Outcome: | The proposed model improves representations of constituent spans using constituency parse trees. |
Copied to clipboard
| Challenge: | An increasing awareness of biased patterns in natural language processing resources such as BERT has motivated many metrics to quantify ‘bias’ and ‘fairness’. |
| Approach: | They combine literature survey, correlation analysis and empirical evaluations to evaluate compatibility of fairness metrics for pre-trained language models and their downstream tasks. |
| Outcome: | The proposed measures are not compatible with each other and highly depend on (i) templates, (ii) attribute and target seeds and (iv) the choice of embeddings. |
Copied to clipboard
| Challenge: | Recent studies show that shallow semantic role labeling (SRL) performance drops under out-of-domain setting. |
| Approach: | They propose to annotate a multi-domain Chinese predicate-argument dataset using a frame-free annotation methodology and strict double annotation for improving data quality. |
| Outcome: | The proposed dataset is compared with a dataset from six different domains. |
Copied to clipboard
| Challenge: | Existing language modeling pretraining objectives do not take structural information of conversational text into account. |
| Approach: | They propose a structure-aware Mutual Information based loss-function DMI for training dialog-representation models that captures the inherent uncertainty in response prediction. |
| Outcome: | The proposed model outperforms strong baseline models on nine diverse tasks. |
Copied to clipboard
| Challenge: | Existing word-level approaches to attack text are limited to a single word . existing methods ignore interactions between consecutive words, resulting in one-to-one attacks . |
| Approach: | They propose a black-box attack framework that misleads the language model by applying variable-length contextualized transformations to the original text. |
| Outcome: | The proposed framework outperforms existing methods on classification and inference tasks. |
Copied to clipboard
| Challenge: | Non-autoregressive translation models suffer from the multi-modality problem when a source sentence corresponds to multiple correct translations. |
| Approach: | They propose to decompose the syntactic multi-modality problem into short- and long-range models and evaluate them on synthesized and real datasets. |
| Outcome: | The proposed loss functions can handle short- and long-range syntactic multi-modalities better than existing models. |
Copied to clipboard
| Challenge: | Current methods for interpolative data augmentation select samples at random, which might make it difficult for the model to generalize better and converge faster. |
| Approach: | They propose a curriculum-based learning method that leverages the relative position of samples in hyperbolic embedding space as a complexity measure to gradually mix up increasingly difficult and diverse samples along training. |
| Outcome: | The proposed method achieves state-of-the-art results over existing methods on 10 benchmark datasets across 4 languages in text classification and named-entity recognition tasks. |
Copied to clipboard
| Challenge: | Existing methods focused on clustering sentences to indicate information saliency and avoid redundancy. |
| Approach: | They propose to group together sub-sentential propositions to generate a representative sentence for each cluster via text fusion. |
| Outcome: | The proposed method improves over the previous state-of-the-art method in the DUC 2004 and TAC 2011 datasets, both in automatic ROUGE scores and human preference. |
Copied to clipboard
| Challenge: | Efficient machine translation models are commercially important as they can increase inference speeds, reduce costs and carbon emissions. |
| Approach: | They compare NAR models with autoregressive models to evaluate their performance . they point out flaws in evaluation methodology and argue for consistent evaluation . |
| Outcome: | The proposed model is faster on GPUs, but slower under more realistic usage conditions. |
Copied to clipboard
| Challenge: | Massively multilingual Transformers (MMTs) have dominated research in multilingual NLP and cross-lingual transfer recently. |
| Approach: | They propose to learn bilingual language pair adapters (BAs) when the goal is to optimize performance for a particular source-target transfer direction. |
| Outcome: | The proposed framework improves performance in three standard downstream tasks and for the majority of low-resource languages. |
Copied to clipboard
| Challenge: | Parody is a figurative device used for mimicking entities for comedic or critical purposes. |
| Approach: | They propose a multi-encoder model that combines three parallel encoders to enrich parody-specific representations with humor and sarcasm information. |
| Outcome: | The proposed model outperforms state-of-the-art methods on a dataset of political parody tweets. |
Copied to clipboard
| Challenge: | Existing models for structural reading comprehension (SRC) only focus on comprehension of plain text, tables, tables or knowledge bases. |
| Approach: | They propose a topological information enhanced model which transforms a token-level task into a tag-level one by introducing a two-stage process. |
| Outcome: | The proposed model outperforms baselines and achieves state-of-the-art performance on the web-based SRC benchmark WebSRC at the time of writing. |
Copied to clipboard
| Challenge: | Using a framework based on Rhetorical Structure Theory, we aim to improve the cohesion and coherence of long-form text generated by language models. |
| Approach: | They propose a framework that utilises Rhetorical Structure Theory to control the discourse structure, semantics and topics of generated text. |
| Outcome: | The proposed framework performs competitively against existing models while offering significantly more controls over generated text than alternative methods. |
Copied to clipboard
| Challenge: | Existing approaches to intent detection rely on epoch wise clustering and classification based on labeled and unlabeled data. |
| Approach: | They propose an end-to-end deep contrastive clustering algorithm that jointly updates model parameters and cluster centers via supervised and self-supervised learning. |
| Outcome: | The proposed approach outperforms baselines on five public datasets and human-in-the-loop variant for practical deployment. |
Copied to clipboard
| Challenge: | Existing datasets for sentence fusion tasks are limited in size and scope . despite recent advances, cross-document tasks such as multi-document summarization have not progressed with the same pace. |
| Approach: | They propose to extend a sentence fusion dataset by almost four times its original size . they relabel the dataset and employ more data sources to improve model performance . |
| Outcome: | The proposed dataset triples the size of an earlier dataset and improves performance . it also includes more complex training instances better reflecting those found in "the wild" |
Copied to clipboard
| Challenge: | Neural Machine Translation models can be optimized to improve latency by constraining the set of output words . lexical shortlisting fails to select the right set of input words for semantically non-compositional phenomena such as idiomatic expressions. |
| Approach: | They propose a model of vocabulary selection that constrains the set of allowed output words . they propose to increase the size of the allowed set to restore translation quality . |
| Outcome: | The proposed model restores translation quality of an unconstrained system, as measured by human evaluations on WMT newstest2020 and idiomatic expressions, at an inference latency competitive with alignment-based selection using aggressive thresholds. |
Copied to clipboard
| Challenge: | Citation context analysis (CCA) is an important task in natural language processing that studies how and why scholars discuss each other’s work. |
| Approach: | They propose to use a dataset of 12.6K citation contexts from 1.2K computational linguistics papers to model three important CCA phenomena. |
| Outcome: | The proposed dataset contains 12.6K citation contexts from 1.2K computational linguistics papers and can model these phenomena. |
Copied to clipboard
| Challenge: | Existing models for event extraction require expensive human annotations. |
| Approach: | They propose a data-efficient event extraction model that formulates event extraction as a conditional generation problem. |
| Outcome: | The proposed model can be trained with only a few labeled examples. |
Copied to clipboard
| Challenge: | Existing methods to train cross-lingual pre-trained language models have shown great success in cross-linguistic sequence labeling tasks. |
| Approach: | They propose a cross-lingual language informative span masking task to eliminate the objective gap between pre-training and fine-tuning stages. |
| Outcome: | The proposed method surpasses the state-of-the-art methods on multiple benchmarks even with limited pre-training data. |
Copied to clipboard
| Challenge: | Named entity recognition (NER) is a fundamental and important task in natural language processing. |
| Approach: | They propose a novel Hero-Gang Neural structure to leverage both global and local information to promote NER by using a Transformer-based encoder and a Gang module. |
| Outcome: | The proposed model can extract local features and position information from the Hero and Gang modules, and it performs on multiple datasets. |
Copied to clipboard
| Challenge: | Existing methods for text classification fail to generalize to unseen classes with very few labeled text instances per class. |
| Approach: | They propose a meta-learning method which performs instance-wise comparison followed by aggregation to generate class-wise matching vectors instead of prototype learning. |
| Outcome: | Experiments show that the proposed method outperforms existing methods under both the standard and generalized FSL settings. |
Copied to clipboard
| Challenge: | Visual Question Answering (VQA) has benefited from increasingly sophisticated models, but has not enjoyed the same level of engagement in terms of data creation. |
| Approach: | They propose a method that automatically derives VQA examples at volume by leveraging existing image-caption annotations combined with neural models for textual question generation. |
| Outcome: | The proposed method improves state-of-the-art zero-shot accuracy by double digits and achieves robustness that lacks in the same model trained on human-annotated VQA data. |
Copied to clipboard
| Challenge: | Using a simple logistic regression algorithm, we combine GEC models for binary classification. |
| Approach: | They propose a logistic regression algorithm that can combine GEC models with binary classification. |
| Outcome: | The proposed method outperforms the state-of-the-art by 4.2 points on the CoNLL-2014 and 7.2 points on BEA-2019 test sets. |
Copied to clipboard
| Challenge: | Existing models for NLP tasks require long text sequences beyond the length limit of pretrained models. |
| Approach: | They propose to pretrain large-size NLP models using the same long-doc corpus and fine tune them for real-world long-context tasks. |
| Outcome: | The proposed models can perform better under standard pretraining paradigms than longformer and Longformer. |
Copied to clipboard
| Challenge: | XML-CNN has been a popular research topic in NLP due to its superior performance . however, the increasing complexity brings difficulties to ensure the true architectural progress . |
| Approach: | They propose to re-examine an influential multi-label text classification method . they propose suitable baselines for multi-level text classification tasks . |
| Outcome: | The proposed method performs better than the original model, the authors show . they show that the re-implementation reveals contradictory results to the original work . |
Copied to clipboard
| Challenge: | Existing methods for Automatic Short Answer Grading (ASAG) ignore structural context and therefore do not perform well. |
| Approach: | They propose a Multi-Relational Graph Transformer to prepare token representations considering the structural context of a sentence. |
| Outcome: | The proposed model outperforms existing state-of-the-art methods on a dataset from an undergraduate computer science course. |
Copied to clipboard
| Challenge: | Experimental results show that a new method for learning event schemas from historical events is effective. |
| Approach: | They propose a new event schema induction framework which captures global dependencies among nodes in event graphs. |
| Outcome: | Experimental results show that the proposed model can learn event schemas with global consistency. |
Copied to clipboard
| Challenge: | CS1QA is a dataset for code-based question answering in the programming education domain. |
| Approach: | They propose a dataset for code-based question answering in the programming education domain. |
| Outcome: | The proposed model can be used as a benchmark for source code comprehension and question answering in the educational setting. |
Copied to clipboard
| Challenge: | Recent successes of NLP systems require large amounts of labelled data for structured prediction tasks. |
| Approach: | They propose a method for unsupervised transfer from multiple input models for structured prediction using a cross-lingual setup. |
| Outcome: | The proposed method produces less noisy labels for the distant supervision. |
Copied to clipboard
| Challenge: | Neural text generation models are typically trained by maximizing log-likelihood with the sequence cross entropy (CE) loss. |
| Approach: | They propose an Edit-Invariant Sequence Loss method which computes the matching loss of a target sequence with all n-grams in the generated sequence. |
| Outcome: | The proposed method outperforms the common CE loss and strong baselines on a wide range of tasks. |
Copied to clipboard
| Challenge: | Exemplification is a process by which writers explain or clarify a concept by providing an example. |
| Approach: | They propose to use a partially-written answer to query a large set of human-written examples extracted from a corpus to determine exemplification quality. |
| Outcome: | The proposed model is able to retrieve human-written examples from a corpus and show that it is more relevant than state-of-the-art models. |
Copied to clipboard
| Challenge: | Existing methods for out-of-scope (OOS) detection use classifier confidence score, but model cannot infer correctly. |
| Approach: | They propose a zero-shot post-processing step that exploits the classification confidence score and the shape of the entire output distribution. |
| Outcome: | The proposed method improves performance when there is no OOS training data and learning procedure when OOS data is available. |
Copied to clipboard
| Challenge: | Existing evaluation metrics for summarization use human annotations as reference. |
| Approach: | They propose a new automatic reference-free evaluation metric that compares semantic distribution between source document and summary by pretrained language models and considers summary compression ratio. |
| Outcome: | The proposed metric is more consistent with human evaluation in terms of coherence, consistency, relevance and fluency. |
Copied to clipboard
| Challenge: | a recent study shows that over-parameterized pre-trained language models are unsuitable for low-capacity devices. |
| Approach: | They propose a transformer-based pre-trained language model that is overparameterized . they use a two-stage knowledge distillation scheme to train the model . |
| Outcome: | The proposed model outperforms state-of-the-art models on well-known NLP benchmarks. |
Copied to clipboard
| Challenge: | Recent large-scale language models have produced human-like responses in open-domain dialogue systems. |
| Approach: | They propose a framework for imposing roles on open-domain dialogue systems . they use few-shot learning to build a Korean dialogue dataset from scratch . |
| Outcome: | The proposed framework meets role specifications while maintaining conversational abilities. |
Copied to clipboard
| Challenge: | named entity recognition (NER) tasks are often dominated by the majority of non-entity tokens in text . a data imbalance problem is causing the NER models to ignore named entities . |
| Approach: | They propose a set of sentence-level resampling methods to reduce data imbalance . they use a training sentence to compute the importance of each training sentence based on its tokens and entities . |
| Outcome: | The proposed methods outperform sub-sentence-level resampling, data augmentation, and loss functions on multiple corpora. |
Copied to clipboard
| Challenge: | Existing word embeddings are high-dimensional and consume considerable computational resources. |
| Approach: | They propose a method to decompose the desiderata of word embeddings into two parts, completeness and soundness, and focus on soundness. |
| Outcome: | The proposed method is extremely efficient and provides minimal means to handle word embeddings. |
Copied to clipboard
| Challenge: | a growing effort in NLP aims to build datasets of human explanations, but it remains unclear whether they serve their intended goals. |
| Approach: | They argue that the term "explanation" is overloaded and refers to a broad range of notions with different properties and ramifications. |
| Outcome: | The proposed datasets examine the diversity of explanations and their use in NLP. |
Copied to clipboard
| Challenge: | a growing popularity of deep-learning models makes model understanding more important . feature attribution methods have shown promising results in computer vision but are not trivial . |
| Approach: | They propose a gradient-based feature attribution method that smooths gradients by aggregating similar reference texts derived from language model embeddings. |
| Outcome: | The proposed method outperforms existing methods on public datasets and key words detection tasks. |
Copied to clipboard
| Challenge: | Existing curriculum learning approaches for relation extraction are lacking in text graphs. |
| Approach: | They propose a generic and trend-aware curriculum learning approach that integrates textual and structural information in text graphs for relation extraction between entities. |
| Outcome: | The proposed model shows improvement over state-of-the-art methods across several datasets. |
Copied to clipboard
| Challenge: | Modern unsupervised machine translation systems reach reasonable translation quality under clean and controlled data conditions. |
| Approach: | They compare unsupervised and supervised machine translation systems of similar quality . they combine the benefits of both methods into a single system . |
| Outcome: | The proposed system improves adequacy and fluency as measured by human evaluators. |
Copied to clipboard
| Challenge: | Existing methods to augment retrieval-augmented generation models with retrievers often rely on spurious cues or generate hallucinations during inference. |
| Approach: | They propose a method to incorporate evidentiality of passages into training a retrieval-augmented generation model. |
| Outcome: | The proposed method outperforms its direct counterpart on all knowledge-intensive tasks. |
Copied to clipboard
| Challenge: | Currently, commonsense reasoning systems are limited by expensive data annotations and overfitting to a specific benchmark. |
| Approach: | They propose to transform a commonsense knowledge graph into synthetic QA-form samples for model training. |
| Outcome: | The proposed framework improves performance with multiple commonsense KGs on five commonsensense reasoning benchmarks. |
Copied to clipboard
| Challenge: | Existing models focus on synthesizing a dialogue with proper knowledge, but neglect that the same knowledge could be expressed differently even under the same context. |
| Approach: | They propose a model that ground dialogue generation by extra knowledge by analyzing the structure of the response and the content style of each part. |
| Outcome: | The proposed model can learn the structure style defined by a few examples and generate responses in desired content style. |
Copied to clipboard
| Challenge: | Existing methods for speaker identification in texts are incomplete and introduce errors that propagate and seriously affect the final output. |
| Approach: | They propose to use speaker identification (SI) in texts to identify the speaker(s) for each utterance in texts. |
| Outcome: | The proposed model can achieve comparable or better than previous state-of-the-art methods on all public SI datasets for Chinese. |
Copied to clipboard
| Challenge: | Existing methods for ED in IE and NLP focus on feature-based models to feature-driven models. |
| Approach: | They propose to use a multilingual dataset to annotate events for 8 different languages . they demonstrate the challenges and transferability of ED across languages in MINION . |
| Outcome: | a new dataset that consistently annotates events for 8 different languages is released . the new dataset will promote future research on multilingual ED . |
Copied to clipboard
| Challenge: | Recent studies show that prompts help models to learn faster in the same way that humans learn faster when provided with task instructions expressed in natural language. |
| Approach: | They experiment with 30 prompts manually written for natural language inference (NLI) they find that models can learn just as fast with many irrelevant or pathologically misleading prompts . |
| Outcome: | The proposed model can learn as fast with irrelevant or pathologically misleading prompts as with instructively “good” prompts. |
Copied to clipboard
| Challenge: | Dense retrieval approaches suffer from the lexical gap and require large amounts of training data. |
| Approach: | They propose an unsupervised method for domain adaptation that uses query generator and pseudo labeling from a cross-encoder to improve retrieval performance. |
| Outcome: | The proposed method outperforms state-of-the-art retrieval methods on domain-specialized datasets by 9.3 points nDCG@10 on six tasks. |
Copied to clipboard
| Challenge: | Existing methods to reduce inference cost by distilling transformer models into lightweight student models are limited for high-volume use cases. |
| Approach: | They propose to distill state-of-the-art transformer models into lightweight student models to reduce computation cost at inference time. |
| Outcome: | The proposed pipeline achieves up to 600x speed-up on GPUs and CPUs on six single-sentence text classification tasks and in domain generalization settings. |
Copied to clipboard
| Challenge: | Existing approaches to pre-training/fine-tuning are focusing on the alignment of pre-trained and fine-tuned PLMs with large-scale discourse structures. |
| Approach: | They propose a novel approach to infer discourse information for arbitrarily long documents using supervised, distantly supervised and simple baselines. |
| Outcome: | The proposed approach shows that the captured discourse information is local and general, even across fine-tuning tasks. |
Copied to clipboard
| Challenge: | Existing methods for relation extraction only implicitly learn to model relevant contexts and entity types while being trained for RE. |
| Approach: | They propose to explicitly teach the model to capture relevant contexts and entity types by supervising and augmenting intermediate steps (SAIS) for RE. |
| Outcome: | The proposed method outperforms the runner-up method on three benchmarks by 5.04% . textual contexts and entity types are the major information sources that lead to the success of previous approaches. |
Copied to clipboard
| Challenge: | To-do texts are often short and under-specified, which poses a challenge for current text representation models. |
| Approach: | They propose a neural multi-task learning framework that extracts representations of English to-do tasks with a multi-head attention mechanism on top of a pre-trained text encoder. |
| Outcome: | The proposed model outperforms baseline models on four downstream tasks and achieves error reduction of 38.7%. |
Copied to clipboard
| Challenge: | a quality summarization dataset requires the production and evaluation of summaries by trained humans and machines. |
| Approach: | They translate a summarization dataset in English and compare its performance to seven languages . they explore equivalence testing as an appropriate statistical paradigm for evaluating correlations between human and automated scoring of summaries . |
| Outcome: | The proposed method could be used in seven languages and compares performance across measures. |
Copied to clipboard
| Challenge: | Mental health disorders are one of the primary causes of disability worldwide . lack of qualified and competent mental health professionals is a major problem . we propose a virtual assistant that can act as the first point of contact and comfort for mental health patients. |
| Approach: | They propose a virtual assistant that can act as the first point of contact and comfort for mental health patients. |
| Outcome: | The proposed system outperforms baselines in the evaluation of 7k dyadic conversations from a peer-to-peer support platform. |
Copied to clipboard
| Challenge: | Existing studies on automatic summary evaluation metrics focus on lexical similarity and require a reference summary which is expensive to obtain. |
| Approach: | They propose to use a weakly supervised summary evaluation approach without the presence of reference summaries to transform existing summarization datasets into corrupted reference summarizers. |
| Outcome: | The proposed method outperforms baselines and shows that it improves linguistic quality over all metrics. |
Copied to clipboard
| Challenge: | Existing approaches to handle knowledge acquisition bottlenecks in multilingual training are limited due to the curse of multilinguality. |
| Approach: | They propose to use large pre-trained monolingual language models in cross lingual zero-shot word sense disambiguation coupled with a contextualized mapping mechanism. |
| Outcome: | The proposed model improves the average F-score by nearly 6.5 points over 17 target languages. |
Copied to clipboard
| Challenge: | a neural machine translation system generates a translation t in the target language, but for any sentence of non-trivial complexity, the translation s is not unique. |
| Approach: | They propose a method to quantify the amount of information missing in a machine translation system. |
| Outcome: | The proposed model captures extra information from a single float representation of the target sentence and reproduces it with two 32-bit floats per target token. |
Copied to clipboard
| Challenge: | Word-in-Context (WiC) task has attracted considerable attention in the NLP community, as demonstrated by the popularity of the recent MCL-Wic SemEval shared task. |
| Approach: | They propose to use lexical resources from word sense disambiguation and target sense verification to reduce the relationship between the two tasks. |
| Outcome: | The proposed methods can be pairwise reduced to each other and therefore work in practice. |
Copied to clipboard
| Challenge: | Pre-trained language models that use subword tokenization schemes can succeed at a variety of language tasks that require character-level information. |
| Approach: | They propose to use word tokenization schemes to probe what word pieces encode . they show that larger models can encode character-level information . |
| Outcome: | The proposed models can encode character-level information and perform better on non-Latin alphabets. |
Copied to clipboard
| Challenge: | Community Question Answering (CQA) fora lack a dataset to produce answer summarizations . a novel dataset of 4,631 CQA threads is used to generate answer summaries . |
| Approach: | They propose a dataset of 4,631 CQA threads for answer summarization curated by professional linguists. |
| Outcome: | The proposed approach boosts summarization performance according to automatic evaluation. |
Copied to clipboard
| Challenge: | Recent studies show that pre-trained transformers perform poorly for multi-candidate inference tasks. |
| Approach: | They propose a pre-training objective that models paragraph-level semantics across multiple input sentences. |
| Outcome: | The proposed model outperforms existing models on three AS2 and one fact verification datasets. |
Copied to clipboard
| Challenge: | Text style transfer (TST) is a task that aims to change the style of a text from source to target while preserving its content. |
| Approach: | They propose a method to incorporate syntactic and semantic information into similarity computation between the source and the converted text. |
| Outcome: | The proposed method is superior in both supervised and unsupervised settings. |
Copied to clipboard
| Challenge: | Recent work has found that multi-task training with a large number of diverse tasks can uniformly improve downstream performance on unseen target tasks. |
| Approach: | They aim to disentangle the effect of scale and relatedness of tasks in multi-task representation learning by increasing the number of tasks and incorporating smaller sets of related tasks. |
| Outcome: | The proposed model improves on unseen target tasks by increasing the scale of multi-task learning to incorporate more tasks and developing similarity metrics to incorporate tasks related to the target task. |
Copied to clipboard
| Challenge: | Existing systems that can perform interactive summarization cannot ingest the full document set or operate at sufficient speed for interactivity. |
| Approach: | They propose two deep reinforcement learning models for interactive summarization task . they use interactive session state and history to refrain from redundancy . |
| Outcome: | The proposed model improves informativeness while preserving positive user experience. |
Copied to clipboard
| Challenge: | Existing models only classify text excerpts as offensive or not, failing to provide information on which words and phrases contribute the most to its offensive tone. |
| Approach: | They propose a model for offensive span detection that uses a pre-trained language model to generate training data. |
| Outcome: | The proposed model can detect offensive spans in a text snippet using a pre-trained language model . the proposed model is able to detect offensive text in simulated training conditions . |
Copied to clipboard
| Challenge: | Experimental results show that neural machine translation engines built via FL can be easily adapted when an FL-based aggregation is applied to fuse different domains. |
| Approach: | They propose to use federated learning to fuse mixed-domain translation models with a centralized aggregation to improve their performance. |
| Outcome: | The proposed model can be easily adapted to a mixed-domain translation model with slight modifications in the training process and perform on par with state-of-the-art training models. |
Copied to clipboard
| Challenge: | Existing studies on text summarization factual consistency are divided into two categories . entailment-based and question answering-based metrics are the most efficient . |
| Approach: | They propose an optimized QA-based metric that improves factual consistency by 14% . they compare entailment-based and QA metrics to find the best fit . |
| Outcome: | The proposed metric outperforms the best performing entailment-based metric on the SummaC factual consistency benchmark. |
Copied to clipboard
| Challenge: | Existing studies of gender bias in NLP focus on extrinsic or intrinsic bias, but the relationship between extrindic and intrinsic bias is relatively unknown. |
| Approach: | They propose a framework to measure extrinsic and intrinsic bias together and propose metric to measure debiasing and intrinsic debiases. |
| Outcome: | The proposed framework provides a comprehensive perspective on bias in NLP models, which can be applied to deploy NLP systems in a more informed manner. |
Copied to clipboard
| Challenge: | a typical approach to natural language processing tasks involves selecting text spans and making decisions about them. |
| Approach: | They propose a grammar-based structured span selection model which learns to make use of partial span annotations. |
| Outcome: | The proposed model improves on two popular span prediction tasks. |
Copied to clipboard
| Challenge: | Semantic typing aims at classifying tokens into semantic categories such as relations, entity types, and event types. |
| Approach: | They propose a unified framework for semantic typing that captures label semantics by projecting both inputs and labels into a joint semantic embedding space. |
| Outcome: | The proposed framework achieves strong performance across three semantic typing tasks. |
Copied to clipboard
| Challenge: | In-context learning is a new paradigm in natural language understanding . large pre-trained language models can be expensive to update . |
| Approach: | They propose an efficient method for retrieving training examples as prompts from annotated data and an LM. |
| Outcome: | The proposed method outperforms prior work and multiple baselines on three sequence-to-sequence tasks. |
Copied to clipboard
| Challenge: | XAI features usually provide a single importance score for each token, but feature attribution methods provide two complementary and theoretically-grounded scores for each utterance. |
| Approach: | They propose a feature attribution method that generates explicit perturbations of the input text, allowing the importance scores themselves to be explainable. |
| Outcome: | The proposed method explain the predictions of hate speech detection models on a set of curated examples from a test suite. |
Copied to clipboard
| Challenge: | Dense retrievers for open domain question answering have been shown to achieve impressive performance by training on large datasets of question-passage pairs. |
| Approach: | They propose to use recurring spans to create pseudo examples for contrastive learning. |
| Outcome: | The proposed model outperforms all pretrained baselines on a wide range of ODQA datasets and is competitive with BM25, a strong sparse baseline. |
Copied to clipboard
| Challenge: | Recent models such as RAG and REALM incorporate retrieval into conditional generation. |
| Approach: | They propose a method that combines retrieval and reranking into a BART-based sequence-to-sequence generation. |
| Outcome: | The proposed model combines retrieval and reranking into a BART-based sequence-to-sequence generation. |
Copied to clipboard
| Challenge: | Current text classifiers are subject to adversarial attacks from adversaries, typically executed using machine learning methods. |
| Approach: | They propose a novel and intuitive defense strategy called Sample Shielding that is attacker and classifier agnostic and does not require reconfiguration of the classifier or external resources. |
| Outcome: | The proposed defense is attacker and classifier agnostic and does not require reconfiguration of the classifier or external resources and is simple to implement. |
Copied to clipboard
| Challenge: | Artificial Intelligence (AI) and Machine Learning (ML) systems are becoming more popular and are causing concerns over user privacy. |
| Approach: | They propose a method for training ML models using positive and negative user feedback and a framework to extract labels on edge to make FL viable. |
| Outcome: | The proposed method improves significantly over a self-training baseline, achieving performance closer to models trained with full supervision. |
Copied to clipboard
| Challenge: | Masked Language Models (MLMs) pre-trained by predicting masked tokens on large corpora have been used successfully in natural language processing tasks for a variety of languages. |
| Approach: | They propose to use English attribute word lists to evaluate bias in eight languages without manually annotating data. |
| Outcome: | The proposed model significantly correlates with the existing English datasets for gender bias. |
Copied to clipboard
| Challenge: | Targeted Sentiment Analysis (TSA) is a task for generating insights from consumer reviews. |
| Approach: | They propose a multi-domain TSA system that augments a given training set with diverse weak labels from assorted domains and augments it with Yelp reviews. |
| Outcome: | The proposed model outperforms manual methods on three evaluation datasets across different domains and shows that it performs well. |
Copied to clipboard
| Challenge: | Neural abstractive summarization models generate factually inconsistent summaries . previous work has introduced the task of recognizing factual inconsistency as a downstream application of natural language inference (NLI). |
| Approach: | They propose a data generation pipeline that enables a task-oriented approach to detect factual inconsistencies in abstractive summarization models. |
| Outcome: | The proposed model improves the state-of-the-art performance across four benchmarks for recognizing factual inconsistency in generated summaries. |
Copied to clipboard
| Challenge: | Named entity recognition (NER) models trained on CoNLL do not transfer well to other domains, even within the same language. |
| Approach: | They propose a token-level gating layer to augment pre-trained multilingual transformers with gazetteers containing named entities (NE) from a target language or domain. |
| Outcome: | The proposed model improves on cross-lingual transfer with an F1 score of 92.92 for English and an average of 89.43 across all languages in CoNLL. |
Copied to clipboard
| Challenge: | Large language models can do in-context learning by conditioning on a few training examples with no parameter updates or task-specific templates. |
| Approach: | They propose a meta-training framework where a pretrained language model is tuned to do in-context learning on a large set of training tasks. |
| Outcome: | The proposed framework outperforms baseline models on 142 NLP datasets and a range of target tasks with domain shifts. |
Copied to clipboard
| Challenge: | Existing conversation models treat knowledge selection as a sentence ranking problem where each sentence is handled individually, ignoring the internal semantic connection between sentences. |
| Approach: | They propose to automatically convert background knowledge documents into document semantic graphs and perform knowledge selection over such graphs. |
| Outcome: | The proposed model improves on the knowledge selection task and the response generation task on HollE and generalizes on unseen topics in WoW. |
Copied to clipboard
| Challenge: | Recent work has shown that language models are susceptible to biases present in the training dataset. |
| Approach: | They propose to use natural sentence prompts to analyze gender-occupation biases in language models. |
| Outcome: | The proposed dataset can be used to analyze gender-occupation biases in language models. |
Copied to clipboard
| Challenge: | Existing work to generate adversarial attacks is costly and not scalable . despite the abundance of research in this area, little attention has been given to adversarials . |
| Approach: | They propose an adversarial attack mechanism that mitigates toxic language generation . they propose a defense mechanism that is scalable and can be generalized . |
| Outcome: | The proposed defense is effective at avoiding toxic language generation even against imperceptible toxicity triggers while preserving conversational flow. |
Copied to clipboard
| Challenge: | Existing methods to protect sensitive data from leaking are over-pessimistic and undifferentiated. |
| Approach: | They propose a new privacy notion, selective differential privacy, to provide rigorous privacy guarantees on the sensitive portion of the data to improve model utility. |
| Outcome: | The proposed privacy-preserving mechanism achieves better utility while remaining safe under various privacy attacks compared to baselines. |
Copied to clipboard
| Challenge: | Distributional models learn representations of words from text but lack grounding or the linking of text to the non-linguistic world. |
| Approach: | They investigate the extent to which trajectories naturally encode verb semantics . they build a procedurally generated agent-object-interaction dataset and compare methods . |
| Outcome: | The proposed model can capture verb semantics by tracing trajectories and self-supervised pretraining. |
Copied to clipboard
| Challenge: | Long-context question answering tasks often require identifying evidence spans (e.g., sentences) prior work showed that jointly training models to perform evidence extraction and question answering is important for achieving high performance. |
| Approach: | They propose a method for equipping long-context QA models with an additional sequence-level objective for better identification of the supporting evidence. |
| Outcome: | The proposed method exhibits consistent improvements on three different strong long-context transformer models, across two challenging question answering benchmarks – HotpotQA and QAsper. |
Copied to clipboard
| Challenge: | Large clinical note corpora are one of the most needed and one of least available resources in biomedical NLP due to patient confidentiality considerations and expert annotation cost. |
| Approach: | They present a corpus of 43,985 clinical patient notes (PNs) written by 35,156 examinees during the USMLE® Step 2 Clinical Skills examination. |
| Outcome: | The corpus of 43,985 clinical patient notes (PNs) written by 35,156 examinees during the high-stakes USMLE® Step 2 Clinical Skills examination is available via a data sharing agreement with NBME . |
Copied to clipboard
| Challenge: | Existing methods to integrate text corpora with knowledge graphs (KGs) have been effective in various NLP tasks such as analyzing and predicting relationships between entities. |
| Approach: | They propose a method that borrows LDPs from entities that co-occur in sentences to represent entities that do not co-exist in a single sentence. |
| Outcome: | The proposed method improves the performance of prior methods such as TransE, DistMult, ComplEx and RotatE. |
Copied to clipboard
| Challenge: | Recent work in entity disambiguation relies on a limited subset of KB facts to link entities . less common entities are prone to missing or inconsistent KB information, which is problematic for models which rely on 'one source' |
| Approach: | They propose an ED model which links entities by reasoning over a symbolic knowledge base in a fully differentiable fashion. |
| Outcome: | The proposed model outperforms state-of-the-art models on six well-established datasets by 1.3 F1 on average. |
Copied to clipboard
| Challenge: | modal dependency parsing is a task of parse a text into its modal dependence structure . the root node of an MDS is always the author of a document, the ultimate source of information sources . |
| Approach: | They propose a modal dependency parser based on priming pre-trained language models and evaluate it on two data sets. |
| Outcome: | The proposed parser improves on two data sets. |
Copied to clipboard
| Challenge: | Document-level relation extraction models are not robust and exhibit bizarre behaviors when non-evidence sentences are removed. |
| Approach: | They propose a document-level relation extraction framework that uses a sentence importance score and a focusing loss to encourage DocRE models to focus on evidence sentences. |
| Outcome: | The proposed framework improves overall performance and makes DocRE models more robust. |
Copied to clipboard
| Challenge: | Existing benchmark datasets contribute little to discriminating top-scoring systems, while those less used datasets exhibit impressive discriminative power. |
| Approach: | They examine the distinguishability of benchmark datasets when comparing different systems . they find that existing benchmark dataset contribute little to discriminating top-scoring systems - whereas those less used datasets exhibit impressive discriminative power. |
| Outcome: | The proposed datasets are released on DataLab. |
Copied to clipboard
| Challenge: | Backdoor attacks are a new threat to neural natural language processing models due to the fragility and lack of interpretability of NLP models. |
| Approach: | They propose a method to perform backdoor attacks without an external trigger . they propose to use clean-labeled examples to generate poisoned clean-labelled examples . |
| Outcome: | The proposed strategy is effective and hard to defend due to its triggerless nature. |
Copied to clipboard
| Challenge: | Large language models (LM) based on transformers generate plausible long texts . a discriminator-guided approach allows to apply constraints more finely and dynamically. |
| Approach: | They propose to use a discriminator-guided approach to generate constrained texts without fine-tuning the LM. |
| Outcome: | The proposed method is easier and cheaper to train than fine-tuning the LM. |
Copied to clipboard
| Challenge: | Existing proof generation tasks require reasoning capabilities, but they usually just request for an answer without the reasoning procedure that would make it interpretable. |
| Approach: | They propose an iterative backward reasoning model to solve the proof generation tasks on rule-based Question Answering. |
| Outcome: | The proposed model improves in-domain performance and cross-domain transferability over existing models. |
Copied to clipboard
| Challenge: | Existing studies on domain-shifting adaptations have focused on domain . |
| Approach: | They propose a self-supervised approach to unsupervised domain adduction using domain puzzles to bridge the source and target domains and retain discriminative representations after adaptation. |
| Outcome: | The proposed approach outperforms baselines and further ablation studies show that it is more stable and effective when performing other data augmentations. |
Copied to clipboard
| Challenge: | Recent years, transformer-based coreference resolution systems have achieved remarkable improvements on the CoNLL dataset. |
| Approach: | They propose to incorporate centering transitions derived from centering theory into a neural coreference model by using a graph. |
| Outcome: | The proposed model improves on pronoun resolution in long documents, formal well-structured text, and clusters with scattered mentions. |
Copied to clipboard
| Challenge: | Recent semi-supervised learning methods have achieved impressive performance . semi-controlled learning can be used to reduce the annotation cost of text classifiers . |
| Approach: | They propose a semi-supervised learning process that builds a standard K-way classifier and a matching network for the input text and the Class Semantic Representation (CSR). |
| Outcome: | The proposed method improves baselines and overall is more stable. |
Copied to clipboard
| Challenge: | Existing unsupervised text style transfer methods suffer from performance degradation when fine-tuning the model in new domains. |
| Approach: | They propose a domain adaptive meta-learning approach with an adversarial style training approach for better content preservation and style transfer. |
| Outcome: | The proposed approach generalizes well on unseen low-resource domains against ten strong baselines. |
Copied to clipboard
| Challenge: | lexical biases in hate speech detection are limited when applied to real-world data, exhibiting limited out-of-distribution robustness and perpetuating harmful social biase. |
| Approach: | They propose to disentangle spurious and authentic artifacts and analyze their impact on out-of-distribution fairness and robustness. |
| Outcome: | The proposed models show that spurious artifacts require different treatments to attain robustness and fairness in hate speech detection. |
Copied to clipboard
| Challenge: | Document-level event argument extraction is a crucial subtask of event extraction. |
| Approach: | They propose to use redundant event information to extract multiple arguments from a document . they propose a loss function to classify Universum class by their open decision boundary . |
| Outcome: | The proposed model outperforms the previous state-of-the-art models by 3.35% in F1-score. |
Copied to clipboard
| Challenge: | Low-resource languages are left out of large-scale pretraining datasets . authors explore how to leverage existing pre-trained models to create low-resourced translation systems for 16 African languages. |
| Approach: | They investigate how large-scale pre-trained models can be used to create low-resource translation systems for 16 African languages. |
| Outcome: | The proposed models can translate between hundreds of languages even though there is little parallel data available for training. |
Copied to clipboard
| Challenge: | Existing studies rely on entity information for sentence-level relation extraction (RE) but this can leak superficial and spurious clues of relations. |
| Approach: | They propose to use entity mentions to extract relations from textual context . they use a causal graph to model dependencies between variables in RE models . |
| Outcome: | The proposed method yields significant gains on both effectiveness and generalization for RE. |
Copied to clipboard
| Challenge: | a new framework to analyze how latent concepts are encoded in representations learned in pre-trained lan-guage models is proposed . conceptX uses clustering to discover the encoded concepts and align them with a large set of human-defined concepts. |
| Approach: | They propose a framework to analyze how latent concepts are encoded in representations learned within pre-trained lan-guage models. |
| Outcome: | The proposed framework explains encoded concepts by aligning with human-defined concepts. |
Copied to clipboard
| Challenge: | DrBoost is a dense retrieval ensemble that is trained in stages to correct retrieval mistakes . it produces representations which are 4x more compact, while delivering comparable retrieval results. |
| Approach: | They propose a dense retrieval ensemble inspired by boosting that is trained in stages . they produce representations which are 4x more compact, while delivering comparable retrieval results . |
| Outcome: | The proposed model performs surprisingly well under approximate search with coarse quantization, reducing latency and bandwidth needs by another 4x. |
Copied to clipboard
| Challenge: | Using a multi-reference multi-source evaluation dataset, Chinese grammatical error correction (CGEC) is relatively scarce. |
| Approach: | They propose a multi-reference multi-source evaluation dataset for Chinese grammar error correction . the dataset contains 7,063 sentences written by Chinese-as-a-Second-Language learners . |
| Outcome: | The proposed dataset can be used to evaluate Chinese grammar errors in Chinese. |
Copied to clipboard
| Challenge: | a new task is proposed to reduce media news framing bias by generating a neutral summary from multiple news articles of the varying political leanings. |
| Approach: | They propose a task that generates a neutral summary from multiple news articles . they find title provides a good signal for framing bias and propose metric and model . |
| Outcome: | The proposed task can neutralize news content in hierarchical order from title to article . scalability remains a bottleneck due to the time-consuming human labor needed for composing the roundup . |
Copied to clipboard
| Challenge: | omitted tokens from the context contribute to incomplete utterance restoration (IUR) understanding conversational interactions through NLP has become important with increasing connectivity and range of capabilities. |
| Approach: | They propose a model for incomplete utterance restoration called JET . they construct a Picker that identifies omitted tokens and two label creation methods to support the picker. |
| Outcome: | The proposed model is better than pretrained T5 and non-generative language model methods on four benchmark datasets in extraction and abstraction scenarios. |
Copied to clipboard
| Challenge: | Semantic parsing is one of the central tasks for natural language understanding (NLU). |
| Approach: | They propose a Segmented Invocation Transformer that utilizes the information from the constituency parse tree of the natural language text and Bash command components to generate Bash commands. |
| Outcome: | The proposed method improves the inference time and reduces the model parameters by 1.8x . |
Copied to clipboard
| Challenge: | Embeddings compress information into low-dimensional vectors, but can leak private information about sensitive attributes of text. |
| Approach: | They propose a method to privatize embeddings based on homomorphic encryption to prevent leakage of sensitive information in the process of text classification. |
| Outcome: | The proposed method can protect embeddings from leakage while preserving their utility on downstream tasks. |
Copied to clipboard
| Challenge: | Recent work on Multi-modal Named Entity Recognition (MNER) relies on image information to model interactions between image and text representations. |
| Approach: | They propose to align image features into the textual space to better utilize attention mechanisms . they use regional object tags, captions and optical characters as visual contexts . |
| Outcome: | The proposed model can achieve state-of-the-art accuracy on multi-modal Named Entity Recognition datasets even without image information. |
Copied to clipboard
| Challenge: | Combination therapies are becoming standard of care for diseases such as cancer, tuberculosis, malaria and HIV. |
| Approach: | They construct an expert-annotated dataset for extracting drug combinations from the scientific literature. |
| Outcome: | The proposed dataset is the first relation extraction dataset consisting of variable-length relations. |
Copied to clipboard
| Challenge: | Existing evaluation methods do not provide insight into how well a language model captures distinct linguistic skills essential for language understanding and reasoning. |
| Approach: | They propose a new format of NLI benchmark for evaluation of broad-coverage linguistic phenomena using a set of datasets and an evaluation procedure for diagnosing how well a language model captures reasoning skills. |
| Outcome: | The proposed model can diagnose model behavior and verify model learning quality. |
Copied to clipboard
| Challenge: | Existing literature has focused on pretrainer-based text-driven brain encoding models . however, few studies have explored the efficacy of task-specific learning of Transformers . |
| Approach: | They propose to use ten popular natural language processing tasks to learn Transformer representations for predicting brain responses. |
| Outcome: | The proposed model predicts brain activity across the whole brain. |
Copied to clipboard
| Challenge: | Recent studies show that abstractive summarization approaches generate summaries that are not factually consistent with the source document. |
| Approach: | They propose a method that decomposes the document and summary into structured meaning representations (MRs) MRs describe core semantic concepts and their relations, aggregating the main content in both document and summary in a canonical form . |
| Outcome: | The proposed method outperforms existing methods on benchmarks for factuality evaluation. |
Copied to clipboard
| Challenge: | Nominalizations can be difficult to interpret because of ambiguous semantic relations between deverbal noun and its arguments. |
| Approach: | They propose to over-generate clausal paraphrases to predict whether a prenominal modifier can be re-written as a noun or adverb in a claual paraphrasability. |
| Outcome: | The proposed method improves paraphrasability prediction and paraphrase generation in English . it shows that the prenominal modifier can be re-written as a noun or adverb in a clausal paraphrase . |
Copied to clipboard
| Challenge: | Entity disambiguation (ED) is a task of assigning mentions to referent entities in a knowledge base. |
| Approach: | They propose a global entity disambiguation (ED) model based on BERT . they train the model using a large entity-annotated corpus obtained from Wikipedia . |
| Outcome: | The proposed model can disambiguate masked entities based on words and non-masked ones at the inference time. |
Copied to clipboard
| Challenge: | Multiple-choice question answering (MCQA) uses text-to-text framework . but, there is an under-utilization of the decoder and knowledge that can be decoded . |
| Approach: | They propose a generative multiple-choice question answering model which generates a clue from the question and leverages it to enhance a reader for MCQA. |
| Outcome: | The proposed model outperforms text-to-text models on multiple MCQA datasets. |
Copied to clipboard
| Challenge: | Rather than pursuing the reachless SOTA accuracy, researchers are focusing on model efficiency and usability. |
| Approach: | They propose an evaluation and a public leaderboard for efficient NLP models that depicts the Pareto Frontier for various language understanding tasks. |
| Outcome: | The proposed model outperforms or performs on par with SOTA compressed and early exiting models. |
Copied to clipboard
| Challenge: | Existing knowledge-grounded dialogue generation models only produce pedantic responses, which lacks emotion and attraction compared with the responses with polite style, positive and negative sentiments. |
| Approach: | They propose a method which generates responses via combing disentangled style templates and content templates. |
| Outcome: | The proposed method improves on evaluation metrics compared with state-of-the-art methods. |
Copied to clipboard
| Challenge: | Existing methods exploit the utterances of all dialogue turns to assign value to slots . this can lead to suboptimal results due to information introduced from irrelevant utterrances . |
| Approach: | They propose a SLot-TUrN Alignment enhanced approach to assign slot value . they explicitly align each slot with its most relevant utterance and then predict the corresponding value based on this aligned utteration. |
| Outcome: | The proposed approach achieves state-of-the-art on three multi-domain task-oriented dialogue datasets. |
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) datasets annotate coarse-grained entities such as a continent, a country, or a city. |
| Approach: | They propose a dataset HarveyNER with fine-grained locations annotated in tweets that characterizes many complex and long location mentions in informal descriptions. |
| Outcome: | The proposed dataset outperforms existing systems on hard cases and improves on the heuristic curricula. |
Copied to clipboard
| Challenge: | Multitask learning with an unbalanced data distribution skews model learning towards high resource tasks. |
| Approach: | They propose to use a temperature heating mechanism and dense pre-training to mitigate this by training models with a fixed model capacity. |
| Outcome: | The proposed techniques improve performance on two multilingual translation benchmarks compared to BASELayers and Dense scaling baselines and in combination, more than 2x model convergence speed. |
Copied to clipboard
| Challenge: | Existing methods to control the protagonist's persona in story generation are implicitly and sparsely embodied in stories, so we propose a planning-based generation model called ConPer to explicitly model the relationship between personas and events. |
| Approach: | They propose a model to control the protagonist's persona in story generation by predicting one target sentence and planning the plot as a sequence of keywords with the guidance of the predicted persona-related events and commonsense knowledge. |
| Outcome: | The proposed model outperforms state-of-the-art models for generating more coherent and persona-controllable stories. |
Copied to clipboard
| Challenge: | CHEF dataset provides evidence retrieval over non-English claims . e-fact-checking is a time-consuming task, which can take journalists several hours or days. |
| Approach: | They construct a dataset of 10K real-world claims that is based on annotated evidence retrieved from the Internet. |
| Outcome: | The proposed dataset provides evidence retrieval as a latent variable and can be used to train and reason over non-English claims. |
Copied to clipboard
| Challenge: | Neural module networks (NMN) have been used in image-grounded tasks such as Visual Question Answering (VQA) however, very limited work on NMN has been studied in the video-ground dialogue tasks. |
| Approach: | They propose to use video as the grounding feature in video-grounded dialogues to model the information retrieval process in videogrounded language tasks as a pipeline of neural modules. |
| Outcome: | The proposed model can achieve promising performance on video-grounded dialogue and QA benchmarks. |
Copied to clipboard
| Challenge: | Dialogue state tracking is a key component of dialogue systems. |
| Approach: | They propose to extend the definition of dialogue state tracking to multimodality . they propose a new synthetic benchmark and a novel baseline for this task . |
| Outcome: | The proposed task is based on a synthetic benchmark and a self-supervised video understanding task. |
Copied to clipboard
| Challenge: | Pre-trained models have not been used to outperform other deep learning models such as CNN in Automated Essay Scoring (AES). |
| Approach: | They propose a novel multi-scale essay representation for BERT that can be jointly learned . they employ multiple losses and transfer learning from out-of-domain essays to further improve performance . |
| Outcome: | The proposed model outperforms existing models in the area of automated essay scoring . the proposed model generalizes well to the CommonLit Readability Prize data set . |
Copied to clipboard
| Challenge: | a new benchmark evaluates coreference resolution systems' ability to recognize singular personal "they" we find that current systems overwhelmingly choose to resolve "they's" correctly to a singular entity or to 'a group' |
| Approach: | They propose to evaluate coreference resolution systems for singular personal "they" they use WinoNB schemas to evaluate whether they can correctly resolve singular "they". |
| Outcome: | The proposed benchmark evaluates coreference resolution systems for singular personal "they" they show that they are biased toward resolving "they", not "them" |
Copied to clipboard
| Challenge: | Recent studies on propaganda detection involve document and fragment-level analyses of news articles. |
| Approach: | They propose a neural approach to detect and categorize propaganda tweets across fine-grained categories . they use a dataset containing tweets weakly annotated with different propaganda techniques . |
| Outcome: | The proposed method outperforms benchmark methods and transfers knowledge to low-resource news domains. |
Copied to clipboard
| Challenge: | Currently, global models are not able to produce personalized responses for individual users, based on their data. |
| Approach: | They propose a scheme for training a single shared model for all users by prepending a fixed, user-specific non-trainable string to each user’s input text. |
| Outcome: | The proposed method outperforms the state-of-the-art model on a suite of sentiment analysis datasets by up to 13 points. |
Copied to clipboard
| Challenge: | Conventional exact or approximate termbased retrieval methods lack the ability of semantic understanding of the clinical as well as language context. |
| Approach: | They combine clinical finding detection with supervised query match learning to train a model . findings are used as queries to train the Sentence-BERT model using triplet loss . |
| Outcome: | The proposed method outperforms existing methods on multiple retrieval benchmarks. |
Copied to clipboard
| Challenge: | Recent work has demonstrated that image captioning is a complex task that requires a large amount of human input. |
| Approach: | They develop a human evaluation protocol for image captioning models based on machine- and human-generated captions on the MSCOCO dataset. |
| Outcome: | The proposed model improves CLIPScore, a recent metric that uses image features, and improves human judgments because it is more sensitive to recall. |
Copied to clipboard
| Challenge: | Recent work on multilingual pre-trained models has focused on pre-training transformers on concatenated corpora of a large number of languages. |
| Approach: | They propose a language-specific module approach that allows for more languages to be trained post-hoc. |
| Outcome: | The proposed model can be pre-trained on multiple languages with no drop in performance . |
Copied to clipboard
| Challenge: | Abstract Meaning Representation (AMR) graphs are compared to gold graphs by the Smatch metric, but lack a well-defined representation and evaluation. |
| Approach: | They propose an algorithm for deriving a unified graph representation using a super-sentential annotation method. |
| Outcome: | The proposed algorithm avoids the pitfalls of over-merging and lacks coherence from under merging. |
Copied to clipboard
| Challenge: | Pretrained language models (PLMs) have made remarkable progress in text generation tasks via fine-tuning. |
| Approach: | They propose a prompt-based method that learns source prompts and transfers them as target prompts to perform target generation tasks. |
| Outcome: | The proposed method can be used to perform text generation tasks in a transferable setting. |
Copied to clipboard
| Challenge: | Recent years have featured a trend towards Transformer based pretrained language models (PLMs) in natural language processing systems. |
| Approach: | They propose to use four evaluation dimensions to evaluate ten widely-used PLMs . they find that pretrained language models are good at different ability tests . |
| Outcome: | The results show that pretrained language models are good at different ability tests and have excellent transferability between tasks. |
Copied to clipboard
| Challenge: | Recent advances on models and metrics should benefit and inform each other, authors argue . bidimensional leaderboards allow for fast, accurate evaluation of language generation models . |
| Approach: | They propose a bidimensional leaderboard that tracks progress in language generation models and metrics for their evaluation. |
| Outcome: | The proposed leaderboards track progress in language generation models and metrics for their evaluation. |
Copied to clipboard
| Challenge: | Existing approaches to improve in-context few-shot learning are pretraining and downstream fewshot evaluation. |
| Approach: | They propose to use self-supervision as an intermediate training stage between pretraining and downstream fewshot usage to train models to perform in-context few shot learning. |
| Outcome: | The proposed model outperforms baseline models on two benchmarks. |
Copied to clipboard
| Challenge: | Recent video-text models can retrieve relevant videos based on text with high accuracy, but to what extent do they comprehend the semantics of the text? |
| Approach: | They propose a framework that probes video-text models with hard negatives . they leverage a pre-trained language model and a set of heuristics to create verb and person entity focused contrast sets. |
| Outcome: | The proposed framework erases the performance gap between CLIP-based methods and the earlier methods. |
Copied to clipboard
| Challenge: | a sonnet is a fourteen-line poem with rigorous meter-and-rhyme constraints. |
| Approach: | They propose a framework which plans the poem sketch before decoding a sonnet without training on poems . they use a rhyme module, polishing module and a constrained decoding algorithm to impose the meter-and-rhyme constraint . |
| Outcome: | The proposed framework generates sonnets that are coherent and poetic without training on poems . the proposed framework is based on a framework that plans the poem sketch before decoding . |
Copied to clipboard
| Challenge: | Recent work on fairness of machine learning models has focused on how to debias, but research on the fairness and performance of biased/debiased models on downstream prediction tasks has been limited. |
| Approach: | They assess intersectional bias - fairness across multiple demographic dimensions . they highlight possible causes and make recommendations for future NLP debiasing research. |
| Outcome: | The proposed approaches fare well in terms of fairness-accuracy trade-off, but are unable to effectively alleviate bias in downstream tasks. |
Copied to clipboard
| Challenge: | Recent work on multilingual language models has demonstrated their capacity for cross-lingual zero-shot transfer on downstream tasks. |
| Approach: | They conduct a large-scale empirical study to isolate the effects of various linguistic properties by measuring zero-shot transfer between four different natural languages. |
| Outcome: | The proposed model exhibits decent cross-lingual zero-shot transfer, with no significant differences in word order and embedding alignment. |
Copied to clipboard
| Challenge: | a recent study shows that gender-neutral pronouns are not associated with processing difficulties . linguistic scholars have observed how technology has altered the course of language evolution . |
| Approach: | They show that gender-neutral pronouns in Danish, English and Swedish are not associated with processing difficulties. |
| Outcome: | a new study shows that gender-neutral pronouns are not associated with human processing difficulties . the findings suggest that such conservativity in language models may limit widespread adoption . |
Copied to clipboard
| Challenge: | Recent work shows the surprising power of continuous prompts to language models for controlled generation and solving a wide range of tasks. |
| Approach: | They propose to extract a discrete (textual) interpretation of continuous prompts faithful to the problem they solve. |
| Outcome: | The proposed model can find prompts that solve a task while being projected to an arbitrary text with a smaller drop in accuracy. |
Copied to clipboard
| Challenge: | Identifying related entities and events within and across documents is fundamental to natural language understanding. |
| Approach: | They propose an approach to entity and event coreference resolution using contrastive representation learning. |
| Outcome: | The proposed method achieves state-of-the-art results on key metrics on the ECB+ corpus and is competitive on others. |
Copied to clipboard
| Challenge: | phonological hierarchies that predict coordinate constructions are often phonetically “natural” . a neural sequence labeling model can learn elaborate expressions in Hmong without using phonology information. |
| Approach: | They propose that coordinate compounds and elaborate expressions can be learned empirically by phonological hierarchies and a neural sequence labeling model can learn the ordering of elaborate expression in Hmong without using phonology. |
| Outcome: | The proposed models beat strong baselines for all three languages and learn hierarchies similar to those proposed by Mortensen. |
Copied to clipboard
| Challenge: | Existing work on generating edits grounded in external knowledge has focused on correcting grammar and reducing repetitive typing. |
| Approach: | They propose a novel task where the goal is to update an existing article given new evidence by using a dataset of 170K distantly supervised data produced from Wikipedia snapshots. |
| Outcome: | The proposed model can update Wikipedia articles faithfully with new capabilities and opens doors to many new applications. |
Copied to clipboard
| Challenge: | Task-oriented dialog (TOD) is arguably one of the most popular natural language processing (NLP) application areas. |
| Approach: | They propose a multilingual multi-domain TOD dataset that spans four languages . they use a framework for multilingual conversational specialization of pretrained language models . |
| Outcome: | The proposed datasets show that they perform better than existing datasets in English . the proposed framework allows for sample-efficient few-shot transfer for TOD tasks . |
Copied to clipboard
| Challenge: | Existing long-range language models lack a meaningful evaluation of their discourse-level language understanding capabilities. |
| Approach: | They propose a dataset that provides an LRLM with a long segment from a narrative that ends at a chapter boundary and asks it to distinguish the beginning of the ground-truth next chapter from n-token segments. |
| Outcome: | The proposed dataset shows that existing models fail to leverage long-range context . |
Copied to clipboard
| Challenge: | Neural information retrieval (IR) methods encode queries and documents into single vectors, but late interaction models produce multi-vector representations at the granularity of each token. |
| Approach: | They propose a retrieval method that couples an aggressive residual compression mechanism with a denoised supervision strategy to improve the quality and space footprint of late interaction. |
| Outcome: | The proposed retriever improves quality and space footprint of late interaction models while reducing space footprint by 6–10x. |
Copied to clipboard
| Challenge: | acoustic models represent linguistic information based on massive amounts of data. |
| Approach: | They examine the model's ability to distinguish low-resource (Dutch) regional varieties by extracting embeddings from hidden layers and dynamic time warping. |
| Outcome: | The proposed model outperforms transcription-based models without phonetic transcriptions on the basis of only six seconds of speech. |
Copied to clipboard
| Challenge: | Existing work uses the same adapter architecture for every dataset regardless of the properties of the dataset or the amount of training data. |
| Approach: | They propose to use adaptable adapters to finetune lightweight neural network layers on top of pretrained weights. |
| Outcome: | The proposed adapters achieve on-par performances with the standard adapter architecture while using a considerably smaller number of adapter layers. |
Copied to clipboard
| Challenge: | Dynamic Adversarial Data Collection (DADC) is a time-consuming and costly approach . DADC is based on training data collected from adversarial and out-of-domain settings . |
| Approach: | They propose a dynamic data collection approach that uses generator-in-the-loop models to provide real-time suggestions that annotators can approve, modify, or reject. |
| Outcome: | The proposed model is more robust in adversarial and out-of-domain settings and harder for humans to fool. |
Copied to clipboard
| Challenge: | Document Information Extraction (DIE) has attracted increasing attention due to its various advanced applications in the real world. |
| Approach: | They propose a multi-modal generation method without predefined label categories for real-world scenarios using a spatial encoder and modal-aware mask module. |
| Outcome: | The proposed method can deal with complex documents that are hard to serialize into sequential order. |
Copied to clipboard
| Challenge: | Existing non-autoregressive neural machine translation models suffer from multimodality problem . multi-modality is not solved by a teacher forcing algorithm, limiting model capability . |
| Approach: | They propose a method that generates multiple reference translations for each source sentence . they compare the NAT output with all references and select the one that best fits the simulated model . |
| Outcome: | The proposed method achieves 29.82 BLEU with only one decoding pass on WMT14 En-De . |
Copied to clipboard
| Challenge: | Existing models that generate rationales before making predictions can ignore noise or adversarially added text by simply masking it out of the generated rationale. |
| Approach: | They propose to use a 'rationalizethen-predict' framework to generate subsets of input to generate rationales and then make predictions using them. |
| Outcome: | The proposed models improve robustness over AddText attacks while struggling in certain scenarios. |
Copied to clipboard
| Challenge: | Recent studies on few-shot intent detection have attempted to formulate the task as a meta-learning problem. |
| Approach: | They propose to modify a few-shot intent detection task to produce a non-trivially strong performance without further domain-specific adaptation. |
| Outcome: | The proposed model improves on the prototypical network variants with task-specific fine-tuning. |
Copied to clipboard
| Challenge: | Recent advances in self-supervised modeling of text and images open new opportunities for computational models of child language acquisition. |
| Approach: | They propose a multimodal language acquisition model trained from image-caption pairs on naturalistic data using cross-modal self-supervision. |
| Outcome: | The proposed model learns word categories and object recognition abilities, the authors show . their model is trained from image-caption pairs on naturalistic data using cross-modal self-supervision . |
Copied to clipboard
| Challenge: | Existing approaches to detect adversarial examples for deep learning based systems focus on image embedding feature spaces . however, existing approaches focus on text features, without considering model embeddable spaces. |
| Approach: | They propose a sentence-embedding “residue” detector to identify adversarial examples from embedded feature spaces. |
| Outcome: | The proposed detector outperforms existing model-focused detectors on many tasks. |
Copied to clipboard
| Challenge: | Existing extraction models memorize and recall already seen triples but cannot generalize effectively for unseen triples. |
| Approach: | They propose a method to generalize existing extraction models by rearranging datasets and augmenting test sets. |
| Outcome: | The proposed method can significantly increase the generalization performance of existing models. |
Copied to clipboard
| Challenge: | Existing models focus more on the structure of summary, not on the personal and logical inconsistency problem. |
| Approach: | They propose a model to solve the problem of personal and logical inconsistency . they use an utterance rewriter to complete the ellipsis content of dialogue content . |
| Outcome: | The proposed model outperforms baseline models on both SAMSum and DialSum datasets. |
Copied to clipboard
| Challenge: | Existing methods for learning sentence embeddings are fine-tuning general-purpose pretrained models with a particular training supervision. |
| Approach: | They propose a method for learning sentence embeddings via contrastive learning between sentences and related entities. |
| Outcome: | The proposed method outperforms baseline methods in multilingual settings on a variety of tasks. |
Copied to clipboard
| Challenge: | Recent work incorporates pre-trained word embeddings into Neural Topic Models (NTMs), generating highly coherent topics. |
| Approach: | They conduct thorough experiments to investigate whether embeddings directly with an appropriate word selection method can generate more coherent and diverse topics than NTMs. |
| Outcome: | The proposed model generates more coherent and diverse topics than traditional NTMs, achieving higher efficiency and simplicity. |
Copied to clipboard
| Challenge: | Existing video QA models lack the capacity for deep video understanding and flexible multistep reasoning. |
| Approach: | They propose a video question answering model which performs dynamic multistep reasoning between questions and videos. |
| Outcome: | The proposed model improves on three widely used video QA datasets and displays better interpretability by backtracing along with the attention mechanisms to the video scene graphs. |
Copied to clipboard
| Challenge: | Grounded text generation systems often generate factual inconsistencies, hindering their real-world applicability. |
| Approach: | They propose a method to assess factual consistency metrics on standardized texts . they recommend NLI and question generation-and-answering-based methods as starting points . |
| Outcome: | The proposed method is more actionable and interpretable than previous methods. |
Copied to clipboard
| Challenge: | Existing large-scale pre-trained language models are mainly trained from scratch individually, ignoring that many well-taught PLMs are available. |
| Approach: | They propose a pre-training framework called knowledge inheritance and propose auxiliary supervision to efficiently learn larger PLMs. |
| Outcome: | The proposed framework can be used to train large-scale language models with huge parameters and a large dataset can be adapted to domain adaptation and knowledge transfer. |
Copied to clipboard
| Challenge: | BLEU scores of 31.16 for ende and 38.37 for deen on the IWSLT14 dataset, 30.78 for entde, 35.15 for de en and 27.17 for zhen . |
| Approach: | They propose a bidirectional pretraining and unidirectional finetuning procedure to boost NMT performance. |
| Outcome: | The proposed method achieves strong translation performance across five datasets. |
Copied to clipboard
| Challenge: | Pre-trained language models (PLMs) can achieve comparable performance to full-parameter fine-tuning by tuning a few soft prompts, but require much more training time than fine-timing. |
| Approach: | They empirically investigate the transferability of soft prompts across different downstream tasks and PLMs to determine what decides prompt transferability. |
| Outcome: | The proposed method can achieve comparable performance to full-parameter fine-tuning by tuning a few soft prompts, but requires much more training time than fine-timing. |
Copied to clipboard
| Challenge: | Existing datasets focus on sentence-level event extraction, but document-level EE is limited due to the lack of large-scale and practical training and evaluation datasets. |
| Approach: | They propose a document-level event extraction dataset with 27,000+ events and 180,000+ arguments. |
| Outcome: | The proposed dataset includes 27,000+ events, 180,000+ arguments and large-scale manual annotations, fine-grained argument types and application-oriented settings. |
Copied to clipboard
| Challenge: | Existing studies show translation artifacts in translations influence performance of cross-lingual tasks. |
| Approach: | They propose a method to reduce translation artifacts by extending an established bias-removal technique. |
| Outcome: | The proposed method reduces translationese at sentence and word level . it is the first study to debias translations on a natural language inference task . |
Copied to clipboard
| Challenge: | Existing methods to train large pretrained language models require more computational resources and are expensive to train in other languages. |
| Approach: | They propose a method to transfer pretrained language models to new languages using subword-based tokenization and embeddings. |
| Outcome: | The proposed method outperforms existing methods on low-resource languages and makes training large models more accessible and less damaging to the environment. |
Copied to clipboard
| Challenge: | Existing knowledge based question answering systems are trained based on labeled reasoning paths, which hinder their performance. |
| Approach: | They propose a KBQA system which leverages multiple reasoning paths’ information and only requires labeled answer as supervision. |
| Outcome: | The proposed system can leverage multiple reasoning paths’ information and only requires labeled answer as supervision. |
Copied to clipboard
| Challenge: | Existing studies on Tabular Natural Language Inference (TNLI) focus on monolingual settings where tabular premise and hypothesis are in the same language. |
| Approach: | They propose a task where tabular premise and hypothesis are in two languages . they translate textual hypotheses from an English-indic TNLI dataset into eleven major languages - english and indic . |
| Outcome: | The proposed model performs well on a bilingual dataset in English and in 11 major Indian languages. |
Copied to clipboard
| Challenge: | Generative methods for biomedical entity linking (EL) use synonyms knowledge from knowledge bases (KB) this is not trivial to inject into a generative method, but it is cost-effective. |
| Approach: | They propose to inject synonyms knowledge into a generative model of biomedical EL by constructing synthetic samples with synonyms and definitions from KB and requiring the model to recover concept names. |
| Outcome: | The proposed method achieves state-of-the-art results on several biomedical EL tasks without candidate selection. |
Copied to clipboard
| Challenge: | Prior research has focused on reducing noise for specific methods to achieve an effective integration. |
| Approach: | They propose to use token substitution and mixup to improve named entity recognition (NER) using a meta-reweighting strategy, which is extensible and requires little effort. |
| Outcome: | The proposed method is extensible, imposing little effort on a specific self-augmentation method. |
Copied to clipboard
| Challenge: | Low-resource languages lack annotated data even for basic syntactic information such as parts of speech. |
| Approach: | They propose an unsupervised cross-lingual approach for POS tagging for low-resource languages of rich morphology . they further investigate morpheme-level alignment and projection and use of linguistic priors for morphological segmentation . |
| Outcome: | The proposed approach outperforms the word-based approach and outperfies word-driven approaches. |
Copied to clipboard
| Challenge: | Existing methods to reduce bias have been shown to be effective over real-world datasets. |
| Approach: | They propose two new training objectives which directly optimise for the widely-used criterion of equal opportunity. |
| Outcome: | The proposed training objectives directly optimise for the widely-used criterion of equal opportunity while maintaining high performance over two classification tasks. |
Copied to clipboard
| Challenge: | Existing text-image approaches use pre-trained vision-language representations for text retrieval . however, these models pose non-trivial memory requirements and substantial indexing time . |
| Approach: | They propose a framework to compress large pre-trained dual-encoders for lightweight text-image retrieval. |
| Outcome: | The proposed model performs better on Flickr30K and MSCOCO benchmarks than the original full model on mobile devices. |
Copied to clipboard
| Challenge: | Existing studies on timeline summarization ignore the information interaction between sentences and dates, and combine them as two separate tasks. |
| Approach: | They propose a joint learning-based heterogeneous graph attention network for timeline summarization (HeterTls) they combine date selection and event detection into a unified framework to improve extraction accuracy . |
| Outcome: | The proposed model outperforms state-of-the-art models on four datasets . it significantly outperformed the baseline models on ROUGE scores and date selection metrics . |
Copied to clipboard
| Challenge: | rumor detection models have been designed with oversimplifcation and evaluated inappropriately on a few datasets where the actual early-stage information is largely missing. |
| Approach: | They propose a new Benchmark dataset for EArly Rumor Detection based on claims from fact-checking websites and a novel model based upon neural Hawkes process for EARD. |
| Outcome: | The proposed model can guide a generic rumor detection model to make timely, accurate and stable predictions. |
Copied to clipboard
| Challenge: | Existing approaches for recognizing feature transitions between utterances extract features for the context at the coarse-grained level. |
| Approach: | They propose a method to recognize feature transitions between utterances that helps understand dialogue flow . they propose empathetic response generation strategy to focus on emotion and keywords related to appropriate features when generating responses. |
| Outcome: | The proposed approach outperforms baseline approaches and improves on multi-turn dialogues. |
Copied to clipboard
| Challenge: | Existing approaches focus on leveraging textual content to identify stances, while they fail to reason with background knowledge or leverage the rich semantic and syntactic textual labels in news articles. |
| Approach: | They propose a political perspective detection approach that leverages news text to enable multi-hop knowledge reasoning and incorporates textual cues as paragraph-level labels. |
| Outcome: | The proposed approach outperforms state-of-the-art methods on two benchmark datasets. |
Copied to clipboard
| Challenge: | Existing approaches to improve IR labels are incomplete and require computational overheads. |
| Approach: | They propose to distill knowledge for informed labeling without high computation overheads at evaluation time. |
| Outcome: | The proposed model outperforms state-of-the-art models while distilling the rankings better. |
Copied to clipboard
| Challenge: | During a conversation, a person’s emotions are influenced by the other speaker’s utterances and their own emotional state over the utterrances. |
| Approach: | They propose a Graph Neural Network based Multi-modal Emotion recognitioN system that leverages local and global information in a conversation. |
| Outcome: | The proposed system gives state-of-the-art results on IEMOCAP and MOSEI datasets and detailed ablation experiments show the importance of modeling information at both levels. |
Copied to clipboard
| Challenge: | Existing methods for OOD detection are based on labeled in-domain data . detecting out-of-domain (OOD) or unknown intents is challenging . |
| Approach: | They propose a novel reassigned contrastive learning method to discriminate IND intents for over-confident OOD and an adaptive class-dependent local threshold mechanism to separate similar IND and OOD intents. |
| Outcome: | The proposed method is effective for both aspects of overconfidence issues. |
Copied to clipboard
| Challenge: | Existing approaches to zero/few-shot slot filling focus on slot descriptions and examples . AISFG model is based on domain-specific labels, which is not capable of transferring to new domains with little or no data. |
| Approach: | They propose a model with a query template that incorporates domain descriptions, slot descriptions, and examples with context. |
| Outcome: | Experimental results show that the proposed model outperforms state-of-the-art approaches in zero/few-shot slot filling task. |
Copied to clipboard
| Challenge: | Negation is a common linguistic feature that is crucial in many language understanding tasks. |
| Approach: | They propose a new approach to detect negation in language models using data augmentation and negation masking. |
| Outcome: | The proposed approach improves negation detection performance and generalizability over the strong baseline NegBERT. |
Copied to clipboard
| Challenge: | Existing Math Word Problem solvers do not generalize well and rely on superficial cues to achieve high performance. |
| Approach: | They propose several data augmentation techniques to increase the size of existing MWP datasets by five folds by deploying them to a benchmark dataset. |
| Outcome: | The proposed methods increase the generalization and robustness of existing solvers by over five percentage points on benchmark datasets. |
Copied to clipboard
| Challenge: | Recent work shows that finetuning pretrained models with contrastive learning makes it possible to learn good sentence embeddings without labeled data. |
| Approach: | They propose an unsupervised contrastive learning framework for learning sentence embeddings . they use a masked language model to mask out the edited sentence . |
| Outcome: | The proposed framework outperforms SimCSE on semantic textual similarity tasks by 2.3 absolute points. |
Copied to clipboard
| Challenge: | Existing approaches to perform aspect and opinion co-extraction are difficult due to the lack of fine-grained annotations. |
| Approach: | They propose a framework to transfer knowledge from a labeled source domain to an unlabeled target domain. |
| Outcome: | The proposed framework is more effective than previous domain adaptation methods on three datasets. |
Copied to clipboard
| Challenge: | Existing QA research on question answering is focused on specific question types, knowledge domains, or reasoning skills. |
| Approach: | They propose a unified QA paradigm that solves various tasks through a single model. |
| Outcome: | The proposed model improves QA-centric ability on 11 QA benchmarks. |
Copied to clipboard
| Challenge: | Using MixUp, additional samples are generated during training by combining random pairs of training samples and their labels. |
| Approach: | They propose a new MixUp strategy that leverages Training Dynamics and allows more informative samples to be combined for generating new data samples. |
| Outcome: | The proposed method achieves competitive performance using a smaller subset of training data compared with strong baselines and yields lower expected calibration error on the pre-trained language model, BERT, on both in-domain and out-of-domain settings. |
Copied to clipboard
| Challenge: | Grapheme-to-phoneme conversion is a task of converting grapheme sequences into phoneme sequence. |
| Approach: | They propose a Thai grapheme-to-phoneme conversion method that uses neural networks to predict the similarity between a candidate and the correct pronunciation. |
| Outcome: | The proposed method can be applied to other languages than Thai . it is comparable to encoder-decoder models in accuracy and accuracy, it shows . |
Copied to clipboard
| Challenge: | Existing approaches to generate adversarial examples for NMT use the meaning-preserving restriction. |
| Approach: | They propose a new definition for adversarial examples based on the Doubly Round-Trip Translation (DRTT) they introduce masked language models to construct bilingual adversarials based upon DRTT . |
| Outcome: | The proposed approach significantly improves the robustness of the NMT model on clean and noisy test sets. |
Copied to clipboard
| Challenge: | Our proposed task, TVShowGuess, builds on the scripts of TV series and takes the form of guessing the anonymous main characters based on the backgrounds of the scenes and the dialogues. |
| Approach: | They propose a task that takes the form of guessing the anonymous main characters based on the backgrounds of the scenes and the dialogues. |
| Outcome: | The proposed models outperform baselines, yet lag behind human performance. |
Copied to clipboard
| Challenge: | Distillation efforts have led to language models that are more compact and efficient without serious drops in performance. |
| Approach: | They propose to augment distillation with a third objective that encourages the student model to imitate the causal dynamics of the teacher through a distillation interchange intervention training objective (DIITO). |
| Outcome: | The proposed method lowers perplexity on the WikiText-103M corpus and improves on the GLUE benchmark, SQuAD, and CoNLL-2003. |
Copied to clipboard
| Challenge: | Using simple linear transformations, Transformer encoders can be sped up with limited accuracy costs by replacing the self-attention sublayers with simple linear mixing mechanisms. |
| Approach: | They propose to replace the self-attention sublayer with a linear transformation that "mixes" input tokens. |
| Outcome: | The proposed model outperforms the “efficient Transformers” on the GLUE benchmark at longer input lengths and on smaller models with a light memory footprint. |
Copied to clipboard
| Challenge: | Current question answering systems assume each question to have one correct answer. |
| Approach: | They propose a problem where answers are partitioned into multiple groups . they construct a comprehensive and non-redundant set of answers by picking one answer from each group . |
| Outcome: | The proposed model performs better than previous models, but it needs further improvements. |
Copied to clipboard
| Challenge: | Spurious correlations are a threat to the trustworthiness of natural language processing systems. |
| Approach: | They propose a definition of spurious correlations in terms of conditional probabilities and a generalized definition of the term . they propose UIs that allow individual input features to be independent of labels. |
| Outcome: | The proposed definition can be generalized from uniformity to independence without affecting the claims of the paper. |
Copied to clipboard
| Challenge: | Existing speaker-follower models are follower-agnostic and fail to take state of follower into account. |
| Approach: | They propose a speaker-follower model that is constantly updated given follower feedback . they optimize the speaker and obtain its training signals by evaluating the follower on labeled data . |
| Outcome: | The proposed model outperforms strong baseline models on room-to-room and room-across-room datasets. |
Copied to clipboard
| Challenge: | Generic unstructured neural networks struggle on out-of-distribution compositional generalization. |
| Approach: | They propose a method to recombinate examples from a model called Compositional Structure Learner and add them to a pre-trained sequence-to-sequence model. |
| Outcome: | The proposed model is even stronger than a T5-CSL ensemble on two real world compositional generalization tasks. |
Copied to clipboard
| Challenge: | Existing models that perform information extraction tasks manually assume heuristic dependency between the task instances and mean-field factorization for the joint distribution of instance labels. |
| Approach: | They propose to induce a dependency graph among task instances to boost representation learning by estimating their joint distribution via Conditional Random Fields. |
| Outcome: | The proposed model outperforms previous models on multiple IE tasks across 5 datasets and 2 languages. |
Copied to clipboard
| Challenge: | Existing models of language understanding are based on explicit representations of hierarchical structure, but there are good reasons to doubt that they can be said to understand language in any meaningful way. |
| Approach: | They examine whether syntactic and semantic graph representations can complement and improve neural language modeling. |
| Outcome: | The proposed model outperforms pretrained models on English WSJ in perplexity and other metrics. |
Copied to clipboard
| Challenge: | Existing methods for Natural Language Understanding focus on textual signals, which hinders models from learning efficiently from limited data samples. |
| Approach: | They propose an Imagination-Augmented Cross-modal Encoder to solve natural language understanding tasks from a novel learning perspective. |
| Outcome: | The proposed learning paradigm bridges the gap between human and agent language understanding in both linguistic and perceptual procedures. |
Copied to clipboard
| Challenge: | linguists J.R. Firth and Zellig Harris are often credited with the invention of "distributional semantics" a close reading of their work uncovers two distinct and in many ways divergent theories of meaning . |
| Approach: | They propose to compare two different theories of meaning that focus on internal workings of linguistic forms with a broader cultural and situational context. |
| Outcome: | The authors examine the differences between their theories of meaning and the internal workings of linguistic forms . they find that Firth could guide the field towards a more culturally grounded notion of semantics . |
Copied to clipboard
| Challenge: | Task-oriented parsing (TOP) aims to convert natural language into machine-readable representations of specific tasks, such as setting an alarm. |
| Approach: | They propose to reduce TOP to abstractive question answering by using canonical paraphrasing to generate linearized parse trees. |
| Outcome: | The proposed technique outperforms state-of-the-art methods in full-data settings while achieving dramatic improvements in few-shot settings. |
Copied to clipboard
| Challenge: | DR.DECR is a cross-lingual information retrieval system trained using multi-stage knowledge distillation (KD) DRDECR demonstrates superior accuracy over direct fine-tuning with labeled CLIR data. |
| Approach: | They propose a cross-lingual information retrieval system with multi-stage knowledge distillation . they teach powerful multilingual representations and CLIR by optimizing two corresponding KD objectives . |
| Outcome: | The proposed system is the best single-model retriever on the XOR-TyDi benchmark . it is based on a multi-stage knowledge distillation process that can be expensive . |
Copied to clipboard
| Challenge: | Existing work on figurative language has not been done on literal language models. |
| Approach: | They propose a Winograd-style task to evaluate figurative phrases with divergent meanings by interpreting paired figurativ phrases with a human input. |
| Outcome: | The proposed task outperforms state-of-the-art models on a nonliteral language understanding task in zero-shot settings. |
Copied to clipboard
| Challenge: | Using co-citations, we can train a model that matches aspects of papers to document level similarity. |
| Approach: | They propose a model that matches fine-grained aspects of papers and aggregates them into a document level similarity model using a naturally-occurring source of supervision: co-citations. |
| Outcome: | The proposed model improves performance on document similarity tasks in four datasets and achieves competitive results. |
Copied to clipboard
| Challenge: | Existing approaches to training dialogue agents are supervised learning, but this is prohibitively expensive and time-consuming. |
| Approach: | They propose offline reinforcement learning methods that can be used to train dialogue agents . offline reinforcement learn methods can be combined with language models to yield realistic dialogue agents. |
| Outcome: | The proposed method can be combined with language models to produce realistic dialogue agents . the results show that the offline method can achieve the goal of the proposed system . |
Copied to clipboard
| Challenge: | Existing methods for learning audio-text connections rely on parallel audio- text data . a new approach allows for the representation of environmental soundscapes without using parallel data - a challenge for many applications . |
| Approach: | They propose a model that induces Audio-Text alignment without using parallel audio-text data. |
| Outcome: | The proposed model outperforms the current state-of-the-art for audio classification tasks with no audio-text data by 2.2% on the ESC50 and US8K tasks. |
Copied to clipboard
| Challenge: | Reinforcement Learning (RL) is dependent on the reward formulation due to the intrinsic difficulty of the task in the high-dimensional discrete action space and the sparseness of the standard reward functions. |
| Approach: | They propose a maximally dense semantic-level unsupervised reward function which mimics human evaluation by considering both sentence fluency and semantic similarity. |
| Outcome: | The proposed reward outperforms the standard sparse reward by 2% on average for in- and out-of-domain settings. |
Copied to clipboard
| Challenge: | Recent work on the emergence of language between artificial agents has not isolated the effect of categorization power on inter-communication ability. |
| Approach: | They propose to use disentangled representations to quantify categorization power of agents to enable differential analysis between combinations of heterogeneous systems. |
| Outcome: | The proposed method reduces signaling accuracy by 40% despite encouraging compositionality in the artificial language. |
Copied to clipboard
| Challenge: | Recent work has leveraged natural language descriptions of schema elements to enable universal dialogue systems; however, descriptions only indirectly convey schema semantics. |
| Approach: | They propose to use schema-guided modeling to prompt seq2seq models with a labeled example dialogue to show schema semantics rather than tell them. |
| Outcome: | The proposed model outperforms models using short examples as schema representations on two popular dialogue state tracking benchmarks. |
Copied to clipboard
| Challenge: | Existing evidence suggests that pre-trained Transformers encode commonsense knowledge . however, the extent to which this knowledge is acquired is unclear . |
| Approach: | They inject verbalized knowledge into pre-training minibatches and evaluate generalization . they find generalization does not improve over the course of pre- training from scratch . |
| Outcome: | The proposed model generalizes to supported inferences after pre-training on the injected knowledge. |
Copied to clipboard
| Challenge: | Previously, paraphrases have been used to probe whether compositionality is accurately captured by BERT, but we believe they can be used to explore many other questions. |
| Approach: | They propose to use paraphrases as a unique source of data to analyze contextualized embeddings, with a particular focus on BERT. |
| Outcome: | The proposed analysis of paraphrases and paraphrase representations using the Paraphrase Database shows that BERT handles polysemous words, but different representations in many cases. |
Copied to clipboard
| Challenge: | Despite the performance gains, NLP models are still fragile and brittle to out-of-domain data, adversarial attacks, or small perturbation to the input. |
| Approach: | They propose a survey of how to define, measure and improve robustness in NLP by connecting multiple definitions of robustness and identifying failures. |
| Outcome: | The proposed models are robust against unseen or challenging scenarios, but are still fragile and brittle to out-of-domain data and adversarial attacks. |
Copied to clipboard
| Challenge: | Recent data augmentation techniques can help to deal with low resource settings, such as BERT, but they can hurt the results. |
| Approach: | They propose a neural approach to automatically learn to generate new examples using a pre-trained sequence-to-sequence model. |
| Outcome: | The proposed approach outperforms existing methods on text classification and natural language inference tasks by 10%. |
Copied to clipboard
| Challenge: | Prior studies suggested pre-trained language models possess limited understanding of commonsense knowledge despite otherwise stellar performance on leaderboards. |
| Approach: | They propose a framework that uses larger models to teach smaller models by distilling knowledge symbolically as text in addition to the neural model. |
| Outcome: | The proposed framework is based on a general language model teacher's commonsense knowledge graphs and a neural commonsensing model surpassing the teacher model's in all three criteria. |
Copied to clipboard
| Challenge: | Existing approaches to open information extraction only work with unrealistically small numbers of entities and relations. |
| Approach: | They propose to use a transformer encoder-decoder model to extract triplets from unstructured text . they use 'generative information extraction' to generate triplet representations of information . |
| Outcome: | The proposed model is state-of-the-art on closed information extraction and generalizes from fewer training data points than baselines. |
Copied to clipboard
| Challenge: | Using a learning approach for entity mentions is a key component of modern entity linking systems for both candidate generation and making linking predictions. |
| Approach: | They propose a training approach that builds minimum spanning arborescences over mentions and entities to explicitly model mention coreference relationships. |
| Outcome: | The proposed approach improves candidate generation recall and link accuracy on the biomedical dataset and on MedMentions, setting a new SOTA result in linking accuracy. |
Copied to clipboard
| Challenge: | Conditional neural text generation models generate high-quality outputs, but often focus on a mode when what we really want is a diverse set of options. |
| Approach: | They propose a search algorithm to construct lattices encoding a massive number of generation options. |
| Outcome: | The proposed algorithm encodes thousands of diverse options that remain grammatical and high-quality into one lattice. |
Copied to clipboard
| Challenge: | Existing models with synthetic indirect answers to yes-no questions are not beneficial when working with real conversations. |
| Approach: | They propose to annotate the underlying direct answers to yes-no questions in real conversations. |
| Outcome: | The proposed model outperforms the majority baseline but the task remains a challenge. |
Copied to clipboard
| Challenge: | a recent study examines the features and limits of LM adaptability to new tasks . many questions about the nature and limits remain unanswered . |
| Approach: | They evaluate adaptability to new tasks using a new benchmark, TaskBench500 . they find adaptation procedures differ dramatically in their ability to memorize small datasets . |
| Outcome: | The proposed benchmark compares 500 procedurally generated sequence modeling tasks to a new benchmark. |
Copied to clipboard
| Challenge: | sexism and hate speech detection models may be over-relying on core features . construct-driven CAD may induce models to ignore context in which core features are used . |
| Approach: | They propose to use construct-driven and construct-agnostic CAD to reduce model bias . sexism and hate speech detection models are trained on counterfactually augmented data . |
| Outcome: | Using a diverse set of CAD—construct-driven and construct-agnostic—reduces unintended bias. |
Copied to clipboard
| Challenge: | In computer vision, the trigger can be a fixed pattern overlaid on the images or videos. |
| Approach: | They propose an attention-based Trojan detector to distinguish Trojaned models from clean ones by observing the attention focus drifting behavior of Trojanes. |
| Outcome: | The proposed detector is based on transformer’s attention and can distinguish Trojan models from clean ones. |
Copied to clipboard
| Challenge: | Existing methods for data augmentation do not fully exploit the potential of DA in NLP. |
| Approach: | They propose an easy and plug-in framework for data augmentation to support effective text classification. |
| Outcome: | The proposed framework outperforms existing methods in most cases, but not using agent networks or pre-trained generation networks. |
Copied to clipboard
| Challenge: | Researchers have shown that many datasets contain statistical biases, or "annotation artifacts" that systems leverage to correctly predict entailment. |
| Approach: | They propose to use edited contexts to examine RoBERTa models' sensitivity to edited context to examine their model's sensitivity. |
| Outcome: | The proposed model can learn to condition on context, despite being trained on artifact-ridden datasets. |
Copied to clipboard
| Challenge: | Pretrained language models are typically learned over a large, static corpus and fine-tuned for various downstream tasks. |
| Approach: | They propose to continuously update a pretrained language model to adapt to emerging data and to keep track of the model's performance. |
| Outcome: | The proposed model can adapt to new corpora while retaining knowledge in earlier domains. |
Copied to clipboard
| Challenge: | a novel AI-empowered chat bot for learning as conversation can be applied to various domains without in-domain dialogue data. |
| Approach: | They propose a novel task where a user does not read a passage but gains information and knowledge through conversation with a teacher bot. |
| Outcome: | The proposed system can be transferred to various domains without in-domain dialogue data and can carry out conversations both informative and attentive to users. |
Copied to clipboard
| Challenge: | Hidden Markov Models (HMMs) and Probabilistic Context-Free Grammars (PCFGs) are widely used structured models. |
| Approach: | They use tensor rank decomposition to reduce computational complexities for a subset of FGGs subsuming HMMs and PCFGs. |
| Outcome: | The proposed model performs better on HMM modeling and unsupervised PCFG parsing than previous work. |
Copied to clipboard
| Challenge: | a survey of the NLP community shows that paper-reviewer matching is a problem . authors lose valuable time and opportunities by writing reviews that are arbitrarily low . |
| Approach: | They propose to use paper-reviewer matching to improve peer review . they identify common issues and perspectives on what factors should be considered . |
| Outcome: | The proposed method improves the quality of peer review and improves interpretable peer review assignments. |
Copied to clipboard
| Challenge: | Recent studies show that Neural Machine Translation models struggle to disambiguate polysemous words without lapsing into their most frequent senses. |
| Approach: | They propose a way to automatically create high-precision sense-annotated parallel corpora . they then propose 'fine-tuning' strategies to exploit these sense annotations during training . |
| Outcome: | The proposed approach achieves higher BLEU scores than its vanilla counterpart in 3 language pairs. |
Copied to clipboard
| Challenge: | Existing studies do not consider semantic information between incomplete utterance and rewritten utterant or model the semantic structure implicitly and insufficiently. |
| Approach: | They propose a query-Enhanced network to bring semantic structural knowledge between incomplete utterance and rewritten utteras . they adopt a fast and effective edit operation scoring network to model the relation between two tokens based on extra information and the well-designed network . |
| Outcome: | The proposed query template explicitly brings semantic structural knowledge between the incomplete utterance and the rewritten utterant making model perceive where to refer back to or recover omitted tokens. |
Copied to clipboard
| Challenge: | Existing methods for domain adaptation of abstractive dialogue summarization lack generalization ability on new domains. |
| Approach: | They propose a domain-oriented prefix-tuning model that uses a prefix module to alleviate domain entanglement and discrete prompts to guide the model to focus on key contents of dialogues. |
| Outcome: | The proposed model can be generalized to two multi-domain dialogue summarization datasets. |
Copied to clipboard
| Challenge: | Existing work on symbol grounding models (grounders) uses lazy few-shot learning to relate open-class words like green and above to their visual percepts; and symbolic reasoning with closed-class word categories like quantifiers and negation. |
| Approach: | They propose a procedure for learning to ground symbols from a sequence of stimuli consisting of an arbitrarily complex noun phrase and its designation in the visual scene. |
| Outcome: | The proposed procedure is based on a visual reference resolution task in which the learner is unaware of concepts that are part of the domain model and how they relate to visual percepts. |
Copied to clipboard
| Challenge: | Quantifiers are pervasive in NLU benchmarks and their occurrence at test time is associated with performance drops. |
| Approach: | They propose a generalized quantifier NLI task to quantify their contribution to the errors of NLU models. |
| Outcome: | The proposed model is based on a generalized quantifier theory and is compared with pre-trained models. |
Copied to clipboard
| Challenge: | Existing methods to evaluate test statistic are Monte Carlo approximations which use a summation over all 2 N possible swaps. |
| Approach: | They propose an exact algorithm for the paired-permutation test for a family of structured test statistics. |
| Outcome: | The proposed algorithm is 10x faster than the Monte Carlo approximation with 20000 samples on a common dataset. |
Copied to clipboard
| Challenge: | Pretraining languages improve cross-lingual transfer for BERT-based models . Interestingly, PLMs exhibit zero-shot cross-linguistic abilities on downstream examples in languages seen only during pretraining. |
| Approach: | They develop a quadratic time complexity method to estimate pretraining languages' relations between linguistic features and two downstream tasks. |
| Outcome: | The proposed method is effective on a diverse set of languages spanning different linguistic features and two downstream tasks. |
Copied to clipboard
| Challenge: | Aspect-based Sentiment Analysis (ABSA) aims to predict sentiment polarity towards aspects in sentences . a novel model for ABSA is proposed, but how to harness it is still a challenge . |
| Approach: | They propose a syntactic and semantic enhanced Graph Convolutional Network (SSEGCN) model for ABSA task using aspect-aware attention mechanism and self-attention. |
| Outcome: | The proposed model outperforms state-of-the-art methods on benchmark datasets. |
Copied to clipboard
| Challenge: | Recent work on controllable text generation has shown promise in successfully altering such text attributes. |
| Approach: | They propose to use empathetic data to reduce the toxicity of generated text by strategically sampling data based on empathy scores. |
| Outcome: | The proposed model significantly reduces the size of fine-tuning data to 7.5-30k samples while making significant improvements over state-of-the-art toxicity mitigation. |
Copied to clipboard
| Challenge: | Social media rumours can cause significant economic and social disruption. |
| Approach: | They propose a rumour detection algorithm that leverages transformers and graph attention networks to jointly model social media conversations and the network of users who engaged in them. |
| Outcome: | The proposed algorithm produces superior performance over four widely used benchmark rumour datasets in English and Chinese. |
Copied to clipboard
| Challenge: | Existing neural machine translation models learn the probability P (y|x) of the target sentence given the source sentence x. |
| Approach: | They propose to replace softmax activation with a multi-label classification layer that can model ambiguity more effectively. |
| Outcome: | The proposed multi-label classification layer can model ambiguity more effectively . it yields consistent BLEU score gains across six translation directions . |
Copied to clipboard
| Challenge: | Existing studies on Skill Extraction (SE) use crowd-sourced labels or annotations from a predefined skill inventory. |
| Approach: | They propose a dataset that contains 14.5K sentences and over 12.5K annotated spans. |
| Outcome: | The proposed model outperforms non-adapted models and single-task outperformed multi-task learning. |
Copied to clipboard
| Challenge: | Existing methods focus on sentencelevel event extraction (SEE), but they are inconsistent with actual situations. |
| Approach: | They propose a document-level event extraction framework which can model relation dependencies by a relation-augmented Attention Transformer. |
| Outcome: | The proposed framework can achieve state-of-the-art performance on two public datasets. |
Copied to clipboard
| Challenge: | Frame semantic parsing is a fundamental NLP task, which consists of three subtasks: frame identification, argument identification and role classification. |
| Approach: | They propose a frame semantic parser with a double-graph to derive knowledge-enhanced representations for frames and FEs. |
| Outcome: | The proposed method outperforms the state-of-the-art method by up to 1.7 F1-score on two FrameNet datasets. |
Copied to clipboard
| Challenge: | Existing approaches to tagging tasks are limited to predefined classes and require large-scale annotated data. |
| Approach: | They propose an Enhanced Span-based Decomposition method for Few-Shot Sequence Labeling to generalize on emerging, resource-scare domains. |
| Outcome: | The proposed method achieves state-of-the-art results on two popular FSSL benchmarks, FewNERD and SNIPS, and is more robust in noisy and nested tagging scenarios. |
Copied to clipboard
| Challenge: | Existing studies aim at extracting event arguments from a single sentence . document-level event extraction still remains under-explored . |
| Approach: | They propose a two-stream abstract meaning representation enhanced extraction model to extract event arguments from an entire document. |
| Outcome: | The proposed model outperforms state-of-the-art in extracting event arguments from documents by 2.54 F1 and 5.13 F1 on public RAMS and WikiEvents datasets. |
Copied to clipboard
| Challenge: | Controlled table-to-text generation is a new approach to generate textual descriptions for highlighted subparts of a table. |
| Approach: | They propose an equivariance learning framework which encodes tables with a structure-aware self-attention mechanism and a positional encoding mechanism to preserve relative position of tokens in the same cell. |
| Outcome: | The proposed framework is free to be plugged into existing table-to-text generation models and has improved T5-based models to offer better performance on ToTTo and HiTab. |
Copied to clipboard
| Challenge: | Existing KG-augmented models for commonsense question answering ignore the effectively fusing and reasoning over question context representations and the KG representations. |
| Approach: | They propose a novel model which combines a logical reasoning and a dynamic pruning mechanism to solve these limitations. |
| Outcome: | The proposed model improves existing models and performs interpretable reasoning on the CommonsenseQA and OpenBookQA datasets. |
Copied to clipboard
| Challenge: | Standard pre-trained language models do not see the characters that compose each token's string representation. |
| Approach: | They probe the embedding layer of pretrained language models and show that models learn the internal character composition of whole word and subword tokens without seeing the characters coupled with the tokens. |
| Outcome: | The embedding layers of RoBERTa and GPT2 hold enough information to accurately spell up to a third of the vocabulary and reach high character ngram overlap across all token types. |
Copied to clipboard
| Challenge: | Existing tasks for evaluating story understanding and generation focus on reasoning plots from context, but they focus on bridging plots with implied morals. |
| Approach: | They propose two understanding tasks and two generation tasks to assess machines' ability to bridge story plots and implied morals. |
| Outcome: | The proposed tasks are based on a dataset of Chinese and English moral stories . they show that the proposed models can perform better than existing models . |
Copied to clipboard
| Challenge: | Existing work on relation extraction focuses on constructing explicit structured features using knowledge graph and dependency tree. |
| Approach: | They propose a method to extract multi-granularity features based solely on the original input sentences. |
| Outcome: | The proposed method outperforms state-of-the-art models that even use external knowledge on three public benchmarks: SemEval 2010 Task 8, Tacred, and Tacred Revisited. |
Copied to clipboard
| Challenge: | Existing approaches for speech translation focus on using additional data from MT and automatic speech recognition (ASR). |
| Approach: | They propose a cross-modal contrastive learning method for end-to-end speech-totext translation. |
| Outcome: | The proposed method outperforms existing methods on a popular benchmark MuST-C. |
Copied to clipboard
| Challenge: | In this paper, we consider mimicking fictional characters as a promising direction for building engaging conversation models. |
| Approach: | They propose a task where only a few utterances of each fictional character are available to generate responses mimicking them. |
| Outcome: | The proposed method generates responses better reflecting the style of fictional characters than baseline methods. |
Copied to clipboard
| Challenge: | Long documents are tedious to read through and can be authored by multiple entities . traditional document navigation is through a Table of Contents (ToC) but there is no way to highlight information relevant to different personas. |
| Approach: | They propose a dynamic table of content-based navigator that highlights sections of interest . DYNAMICTOC is augmented with short questions to assist users in understanding underlying content . |
| Outcome: | The proposed navigator highlights sections of interest in documents as per the aspects relevant to different personas. human and automatic evaluations suggest the efficacy of both end-to-end pipeline and different components. |
Copied to clipboard
| Challenge: | Pre-trained language models (PLMs) have proved to be effective on various natural language understanding tasks. |
| Approach: | They propose a domain adaption framework which modulates the intermediate hidden representations of PLMs with domain knowledge, consisting of entities and their relational facts. |
| Outcome: | The proposed framework outperforms adaptive pre-training on question answering and named entity recognition tasks on multiple datasets across different domains. |
Copied to clipboard
| Challenge: | Recent studies on large-scale in-context language models have reported successful in-const zero- and few-shot learning ability. |
| Approach: | They investigate the effects of the pretraining corpus on in-context learning in a Korean-centric model. |
| Outcome: | The study shows that pretraining corpus size does not determine in-context learning ability . the findings suggest that in-constext learning is not always competitive . |
Copied to clipboard
| Challenge: | Existing models for long sequences are not efficient due to the quadratic space and time complexity of the self-attention modules. |
| Approach: | They propose to reduce the quadratic complexity to linear (modulo logarithmic factors) by low-dimensional projection and row selection. |
| Outcome: | The proposed methods outperform transformer-based models with smaller time/space footprint on the Long Range Arena benchmark. |
Copied to clipboard
| Challenge: | Existing frameworks that focus on self personas ignore the value of partner persona . experimental results show that our framework generates relevant, interesting, coherent and informative partner personages even compared to ground truth partner personagers. |
| Approach: | They propose a framework that leverages automatic partner personas generation to enhance dialogue response generation. |
| Outcome: | The proposed framework generates relevant, interesting, coherent and informative partner personas even compared to ground truth partner person . it surpasses baselines that condition on ground truth persona . |
Copied to clipboard
| Challenge: | Existing approaches to slang interpretation rely on context but ignore semantic extensions common in slings . a semantically informed slapping framework can be applied to enhancing machine translation of informal language . |
| Approach: | They propose a semantically informed slang interpretation framework that considers contextual and semantic appropriateness of a candidate interpretation for a query s. |
| Outcome: | The proposed framework achieves state-of-the-art accuracy in slang interpretation in English and in other languages. |
Copied to clipboard
| Challenge: | Existing fact extraction and verification tasks only consider evidence of a single format . Existing models convert evidence into either sentences or tables, thus losing context information . |
| Approach: | They propose a Dual Channel Unified Format fact verification model which unifies various evidence into parallel streams, i.e., natural language sentences and a global evidence table, simultaneously. |
| Outcome: | The proposed model outperforms existing models in two formats by a large margin . it makes the most of existing tables and tables to absorb evidence of two formats . |
Copied to clipboard
| Challenge: | Existing data augmentation methods miss the important characteristic of compositionality, meaning of a complex expression is built from its sub-parts. |
| Approach: | They propose a compositional data augmentation approach for natural language understanding called TreeMix that leverages constituency parsing tree to decompose sentences into constituent sub-structures and the Mixup data enhancing technique to recombine them to generate new sentences. |
| Outcome: | The proposed approach outperforms current state-of-the-art methods on text classification and SCAN. |
Copied to clipboard
| Challenge: | In this paper we examine patterns of colexification as an aspect of lexical-semantic organization, and compare several approaches to build large scale graphs across 499 world languages. |
| Approach: | They propose to use patterns of colexification as an aspect of lexical-semantic organization to build large scale synset graphs across a typologically diverse set of 499 world languages. |
| Outcome: | The proposed models are evaluated against human judgments on a semantic similarity task for nine languages. |
Copied to clipboard
| Challenge: | Existing knowledge-grounded conversational benchmarks produce factually invalid statements, a phenomenon commonly called hallucination. |
| Approach: | They conduct a human study on knowledge-grounded conversational benchmarks and state-of-the-art models. |
| Outcome: | The findings raise important questions on the quality of existing datasets and models. |
Copied to clipboard
| Challenge: | Recursive noun phrases have interesting semantic properties, yet it is unknown whether language models have such knowledge. |
| Approach: | They propose a dataset of three textual inference tasks targeting recursive noun phrases . they show that such knowledge is learnable with appropriate data . |
| Outcome: | The proposed model achieves strong zero-shot performance on an extrinsic Harm Detection task. |
Copied to clipboard
| Challenge: | Existing work on translationese neglects important factors and conclusions are mostly correlational but not causal. |
| Approach: | They use a dataset where MT training data are also labeled with human translation directions to examine the impact of translationese on machine translation evaluation. |
| Outcome: | The proposed model learns in the same direction as human translation directions. |
Copied to clipboard
| Challenge: | Fig. 1 shows how text-only and image-only models can capture commonsense visual attributes, but reporting bias affects their performance. |
| Approach: | They use a Visual Commonsense Tests dataset to validate their findings . they find multimodal models better reconstruct attribute distributions, but are still subject to reporting bias . |
| Outcome: | The proposed model improves on the unimodal and multimodal models, but is still subject to reporting bias. |
Copied to clipboard
| Challenge: | Existing models for natural language understanding are limited to processing only a few hundred words at a time. |
| Approach: | They propose a dataset with context passages in English that have an average length of 5,000 tokens. |
| Outcome: | a new dataset with long-text comprehension questions is used to test models on long-document comprehension . the questions are validated by contributors who have read the entire passage, not just excerpts . only half of the questions can be answered by annotators working under tight time constraints . |
Copied to clipboard
| Challenge: | Interpretability methods are developed to understand the working mechanisms of black-box models. |
| Approach: | They propose a mathematical framework for quantifying model understanding with an explanation summary. |
| Outcome: | The proposed framework highlights limitations in the current practice and reveals easily overlooked properties of the model. |
Copied to clipboard
| Challenge: | AMR parsing has experienced an unprecendented increase in performance in the last three years due to a mixture of effects including architecture improvements and transfer learning. |
| Approach: | They propose to combine Smatch-based ensembling techniques with ensemble distillation to overcome this diminishing returns of silver data. |
| Outcome: | The proposed technique can produce gains rivaling those of human annotated data for QALD-9 and achieve a new state-of-the-art for BioAMR. |
Copied to clipboard
| Challenge: | Recent studies show that models encode syntactic information redundantly . this allows researchers to boost models' performance by injecting syntaktic information into embeddings . |
| Approach: | They propose a new probe design that guides probes to consider all syntactic information present in embeddings. |
| Outcome: | The proposed model improves performance by injecting syntactic information into models. |
Copied to clipboard
| Challenge: | Existing work on document-level relation extraction has focused on end-to-end setting that extracts global entities and relations jointly. |
| Approach: | They propose to introduce a two-way interaction between COREF and RE that is specifically designed to leverage task characteristics, bridging decisions of two tasks for direct task interference. |
| Outcome: | The proposed model achieves the best performance by up to 2.3/5.1 F1 over the baseline. |
Copied to clipboard
| Challenge: | Large language models can perform semantic parsing with little training data, when prompted with in-context examples. |
| Approach: | They propose to map natural language to a controlled natural language-like representation . they find that OpenAI Codex performs better on such tasks than equivalent GPT-3 models . |
| Outcome: | The proposed model performs better on large parsing tasks than GPT-3 models on Overnight and SMCalFlow. |
Copied to clipboard
| Challenge: | Academic research is an exploratory activity to discover new solutions to problems . prior work focused on the sentence as the basic unit of generation, neglecting that related work sections consist of variable length text fragments derived from different information sources. |
| Approach: | They propose a Citation Oriented Related Work Annotation dataset that labels citation text fragments . they propose linguistically-motivated framework for human-in-the-loop, abstractive related work generation . |
| Outcome: | The proposed framework is based on a Citation Oriented Related Work Annotation dataset . it automatically tags unlabeled related work sections on the dataset based upon the proposed model . |
Copied to clipboard
| Challenge: | Existing work on lifelong learning requires incremental memory space to learn a model . existing work on experience replay or elastic weighted consolidation requires incremental space . |
| Approach: | They propose a framework that leverages a recall optimization mechanism to memorize parameters of previous tasks via regularization and a domain drift estimation algorithm to compensate the drift between different domains in the embedding space. |
| Outcome: | The proposed framework outperforms SOTA models on paraphrase and dialog response generation tasks. |
Copied to clipboard
| Challenge: | Experimental results show that MACLR achieves superior performance compared to other baseline methods. |
| Approach: | They propose to pre-train Transformer-based encoders with self-supervised contrastive losses to learn the semantic embeddings of instances and labels with raw text. |
| Outcome: | The proposed method improves on the EZ-XMC model with a limited number of ground-truth positive pairs. |
Copied to clipboard
| Challenge: | Traditionally, researchers used manual coding to track conflict processes worldwide, but the high costs and slow pace of domain experts make it difficult and costly to monitor complex and rapidly changing conflicts. |
| Approach: | They propose a domain-specific pre-trained language model for conflict and political violence that can be used to train a language model from scratch and continue training. |
| Outcome: | The proposed model outperforms BERT in conflict research. |
Copied to clipboard
| Challenge: | Prompt-based learning is an emerging paradigm for exploiting knowledge learned by a pretrained language model. |
| Approach: | They propose a method to automatically select label mappings for few-shot text classification with prompting. |
| Outcome: | The proposed method achieves competitive performance on the GLUE benchmark without human effort or external resources. |
Copied to clipboard
| Challenge: | Pre-trained language models have shown successful progress in many text understanding benchmarks. |
| Approach: | They propose a strategy to re-rank language model predictions based on interaction and feedback from the environment. |
| Outcome: | The proposed approach shows competitive performance on subgoal prediction and task completion in the ALFRED benchmark compared to prior methods that assume more subgoals supervision. |
Copied to clipboard
| Challenge: | Existing prompt tuning methods use a fixed prompt in each input instance during the model training stage. |
| Approach: | They propose a conditional prompt generation method to generate prompts for each input instance. |
| Outcome: | The proposed method outperforms other prompt tuning methods while tuning fewer parameters. |
Copied to clipboard
| Challenge: | Existing methods for fine-tuning pre-trained language models can cause severe over-fitting. |
| Approach: | They propose an Embedding Hallucination method which generates auxiliary embedding-label pairs to expand the fine-tuning dataset. |
| Outcome: | The proposed method outperforms current fine-tuning methods in a wide range of language tasks. |
Copied to clipboard
| Challenge: | speculative trading of highly volatile assets such as cryptocurrencies and meme stocks presents a new challenge in the financial realm. |
| Approach: | They propose a multi-span bubble detection task based on social media hype and a set of sequence-to-sequence hyperbolic models . they use data from 9 exchanges over five years to test their models based upon the power-law dynamics of cryptocurrencies and user behavior on social networks. |
| Outcome: | The proposed model is able to detect bubbles on a set of reddit and twitter posts spanning over two million tweets over five years . |
Copied to clipboard
| Challenge: | k-nearest-neighbor machine translation (kNN-MT) is a state-of-the-art machine translation technique . however, it requires conducting kNN searches for each decoding step, which increases the cost of decoding . |
| Approach: | They propose to move the time-consuming kNN search forward to the preprocessing phase and introduce k Nearest Neighbor Knowledge Distillation (kNN-KD) that trains the base NMT model to directly learn the knowledge of kN. |
| Outcome: | The proposed method improves over the state-of-the-art model while maintaining the same training and decoding speed as the standard model. |
Copied to clipboard
| Challenge: | Extensive experiments with autoregressive transformer LMs show that DEMix layers reduce test-time perplexity and increase training efficiency. |
| Approach: | They introduce a new domain expert mixture layer that enables conditioning a language model on the domain of the input text. |
| Outcome: | Experiments with 1.3B LMs show that DEMix layers reduce test-time perplexity, increase training efficiency, and enable rapid adaptation. |
Copied to clipboard
| Challenge: | a recent study has shown that GPT-3 fine-tuning models with limited examples is effective . a contrastive learning framework clusters inputs from the same class under different augmented “views” and repels those from different classes. |
| Approach: | They propose a supervised contrastive framework that clusters inputs from the same class under different augmented "views" they combine a contrastive loss with the standard masked language modeling loss in prompt-based few-shot learners . |
| Outcome: | The proposed framework improves on the state-of-the-art methods in a diverse set of 15 language tasks. |
Copied to clipboard
| Challenge: | Recent work in this area has harnessed the language-invariant qualities of pre-trained Multi-lingual Language Models. |
| Approach: | They propose to use adversarial language adaptation to train a model to detect events in a target language. |
| Outcome: | The proposed model achieves state-of-the-art on 8 different language pairs, using 4 languages from unrelated families. |
Copied to clipboard
| Challenge: | Existing datasets displaying high degree of implicit abuse are biased . current methods focus on explicit abuse, but there is little work on implicit forms of abuse . |
| Approach: | They propose to model atomic negative sentences to address implicit abuse by addressing its different subtypes and then separate them into subtype. |
| Outcome: | The proposed approach generalizes across different identities and languages. |
Copied to clipboard
| Challenge: | Existing work on semantic role labeling treats symbolic labels as symbolic . labeled data is costly and often lacking in many tasks, domains, and languages. |
| Approach: | They propose to retrieve and leverage semantic role labels from annotation guidelines . argument classification is at the core of Semantic Role Labeling . |
| Outcome: | The proposed model achieves state-of-the-art on a CoNLL09 dataset injected with label definitions given the predicate senses. |
Copied to clipboard
| Challenge: | Existing studies on text classification of the Dark Web have been ineffective due to its inherent characteristics. |
| Approach: | They propose a publicly available Dark Web dataset tailored towards text-based analysis. |
| Outcome: | The proposed method compares with an existing public Dark Web dataset and evaluates its suitability for various use cases. |
Copied to clipboard
| Challenge: | Existing methods to control for text-based confounders rely on assumption that there is no treatment leakage . prior literature has assumed that documents only contain information about confounder, but not about treatment assignment. |
| Approach: | They define the treatment leakage problem and propose methods to mitigate it . they remove treatment-related signal from text in a pre-processing step . |
| Outcome: | The proposed method can mitigate the problem of treatment leakage by removing the treatment-related signal from the text. |
Copied to clipboard
| Challenge: | Existing methods for regularizing a model are agnostic to the training model and may not be effective for perturbed inputs. |
| Approach: | They propose an augmentation method of adding a discrete noise that would incur the highest divergence between predictions by replacing tokens while keeping original semantics. |
| Outcome: | The proposed method outperforms baselines on semi-supervised text classification tasks and a robustness benchmark. |
Copied to clipboard
| Challenge: | Factual inconsistencies in generated summaries severely limit the practical applications of abstractive dialogue summarization. |
| Approach: | They propose a typology of factual errors to better understand hallucinations generated by current models and a contrastive fine-tuning strategy to improve the factual consistency and overall quality of summaries. |
| Outcome: | The proposed model significantly reduces all kinds of factual errors on both SAMSum dialogue summarization and AMI meeting summarizing datasets. |
Copied to clipboard
| Challenge: | Emotion recognition in conversation is inaccurate if the previous utterances are not taken into account, so many studies reflect the dialogue context to improve the performance. |
| Approach: | They propose a method that combines pre-trained memory with the context model to improve the performance of the context models. |
| Outcome: | The proposed method achieves the first or second performance on all data and is state-of-the-art among systems that do not leverage structured data. |
Copied to clipboard
| Challenge: | Existing pre-trained summarization models produce text that is factually inconsistent with the input. |
| Approach: | They present a scale-based scale for Likert rating and a scoring algorithm for Best-Worst Scaling to improve crowdsourcing reliability. |
| Outcome: | The proposed model is more reliable than existing models on two news summarization datasets. |
Copied to clipboard
| Challenge: | Current models for dialogue summarization have flaws that may not be well exposed by frequently used metrics such as ROUGE. |
| Approach: | They propose to re-evaluate 18 categories of metrics in terms of four dimensions: coherence, consistency, fluency and relevance, as well as a unified human evaluation of various models for the first time. |
| Outcome: | The proposed dataset will be used to evaluate 18 categories of metrics in terms of coherence, consistency, fluency and relevance, and a unified human evaluation of various models for the first time. |
Copied to clipboard
| Challenge: | Keyphrase extraction is a fundamental task in natural language processing that aims to extract a set of phrases with important information from a source document. |
| Approach: | They propose a hyperbolic matching model to explore keyphrase extraction in hyperbolical space using word embeddings from RoBERTa to capture hierarchical syntactic and semantic structures. |
| Outcome: | The proposed model outperforms the state-of-the-art models on six benchmark datasets and outperformed previous models. |
Copied to clipboard
| Challenge: | Prompt-based methods have been successfully applied in few-shot learning tasks . however, when applied to token-level labeling tasks, it would be time-consuming to enumerate the template queries over all potential entity spans. |
| Approach: | They propose a method to reformulate NER tasks as LM problems without templates. |
| Outcome: | The proposed method is 30.12 times faster than the template-based method under few-shot settings. |
Copied to clipboard
| Challenge: | Existing benchmarks for relation extraction are built on sentence-level corpora, but document-level ones provide more realism. |
| Approach: | They propose a few-shot document-level relation extraction benchmark based on document-based corpora. |
| Outcome: | The proposed benchmark is based on two existing supervised learning data sets, DocRED and sciERC. |
Copied to clipboard
| Challenge: | Existing approaches to model long-term dependencies are limited to long texts with thousands of words. |
| Approach: | They propose a look-ahead memory that augments the recurrence memory by attending to the right-side tokens and interpolating with the old memory states to maintain long-term information in the history. |
| Outcome: | Experiments on widely used language modeling benchmarks show that LaMemo outperforms baseline models with recurrence memory. |
Copied to clipboard
| Challenge: | Existing models for text generation do not need syntactic information such as constituency parses or semantic information such a paraphrase pairs. |
| Approach: | They propose a generative model which exhibits disentangled latent representations of syntax and semantics by using Attention in its decoder. |
| Outcome: | The proposed model outperforms supervised models on syntax/semantics transfer and shows that it can read latent variables with keys and values. |
Copied to clipboard
| Challenge: | Existing approaches to lexically constrained neural machine translation suffer from high latency. |
| Approach: | They propose a plug-in algorithm for non-autoregressive translation for this problem . they propose ACT to familiarize the model with the source-side context of constraints . |
| Outcome: | The proposed model improves over the backbone constrained NAT model in constraint preservation and translation quality, especially for rare constraints. |
Copied to clipboard
| Challenge: | Recent research reveals that transformer-based models are biased towards extracting knowledge about object relations. |
| Approach: | They propose to use transformer-based models to extract knowledge about object relations to investigate whether they can be used to extract object relations. |
| Outcome: | The proposed models outperform static models in many respects and perform much worse than similarity measures and classifiers. |
Copied to clipboard
| Challenge: | Existing personalized dialogue systems extract user profiles from dialogue history to guide personalized response generation. |
| Approach: | They propose to refine the user dialogue history on a large scale to obtain more persona information from the dialogue history and leverage other similar users' data to enhance personalization. |
| Outcome: | The proposed model can handle more dialogue history and obtain more abundant and accurate persona information. |
Copied to clipboard
| Challenge: | Covid-19 infodemic has led to low quality information leading to poor health decisions . authors propose a framework for analyzing false claims and reasoning about the decisions a person makes . |
| Approach: | They propose a framework linking stance and reason analysis and moral sentiment analysis. |
| Outcome: | The proposed framework provides reliable predictions even in low-supervision settings. |
Copied to clipboard
| Challenge: | Recent studies show pre-trained language models contain matching subnetworks that have similar transfer learning performance as the original PLM. |
| Approach: | They propose to prune matching subnetworks using magnitude-based pruning . they propose to optimize the subnetwork structure towards the pre-training objectives . |
| Outcome: | The proposed method is more efficient in searching subnetworks and advantageous when fine-tuning within a range of data scarcity. |
Copied to clipboard
| Challenge: | Social chatbots evolve rapidly with large pretrained language models. |
| Approach: | They propose effective defense objectives to protect persona leakage from hidden states by a simple neural network. |
| Outcome: | The proposed defense objectives reduce the attack accuracy from 37.6% to 0.5% while preserving language models’ powerful generation ability. |
Copied to clipboard
| Challenge: | Existing frameworks for dialogue model evaluation are lacking to investigate these biases . a number of dialogue metrics are biased and can cause unforeseen problems . |
| Approach: | They propose an adversarial test-suite which generates problematic variations of various dialogue aspects using automatic heuristics. |
| Outcome: | The proposed test-suite generates problematic variations of various dialogue aspects using automatic heuristics. |
Copied to clipboard
| Challenge: | toxicity annotations are often ignored because of its subjective nature and lack of nuance. |
| Approach: | They examine the effect of annotator identities and beliefs on toxic language annotations by considering posts with three characteristics: anti-Black language, African American English (AAE) dialect, and vulgarity. |
| Outcome: | The findings show strong associations between annotator identity and beliefs and ratings of toxicity. |
Copied to clipboard
| Challenge: | Existing methods to correct ASR errors focus on fixed-length corrections, but rarely consider variable-length ones. |
| Approach: | They propose a non-autoregressive method to correct Chinese ASR errors . they use phonological tokens to extend the source sentence for variable-length correction . |
| Outcome: | The proposed method improves word error rate and speeds up inference by 6.2 times compared with the autoregressive model. |
Copied to clipboard
| Challenge: | Existing datasets and models target hate speech but ignore context . Existing models target either hate speech or hate and counter speech but disregard context - a new study shows that context is critical to identify hate and anti-hate speech. |
| Approach: | They propose to use context to identify hate and counter speech in a reddit conversation thread. |
| Outcome: | The proposed model improves when and why context is taken into account. |
Copied to clipboard
| Challenge: | a large corpus of documents is available for summarization tasks in English . supervised methods require adequate corpora for summarizing . |
| Approach: | They describe a corpus of catalan and spanish newspapers that can be used to train summarization models for Catalan, Spanish and other languages. |
| Outcome: | The proposed corpus can be used to train summarization models for Catalan and Spanish. |
Copied to clipboard
| Challenge: | a pretrained model is optionally adapted through domain-specific pretraining, followed by task-specific finetuning. |
| Approach: | They establish a suite of eight tasks across different domains to quantify the effects of temporal misalignment in modern NLP systems. |
| Outcome: | The proposed tasks are based on eight domains and periods of time spanning five years or more and show that they have stronger effects than previous studies. |
Copied to clipboard
| Challenge: | Existing approaches to learning semantically meaningful sentence embeddings are limited by the complexity of pre-trained models. |
| Approach: | They propose a sentence embedding learning approach that exploits both visual and textual information via a multimodal contrastive objective. |
| Outcome: | The proposed approach improves the state-of-the-art average Spearman’s correlation by 1.7% on a variety of semantic textual similarity tasks. |
Copied to clipboard
| Challenge: | Existing methods to extract relational feature signals from natural language sentences use self-supervised clustering and classification that cause gradual drift problems. |
| Approach: | They propose a framework that derives hierarchical signals from relational feature space using cross hierarchy attention and effectively optimizes relation representation of sentences under exemplar-wise contrastive learning. |
| Outcome: | The proposed framework can extract the relationship between entities from natural language sentences without prior knowledge on relation scope or distribution. |
Copied to clipboard
| Challenge: | Existing models claim to be able to align object tokens with specific visual targets, but there are non-negligible gaps between the two. |
| Approach: | They conduct diagnostic experiments to examine how the agents perceive multimodal input by ablation diagnostics input data. |
| Outcome: | The results show that indoor and outdoor navigation agents refer to object and direction tokens when making decisions. |
Copied to clipboard
| Challenge: | Social value alignment is the ability to create agents that act in alignment with socially beneficial norms and values in interactive narratives or text-based games. |
| Approach: | They introduce a game-value ALignment agent that uses social commonsense to restrict its action space to actions that are aligned with socially beneficial values. |
| Outcome: | The proposed agent improves state-of-the-art task performance by 4% while reducing the frequency of socially harmful behaviors by 25% compared to strong contemporary value alignment approaches. |
Copied to clipboard
| Challenge: | despite being a common figure of speech, hyperbole is under-researched in Figurative Language Processing . we use an unsupervised method to generate hyperbolic paraphrases from literal sentences . |
| Approach: | They propose an unsupervised method for hyperbole generation that does not require parallel literal-hyperbole pairs. |
| Outcome: | The proposed method outperforms baseline systems and is based on a large-scale English hyperbole corpus. |
Copied to clipboard
| Challenge: | a method for learning an NLI model is time-consuming and resource-intensive, but it can save time and resources. |
| Approach: | They propose a method for predicting model performance without fine-tuning it . they compare sentence embeddings with cosine similarity to classifiers . |
| Outcome: | The proposed method can save time and resources by comparing pre-trained models to real-world datasets. |
Copied to clipboard
| Challenge: | Existing definitions of system-level correlations are inconsistent with how they are used to evaluate systems. |
| Approach: | They propose to calculate correlations only on pairs of systems separated by small differences in automatic scores . they propose to use the full test set instead of the subset of summaries judged by humans . |
| Outcome: | The proposed changes improve the accuracy of the estimated correlations on pairs of systems separated by small differences in automatic scores. |