Papers with model
Copied to clipboard
| Challenge: | Metaphors are commonly found in advertising and internet memes, but lack of high-quality textual data is a challenge for language models . a new framework for multi-modal metaphor detection is being developed to address these challenges . |
| Approach: | They propose a framework that extracts and integrates knowledge from Large Language Models into smaller ones to improve model performance. |
| Outcome: | The proposed framework outperforms existing models on the MET-MEME dataset. |
Copied to clipboard
| Challenge: | Existing abstractive summarization models rely heavily on reference summaries and lack control over their performance. |
| Approach: | They propose a BRIO paradigm to reduce the dependence on reference summaries by fine-tuning pre-trained language models and training them with the paradigm. |
| Outcome: | The proposed paradigm outperforms existing models on Vietnamese and CNNDM datasets while maintaining the main content of the original text. |
Copied to clipboard
| Challenge: | Manually labeled training data is expensive, noisy, and often scarce . semi-supervised learning methods can be used to improve model performance . |
| Approach: | They explore different methods for consistency training on unlabeled data . they use human paraphrasing, back-translation, and dropout to augment unlabed data. |
| Outcome: | The proposed methods outperform purely supervised learning on unlabeled data. |
Copied to clipboard
| Challenge: | Synthetic data suffer from poor diversity, which leads to performance limitations. |
| Approach: | They propose a graph-propagated data augmentation framework for named entity recognition that uses graph propagation to build relationships between labeled data and unlabeled natural texts. |
| Outcome: | The proposed framework improves on a low-resource named entity recognition dataset. |
Copied to clipboard
| Challenge: | TextAttack provides implementations of 16 adversarial attacks from the literature and supports a variety of models and datasets. |
| Approach: | They introduce a Python framework for adversarial attacks, data augmentation, and adversarially training in NLP. |
| Outcome: | This paper introduces a Python framework for adversarial attacks, data augmentation, and adversarially training in NLP. |
Copied to clipboard
| Challenge: | Pretrained language models provide high-quality contextualized word embeddings, but training question answering models requires large amounts of annotated data for specific domains. |
| Approach: | They propose a framework for automatically generating more non-trivial question-answer pairs to improve model performance. |
| Outcome: | The proposed framework outperforms state-of-the-art (SOTA) pretrained language models and transfer learning approaches on standard question-answering benchmarks. |
Copied to clipboard
| Challenge: | Existing approaches to lifelong language learning rely on sparse experience replay to prevent catastrophic forgetting. |
| Approach: | They propose to use a selective memory population to store a uniform number of samples from the entire data stream to improve model performance. |
| Outcome: | The proposed methods show that they are relevant for lifelong language learning tasks, especially for low memory size, and consistent with computer vision studies. |
Copied to clipboard
| Challenge: | Existing researches on conversation-based QA focus on document-based tasks . current researche focuses on document based tasks, but there is a lack of researche on conversation based qa . |
| Approach: | They propose a multi-span extraction model on conversation-based QA and introduce continual pre-training and multi-task learning schemes to further improve model performance. |
| Outcome: | The proposed model outperforms baseline on two Chinese datasets and will be released for research purposes. |
Copied to clipboard
| Challenge: | Attention regularisation aims to supervise the attention patterns in language models like BERT. |
| Approach: | They compare regularisation on human rationales with random tokens to find that human-annotated rationale is better at reducing model sensitivity to spurious correlations. |
| Outcome: | The proposed regularisation method improves model performance and model robustness, but not with human-annotated rationales. |
Copied to clipboard
| Challenge: | Adversarial attacks against Language models (LMs) are a significant concern. |
| Approach: | They propose an approach to automatically learn a policy to generate challenging examples that improve the model’s performance. |
| Outcome: | The proposed approach outperforms baselines and exhibits generalizability across classifiers and datasets. |
Copied to clipboard
| Challenge: | Existing work on NN models with output constraints has not been able to categorize them in a unified manner. |
| Approach: | They propose new algorithms to integrate the information of main task and constraint injection . they use the H-score as a metric for considering main task metric and constrain infringement simultaneously . |
| Outcome: | The proposed algorithms integrate the information of main task and constraint injection, inspired by continual-learning algorithms. |
Copied to clipboard
| Challenge: | Existing contrastive learning-based methods struggle with data sparsity in real-world recommendations . Graph collaborative filtering incorporates contrastive training as an auxiliary task to improve performance . |
| Approach: | They propose a perturbation-driven dual auxiliary contrastive learning task for collaborative filtering . structure perturbation and weight perturbation are used to construct two graphs . |
| Outcome: | The proposed model outperforms benchmark models on multiple public datasets. |
Copied to clipboard
| Challenge: | Existing methods to segment unformatted text and transcripts explicitly train to predict segment boundaries, but they fail to provide a large annotated dataset. |
| Approach: | They propose a method to generate hierarchical segmentation structures based on Wikipedia annotations by using a neural conditional random field. |
| Outcome: | The proposed method outperforms or achieves competitive performance when compared to previous state-of-the-art algorithms. |
Copied to clipboard
| Challenge: | Recent introduced instruction-paradigm empowers non-expert users to leverage NLP resources by defining a new task in natural language. |
| Approach: | They propose to define a task in natural language without creating task-specific datasets or building models. |
| Outcome: | The proposed model outperforms multitask learning models but is far from state-of-the-art task-specific models. |
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have achieved impressive performance on many language tasks. |
| Approach: | They synthesize a high-quality dataset of error correction pairs to evaluate and improve LLMs for mobile applications by reweighting the sample. |
| Outcome: | The proposed model improves on offline evaluation and live A/B testing, given the LLM performance on offline data and scores from a small privacy-preserving on-device language model. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown compelling abilities in reasoning, decision-making, and instruction following. |
| Approach: | They propose a benchmark to evaluate the proficiency of large language models (LLMs) in judging and identifying safety risks given agent interaction records. |
| Outcome: | The proposed model outperforms the best-performing model, GPT-4o, while no other models significantly exceed the random. |
Copied to clipboard
| Challenge: | Existing work has attempted to model individual annotation behaviour rather than predicting aggregated labels. |
| Approach: | They propose to model individual annotator behaviour rather than predicting aggregated labels by adding group-specific layers to multi-annotator models to account for sociodemographics. |
| Outcome: | The proposed model does not significantly improve on toxic content detection tasks. |
Copied to clipboard
| Challenge: | Existing studies on individual differences and language representations focused on predicting selected attributes from text or conditioning text representations on author attributes. |
| Approach: | They propose a self-supervised approach to learning language-based user encodings using transformers. |
| Outcome: | The proposed model can pick up on complex linguistic signatures of users and infer rich information about them. |
Copied to clipboard
| Challenge: | Among social media platforms, Reddit has emerged as the most promising one due to its anonymity and its focus on topic-based communities (subreddits) . a challenge for previous work on suicide risk assessment has been the small amount of labeled data. |
| Approach: | They propose to use social media to collect user data from r/SuicideWatch subreddit and annotate it with user-level suicide risk: no-risk, low-risk and high-risk. |
| Outcome: | The proposed model improves by using pseudo-labeling based on related issues around mental health (e.g., anxiety, depression) |
Copied to clipboard
| Challenge: | Recent work shows document-level contexts can significantly improve Named Entity Recognition models. |
| Approach: | They propose to find external contexts of a sentence by retrieving and selecting a set of semantically relevant texts through a search engine with the original sentence as the query. |
| Outcome: | The proposed approach can achieve new state-of-the-art performance on 8 NER data sets across 5 domains. |
Copied to clipboard
| Challenge: | Modern NLP systems require high-quality annotations, but experts are expensive and lay annotators may not have the knowledge to provide high- quality annotations. |
| Approach: | They propose to directly model instance difficulty to improve model performance and to route instances to appropriate annotators. |
| Outcome: | The proposed model improves performance on a biomedical information extraction task using expert and lay annotations. |
Copied to clipboard
| Challenge: | Existing methods for fine-tuning pre-trained models are time-consuming and memory-inefficient. |
| Approach: | They propose a method that inserts learnable vectors into each Transformer layer . they propose SL to encourage diversity in prefix tokens . |
| Outcome: | Extensive experiments validate the effectiveness of Prefix Tuning in sentence and token classification tasks. |
Copied to clipboard
| Challenge: | Existing methods to generate training data using weakly labeled data are costly and limited . |
| Approach: | They propose a method for acquiring and labeling affective events with multiple view co-prompting using pre-trained language models. |
| Outcome: | The proposed approach improves state-of-the-art affective event classifier on two datasets. |
Copied to clipboard
| Challenge: | evaluating model robustness to adversarial attacks can provide deeper understanding of how deep neural networks work and what kind of linguistic information is actually captured by neural networks. |
| Approach: | They propose a method for strategic sentence-level perturbations to evaluate model robustness to adversarial attacks using character and word perturbations. |
| Outcome: | The proposed model improves model performance during adversarial attacks by using ensembles and predicts errors in adversarials. |
Copied to clipboard
| Challenge: | Existing work on front-end code generation fails to provide visual fidelity and rendering quality for front- end developers. |
| Approach: | They propose a three-stage pipeline to enhance front-end code generation capabilities in LLMs . they use synthetic data, quality-controlled supervised fine-tuning, and reinforcement learning . |
| Outcome: | The proposed model achieves competitive performance with frontier models while maintaining generation efficiency. |
Copied to clipboard
| Challenge: | Experimental results show that data augmentation improves accuracy over strong baselines. |
| Approach: | They propose to use translationese as input for GEC data augmentation to overcome stylistic discrepancies . they propose to obtain human-translated texts with a more similar style to non-native texts . |
| Outcome: | The proposed method improves correction accuracy over strong baselines on four GEC benchmarks. |
Copied to clipboard
| Challenge: | Recent studies for relation extraction (RE) leverage the dependency tree of the input sentence to improve performance. |
| Approach: | They propose to use a graph convolutional network to build a context graph without dependency parsers. |
| Outcome: | The proposed approach improves neural RE methods without dependency parsers on English benchmark datasets. |
Copied to clipboard
| Challenge: | Existing neural approaches to transliterate names from English to Arabic are limited and focus on leveraging the phonemic association between English and Arabic. |
| Approach: | They propose a model for English-Arabic transliteration using a memory module modeling the phonemic association between English and Arabic to guide the transliterations process. |
| Outcome: | The proposed model improves on EANames corpus, which better represents names in the general public than linked Wikipedia entries that are always names of famous people. |
Copied to clipboard
| Challenge: | a large number of language models struggle to handle disfluencies, authors say . when a speaker hesitates, interrupts themselves, repeats or corrects words, or abandons phrases, it can make their speech fragmented. |
| Approach: | They propose to use disfluent queries to “clean” spontaneous speech . they propose to apply disfluencies to models that use different types of speech repairs . |
| Outcome: | The proposed model improves on a reading comprehension task using disfluent queries . the results suggest that disfluencies can improve model performance, rather than their removal . |
Copied to clipboard
| Challenge: | a study examines whether readers can distinguish between two types of reading goals: information seeking and ordinary reading for comprehension. |
| Approach: | They propose a method to distinguish between two types of reading goals: information seeking and ordinary reading for comprehension. |
| Outcome: | The proposed model solves the reading goal-oriented task with the most accurate predictions in real time, the authors say . |
Copied to clipboard
| Challenge: | Multiple-choice MRC is one of the most studied tasks in MRC due to the convenience of evaluation and the flexibility of answer format. |
| Approach: | They propose to use multiple-choice MRC to explain a trained model and reveal how it arrives at the prediction by punishing illogical attributions. |
| Outcome: | The proposed method improves model performance without external information and model structure change without any external information. |
Copied to clipboard
| Challenge: | Large Language Models (LMs) have achieved state-of-the-art performance on many NLP benchmarks. |
| Approach: | They propose to decompose a hard question into simpler questions that are easier for models to answer. |
| Outcome: | The proposed approach significantly improves model performance (24% for GPT3 and 29% for RoBERTa-SQuAD along with a symbolic calculator) by decomposing a hard question into simpler questions that are easier for models to answer. |
Copied to clipboard
| Challenge: | a recent study found that models prefer acceptable inputs over acceptable ones. |
| Approach: | They find that model judgements are generally robust when placed in randomly sampled linguistic contexts, but unstable when contexts match the test stimuli in syntactic structure. |
| Outcome: | The proposed model performance improves when contexts match syntactic structure, and declines when they are unacceptable. |
Copied to clipboard
| Challenge: | Semi-supervised learning (SSL) is a popular setting to make use of unlabelled data . Currently, there are two popular approaches to make effective use of the unlabelled datasets . |
| Approach: | They compare semi-supervised learning (SSL) and task-adaptive pre-training (TAPT) they find TAPT is a stronger and more robust SSL learner, even when using just a few hundred unlabelled samples . |
| Outcome: | The proposed methods improve model performance across different NLP tasks and data sizes. |
Copied to clipboard
| Challenge: | Existing pruning and quantization algorithms could compress LLMs to 4 bits without retraining. |
| Approach: | They propose a method to prune models using a minimal set of critical weights . they compare the method to a 2:4 structured sparsity method . |
| Outcome: | The proposed method outperforms other methods on LLAMA2 and OPT models while maintaining the efficiency of the compression. |
Copied to clipboard
| Challenge: | In speech translation, multimodal data to address limitations of individual modalities has shown significant effectiveness. |
| Approach: | They propose a cross-modal model which supports three input modalities for speech, text and fused speech-text. |
| Outcome: | The proposed model achieves an average of 34.0 BLEU on MuST-C, GigaST and newstest benchmark. |
Copied to clipboard
| Challenge: | Existing methods focus on distantly-labeled rationales, ignoring the potential important non-rationale words and not distinguishing the importance of different rationale words. |
| Approach: | They propose two novel auxiliary loss functions to make better use of distantly-labeled rationales, which encourage models to maintain their focus on important words beyond labeled rationals (PINs) and alleviate redundant training on non-helpful rationale (NoIRs). |
| Outcome: | The proposed methods outperform existing methods on two representative classification tasks while maintaining the ability to spread focus to other unlabeled important words. |
Copied to clipboard
| Challenge: | Recent studies show that pre-trained models are beneficial to Chinese Word Segmentation (CWS). However, these models lack task-specific prior segmentation knowledge. |
| Approach: | They propose a pre-trained Chinese word segmentation model MetaSeg which incorporates meta learning into a multi-criteria pre-training task. |
| Outcome: | Empirical results show that MetaSeg can achieve new state-of-the-art performance on twelve widely-used CWS datasets and significantly improve model performance in low-resource settings. |
Copied to clipboard
| Challenge: | Multimodal retrieval-augmented generation (mRAG) pipelines are becoming more popular for vision-centric tasks. |
| Approach: | They propose to use a visual asset as a trigger to leak data from a model prompt. |
| Outcome: | The proposed pipelines can connect private datasets and improve model performance, but they can leak private information from them. |
Copied to clipboard
| Challenge: | Existing approaches to sentence representation learning often encounter semantic inconsistencies and feature suppression. |
| Approach: | They propose a method for generating syntactically aligned negative (SAN) samples using a semantic importance-aware Masked Language Model (MLM) approach. |
| Outcome: | The proposed method produces negative samples with substantial textual overlap with the original sentences while conveying different meanings. |
Copied to clipboard
| Challenge: | Recent advances in in-context learning (ICL) have limited customization and inadequate error coverage. |
| Approach: | They propose a method to retrieve in-context principles from mistakes to improve model performance. |
| Outcome: | The proposed framework enhances model performance when applied to various prompting strategies. |
Copied to clipboard
| Challenge: | Existing large language models have limited abilities to solve deductive reasoning problems . performance differences between conditions do not improve overall performance . |
| Approach: | They investigate whether several large language models can solve a deductive reasoning problem in their conventional form. |
| Outcome: | The proposed models can solve a classic type of deductive reasoning problem in their conventional form. |
Copied to clipboard
| Challenge: | Recent studies have shown that effective filters can be created by utilising Large Language Models to synthetically label data, which is then used to train smaller neural models for filtering purposes. |
| Approach: | They extend this approach to languages beyond English to train neural models for filtering purposes. |
| Outcome: | The proposed approach is effective at filtering parallel text for translation quality and filtering for domain specificity. |
Copied to clipboard
| Challenge: | Prior methods producing useful task rankings are infeasible for large source pools . Embedding space maps (ESMs) reduce execution time and disk space usage . |
| Approach: | They introduce Embedded Space Maps (ESMs) that approximate the effect of fine-tuning a language model. |
| Outcome: | The proposed method reduces execution time and disk space usage by 10 and 278, respectively, while retaining high selection performance. |
Copied to clipboard
| Challenge: | Existing models of layout reading order do not convey the complete reading order information in the layout. |
| Approach: | They propose to model layout reading order as ordering relations over layout elements . they propose a reading-order-relation-enhancing pipeline to improve model performance . |
| Outcome: | The proposed model outperforms existing models on a visual-rich document dataset and on eight cross-domain VrD-IE/QA tasks without targeted optimization. |
Copied to clipboard
| Challenge: | Information extraction (IE) tasks require a limited number of example instructions to achieve effective performance. |
| Approach: | They propose two strategies to find spurious associations in large language models (LLMs) they use forward label extension and backward label validation to leverage extended labels to improve model performance. |
| Outcome: | The proposed methods improve performance on Chinese and English datasets and 9.55%, 11.42%, and 21.27% in F1 scores on SciERC, ACE05, and DuEE datasets. |
Copied to clipboard
| Challenge: | In this paper we compare the performance of our deep transformer based language models to existing models. |
| Approach: | They present a set of BERT and ELECTRA based German language models, GBERT and GELECTRE. |
| Outcome: | The proposed models outperform the previous best models on NER and classification tasks but are prohibitively large for many. |
Copied to clipboard
| Challenge: | Abstractive summarization systems have a severe mismatch between training and inference, i.e., exposure bias. |
| Approach: | They propose a multi-level contrastive learning framework for abstractive summarization and a tailored sparse decoder self-attention pattern to bridge the gap between training and inference. |
| Outcome: | The proposed framework outperforms the state-of-the-art models on two summarization datasets while adding relatively low overhead. |
Copied to clipboard
| Challenge: | Using sparse autoencoders, we explore how bilingual language models develop complex internal representations. |
| Approach: | They employ sparse autoencoders to analyze bilingual language models' internal representations. |
| Outcome: | The proposed method integrates decomposed representations from a fully trained model into a mid-training model. |
Copied to clipboard
| Challenge: | Existing studies on cognitive distortion have limited generalizability and performance of models in large-scale and cross-linguistic contexts. |
| Approach: | They propose a multi-task learning model based on teacher student architecture solution which improves generalization performance. |
| Outcome: | The proposed model improves generalizability and interpretability of the proposed model. |
Copied to clipboard
| Challenge: | Aspect-based sentiment analysis (ABSA) has seen success with English texts, but real-world social media interactions often involve multiple languages. |
| Approach: | They propose a framework for cross-lingual ABSA that incorporates code-switched bilingual sentences into the language discriminator and consistency training modules to enhance cross-linguistic alignment. |
| Outcome: | The proposed framework achieves cross-lingual sentence-level and aspect-level alignment, aligning features of aspect terms in different contextual environments. |
Copied to clipboard
| Challenge: | Existing methods for instruction selection rely on external models or rules, overlooking the intrinsic association between pre-trained model and instruction data. |
| Approach: | They propose a method that utilizes noise injection to identify the quality of instruction data without relying on external models. |
| Outcome: | The proposed method outperforms the model trained on the entire dataset and established baselines. |
Copied to clipboard
| Challenge: | Existing benchmarks for multimodal large language models do not capture real-world clinical complexity. |
| Approach: | They evaluate multilingual, multimodal multimodal models of clinical cases with up to 7 distinct visual clinical evidence types per case. |
| Outcome: | The proposed model outperforms human models on differential diagnosis (DDx) generation and final diagnosis (FDx) selection. |
Copied to clipboard
| Challenge: | a new study examines the potential of retrieval-augmented generation (RAG) with foundation models to enhance expert-level reasoning. |
| Approach: | They introduce PhoPile, a high-quality multimodal dataset specifically designed for Olympiad-level physics. |
| Outcome: | The proposed model can be used to solve Olympiad-level physics problems. |
Copied to clipboard
| Challenge: | Knowledge graph embedding (KGE) is an important task for many downstream applications. |
| Approach: | They propose to use self-knowledge distillation to learn a low-dimensional model from a pre-trained high-dimensional one. |
| Outcome: | The proposed model can improve model performance while maintaining lightweight structure. |
Copied to clipboard
| Challenge: | Existing approaches to scaling test-time compute rely on static compute allocation or sample from fixed generation distributions. |
| Approach: | They propose a test-time compute allocation framework that jointly adapts where computation is spent and how generation is performed. |
| Outcome: | The proposed approach outperforms baselines while consuming less inference-time compute. |