Papers with RoBERTa

197 papers
Masked Measurement Prediction: Learning to Jointly Predict Quantities and Units from Textual Context (2022.findings-naacl)

Copied to clipboard

Challenge: Current benchmarks do not evaluate numeracy of pretraining language models on measurements.
Approach: They propose a new task where a model learns to reconstruct a number with its associated unit given masked text.
Outcome: The proposed model significantly underperforms pre-trained model with baselines and ablations.
Reading Comprehension as Natural Language Inference:A Semantic Analysis (2020.starsem-1)

Copied to clipboard

Challenge: In recent past, Natural language Inference (NLI) has gained significant attention, but its true impact has not been well studied.
Approach: They propose to transform a large RACE dataset into an NLI model and compare it to a state-of-the-art model.
Outcome: The proposed model outperforms the previous model on a question-answer concatenation form and a coherent entailment form.
INSIGHTBUDDY-AI: Medication Extraction and Entity Linking using Pre-Trained Language Models and Ensemble Learning (2025.naacl-srw)

Copied to clipboard

Challenge: InsightBuddy-AI is a system for extracting medication mentions and their associated attributes.
Approach: They propose a system for extracting medication mentions and their associated attributes . they use stacked and voting ensembles built upon pre-trained language models .
Outcome: The proposed system outperforms fine-tuned models in the extraction of medication mentions and associated attributes.
Have Attention Heads in BERT Learned Constituency Grammar? (2021.eacl-srw)

Copied to clipboard

Challenge: Recent pre-trained language models have gained great success in many tasks, but what they have learned, and when they perform well remain unknown.
Approach: They employ the syntactic distance method to extract implicit constituency grammar from attention weights of attention heads of BERT and RoBERTa.
Outcome: The proposed models induce some grammar types much better than baselines, suggesting some heads act as a proxy for constituency grammar.
When Choosing Plausible Alternatives, Clever Hans can be Clever (D19-60)

Copied to clipboard

Challenge: Pretrained language models have shown large improvements in the commonsense reasoning benchmark COPA, but recent work has identified superficial cues in benchmark datasets which are predictive of the correct answer.
Approach: They propose an extension of COPA that does not suffer from easy-to-exploit single token cues and exploits them.
Outcome: The proposed extension of COPA does not suffer from easy-to-exploit single token cues.
CharBERT: Character-aware Pre-trained Language Model (2020.coling-main)

Copied to clipboard

Challenge: Pre-trained language models (PLMs) construct word representations at subword level with Byte-Pair Encoding (BPE) or its variations . but these methods split a word into subword units and make it incomplete and fragile .
Approach: They propose a character-aware pre-trained language model to tackle OOV problems . they construct contextual word embedding for each token from sequential character representations .
Outcome: The proposed model improves on the existing models on multiple NLP benchmarks.
Improving Numeracy by Input Reframing and Quantitative Pre-Finetuning Task (2023.findings-eacl)

Copied to clipboard

Challenge: Innumeracy is a problem in pretrained language models, but it is not discussed in this paper . Numerals are an indispensable part of narratives and provide much fine-grained information.
Approach: They propose a method to solve innumeracy in pretrained language models by exploring the notation of numbers.
Outcome: The proposed method improves performance in three benchmark datasets containing quantitative-related tasks.
UnMASKed: Quantifying Gender Biases in Masked Language Models through Linguistically Informed Job Market Prompts (2024.eacl-srw)

Copied to clipboard

Challenge: Language models (LMs) often include societal biases encoded in the human-produced datasets used for their training.
Approach: They evaluated six prominent language models: BERT, RoBERTa, DistilBERT, BERT- multilingual, XLM-RoBERT and DistilberT- multilinguistic.
Outcome: The results show that the models generated by the models were stereotypically gendered and with a reduced bias in multilingual variants.
AdapterHub: A Framework for Adapting Transformers (2020.emnlp-demos)

Copied to clipboard

Challenge: AdapterHub framework enables dynamic “stiching-in” of pre-trained adapters for different tasks and languages.
Approach: They propose a framework that allows dynamic "stiching-in" of pre-trained adapters for different tasks and languages.
Outcome: The proposed framework allows dynamic “stiching-in” of pre-trained adapters for different tasks and languages.
YATO: Yet Another deep learning based Text analysis Open toolkit (2023.emnlp-demo)

Copied to clipboard

Challenge: YATO is an open-source toolkit for text analysis with deep learning . it supports free combinations of three types of widely used features .
Approach: They introduce YATO, an open-source toolkit for text analysis with deep learning.
Outcome: YATO is an open-source toolkit for text analysis with deep learning . the toolkit supports free combinations of three types of widely used features .
Do Syntactic Probes Probe Syntax? Experiments with Jabberwocky Probing (2021.naacl-main)

Copied to clipboard

Challenge: a study of neural language models shows that syntactic probes do not properly isolate syntax.
Approach: They show that syntactic probes do not properly isolate syntax . they train two probes trained on normal data and find they perform worse .
Outcome: The proposed method outperforms the baseline models on the most popular models, but their lead is reduced by 53%.
To Clarify or not to Clarify: A Comparative Analysis of Clarification Classification with Fine-Tuning, Prompt Tuning, and Prompt Engineering (2024.naacl-srw)

Copied to clipboard

Challenge: Xu et al., 2019) show that pre-trained language model fine-tuning and prompt tuning are better than manual prompt engineering for clarification identification.
Approach: They propose to use pre-trained language model fine-tuning, prompt tuning and manual prompt engineering to model clarification identification.
Outcome: The proposed model outperforms pre-trained language model fine-tuning, prompt tuning and manual prompt engineering on the task of clarification identification.
Do we need Label Regularization to Fine-tune Pre-trained Language Models? (2023.eacl-main)

Copied to clipboard

Challenge: Knowledge Distillation (KD) is a label regularization technique that can be replaced with lighter teacher-free variants such as the label-smoothing technique.
Approach: They propose to use knowledge distillation to train student models by deploying the teacher network during training.
Outcome: The proposed method can be replaced with lighter teacher-free variants on PLMs with more than 600 distinct trials and ran each configuration five times.
jiant: A Software Toolkit for Research on General-Purpose Text Understanding Models (2020.acl-demos)

Copied to clipboard

Challenge: jiant is an open source toolkit for conducting multitask and transfer learning experiments on English NLU tasks.
Approach: They introduce jiant, an open source toolkit for conducting multitask and transfer learning experiments on English NLU tasks.
Outcome: The proposed toolkit reproduces published performance on GLUE and SuperGLUE tasks.
The Microsoft Toolkit of Multi-Task Deep Neural Networks for Natural Language Understanding (2020.acl-demos)

Copied to clipboard

Challenge: MT-DNN is an open-source natural language understanding toolkit . it allows researchers and developers to train customized deep learning models .
Approach: They present MT-DNN, an open-source natural language understanding toolkit . it is designed to facilitate rapid customization for a broad spectrum of NLU tasks . MT supports multi-task knowledge distillation, which can substantially compress a deep neural model without significant performance drop.
Outcome: The proposed model can significantly compress a large model without significant performance drop.
Newspaper Signaling for Crisis Prediction (2024.naacl-demo)

Copied to clipboard

Challenge: Existing systems for detecting crisis-related signals are limited due to unstructured data, media, and cultural bias, and multiple languages.
Approach: They propose a model for multi-lingual and open-domain newspaper signaling for detecting crisis-related indicators in newspaper articles.
Outcome: The proposed model can detect crisis-related indicators in multiple languages and can be used in open crisis domains in real-time.
TeluguNER: Leveraging Multi-Domain Named Entity Recognition with Deep Transformers (2022.acl-srw)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a successful and well-researched problem in English due to the availability of resources.
Approach: They propose to use two annotated NER datasets for the Telugu language . they compare the finetuned Telugus model with the existing model in NER .
Outcome: The proposed models outperform existing models on a large dataset of 38,363 sentences on telugu and other languages.
NSIT@NLP4IF-2019: Propaganda Detection from News Articles using Transfer Learning (D19-50)

Copied to clipboard

Challenge: In this paper, we describe our approach and system description for NLP4IF 2019 Workshop: Shared Task on Fine-Grained Propaganda Detection.
Approach: They propose to use document Embeddings and LSTM to detect whether a sentence contains a propagandistic agenda.
Outcome: The proposed approach ranked 21st in the NLP4IF 2019 Workshop: Shared Task on Fine-Grained Propaganda Detection.
Calibration of Pre-trained Transformers (2020.emnlp-main)

Copied to clipboard

Challenge: Pre-trained Transformers dominate benchmark tasks but use a large number of self-attention heads across many layers in a way that is difficult to unpack.
Approach: They analyze pre-trained Transformer models' posterior probabilities to determine whether they are calibrated for three tasks: natural language inference, paraphrase detection, and commonsense reasoning.
Outcome: The models are calibrated in-domain and out-of-domain, and their calibration error out-domain can be as much as 3.5x lower.
Disambiguating Emotional Connotations of Words Using Contextualized Word Representations (2024.starsem-1)

Copied to clipboard

Challenge: BERT, RoBERTa, XLNet, and GPT-2 models effectively discern emotional connotations of words, demonstrating superior performance and greater resilience against biases.
Approach: They propose to use contextualized word representations to examine how words can be used to distinguish emotional connotations across contexts.
Outcome: The proposed models show that they can distinguish emotional connotations of words in different contexts.
Naturalistic Causal Probing for Morpho-Syntax (2023.tacl-1)

Copied to clipboard

Challenge: Existing methods for probing are limited and lack understanding of their limitations and weaknesses.
Approach: They propose a strategy for input-level intervention on naturalistic sentences . they use morpho-syntactic features of a sentence to intervene on the rest of the sentence .
Outcome: The proposed approach allows for input-level intervention on naturalistic sentences while keeping the rest of the sentence unchanged.
The Shape of Vulnerability: How Adversarial Perturbations Reshape the Topology of Language Model Latent Spaces (2026.acl-srw)

Copied to clipboard

Challenge: Large Language Models (LLMs) have unprecedented capabilities, but they pose security concerns . current adversarial attacks exploit vulnerabilities in the embedding space of language models, allowing attackers to bypass safety guardrails and cause significant harmful consequences.
Approach: They propose to use topological data analysis to characterize how adversarial perturbations act on text inputs by computing persistent homology metrics from attention maps across different model architectures.
Outcome: The proposed visualizations show that adversarial perturbations alter higher-dimensional topological features in ways that distinguish them from clean, non-adversarial inputs.
Comparing Probabilistic, Distributional and Transformer-Based Models on Logical Metonymy Interpretation (2020.aacl-main)

Copied to clipboard

Challenge: Logical metonymies are type clashes between an event-selecting verb and an entity-denoting noun . they are typically interpreted by inferring a hidden event on the basis of contextual cues .
Approach: They propose to use probabilistic and distributional models to model logical metonymy interpretation . they compare models with the best Transformer-based models and some traditional distributional ones .
Outcome: The proposed models perform well on a complex scenario, but low performance on some datasets suggests that logical metonymy is still a challenging phenomenon for computational modeling.
Exploring Semantics in Pretrained Language Model Attention (2024.starsem-1)

Copied to clipboard

Challenge: Abstract Meaning Representations (AMRs) encode the semantics of sentences in the form of graphs.
Approach: They propose to use attention heads of two LMs to detect semantic relations encoded in AMRs.
Outcome: The proposed models detect semantic relations without fine tuning, using both unsupervised and supervised learning techniques.
Do ever larger octopi still amplify reporting biases? Evidence from judgments of typical colour (2022.aacl-short)

Copied to clipboard

Challenge: Language models trained on text-only corpora have no direct access to the physical world and thus suffer from reporting bias.
Approach: They investigate reporting bias from the perspective of colour in larger language models such as PaLM and GPT-3.
Outcome: The proposed models outperform smaller models on the basis of colour and more closely track human judgements than smaller models.
‘Am I the Bad One’? Predicting the Moral Judgement of the Crowd Using Pre–trained Language Models (2022.lrec-1)

Copied to clipboard

Challenge: Existing studies on NLP touch upon moral contexts in text.
Approach: They construct a dataset that can be used for moral judgement tasks on a popular reddit subreddit.
Outcome: The proposed model passes moral judgements on posts from a popular reddit subreddit . it shows that the model can be fine tuned and improves across the datasets .
T-MAD: Target-driven Multimodal Alignment for Stance Detection (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for Multimodal Stance Detection struggle with generalizing to unseen targets and handling modality inconsistencies.
Approach: They propose a multimodal stability detection model which captures target-specific relationships and balances modality contributions by iterative reasoning.
Outcome: Experiments on the MMSD and MultiClimate datasets show that the proposed model outperforms state-of-the-art models with optimal results achieved using RoBERTa, ViT, and an iterative depth of 5.
Cross-lingual Evidence Improves Monolingual Fake News Detection (2021.acl-srw)

Copied to clipboard

Challenge: Existing methods focused on one language and do not use multilingual information.
Approach: They propose a new technique based on cross-lingual evidence that can be used for fake news detection . they compared their proposed technique with strong baselines on two datasets of general-topic news .
Outcome: The proposed technique improves existing methods and can be used on real and fake news datasets.
Farewell to Aimless Large-scale Pretraining: Influential Subset Selection for Language Model (2023.findings-acl)

Copied to clipboard

Challenge: Pretrained language models have achieved remarkable success in various natural language processing tasks.
Approach: They propose to use end-task knowledge to select a tiny subset of pretraining corpus to influence performance.
Outcome: The proposed model outperforms pretrained models on eight datasets covering four domains with 0.45% of the data and a three-orders-of-magnitude lower computational cost.
Beyond Reptile: Meta-Learned Dot-Product Maximization between Gradients for Improved Single-Task Regularization (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to improve generalization of neural models use a small component of the gradient for maximizing dot-product between batches.
Approach: They propose to use a finite differences first-order algorithm to calculate a gradient from dot-product of gradients and regularize it.
Outcome: The proposed method outperforms previous approaches of Reptile and MAML when used as a regularization technique.
On the Robustness of Reading Comprehension Models to Entity Renaming (2022.naacl-main)

Copied to clipboard

Challenge: SpanBERT model is more robust than RoBERTa, despite having similar accuracy on unperturbed test data.
Approach: They propose a pipeline to replace entity names with names from a variety of sources.
Outcome: The proposed model performs worse when entities are renamed, the authors show . SpanBERT, which is pretrained with span-level masking, is more robust than RoBERTa .
Sentiment Analysis of Yelp Review Dataset: A Comparative Study of Machine Learning Methods (2026.acl-srw)

Copied to clipboard

Challenge: Existing methods for sentiment analysis are inconsistent and require manual processing.
Approach: They use natural language processing and machine learning to classify Yelp reviews' sentiments.
Outcome: The proposed model outperforms other models on Yelp reviews.
Chemical Language Understanding Benchmark (2023.acl-industry)

Copied to clipboard

Challenge: CLUB datasets are used to facilitate NLP research in the chemical industry.
Approach: They introduce a benchmark dataset called CLUB to facilitate NLP research in the chemical industry.
Outcome: The CLUB datasets are a new benchmark dataset for NLP in the chemical industry.
Context-Aware Transformer Pre-Training for Answer Sentence Selection (2023.acl-short)

Copied to clipboard

Challenge: Existing approaches to perform Answer Sentence Selection (AS2) using only the candidate sentence are sub-optimal.
Approach: They propose to use pre-trained transformers to perform contextual AS2 fine-tuning . they propose to apply pre-training objectives to local contextual AS2.
Outcome: The proposed methods improve baseline AS2 accuracy by up to 8% on some datasets.
Beat the AI: Investigating Adversarial Human Annotation for Reading Comprehension (2020.tacl-1)

Copied to clipboard

Challenge: Innovations in annotation methodologies have been a catalyst for Reading Comprehension (RC) datasets and models.
Approach: They propose to use a model-in-the-annotation-loop approach to train adversarial models in three different settings to explore reproducibility of the adversarial effect, transfer from data collected with varying model- in-the loop strengths, and generalization to data collected without a modeling model.
Outcome: The proposed approach achieves 39.9F1 on questions it cannot answer when trained on SQUAD, but lower than when trained using RoBERTa itself (41.0F1).
Using Adversarial Attacks to Reveal the Statistical Bias in Machine Reading Comprehension Models (2021.acl-short)

Copied to clipboard

Challenge: Pre-trained language models have achieved human-level performance on many Machine Reading Comprehension (MRC) tasks, but it remains unclear whether these models truly understand language or answer questions by exploiting statistical biases in datasets.
Approach: They propose a method to attack MRC models by exposing statistical biases in a RACE dataset and propose an augmented training method that can greatly reduce models’ statistical bias.
Outcome: The proposed method can reduce models’ statistical biases from human-level performance to chance-level.
Light-Weight Hallucination Detection using Contrastive Learning for Conditional Text Generation (2025.acl-srw)

Copied to clipboard

Challenge: Existing methods for hallucination detection are limited to the scenario where we can access the LLMs that have generated the outputs.
Approach: They propose a hallucination detection method that uses contrastive learning to pull faithful outputs and input contexts together while pushing hallucinous outputs apart.
Outcome: The proposed method outperforms GPT-4o prompting in binary hallucination detection.
ColD Fusion: Collaborative Descent for Distributed Multitask Finetuning (2023.acl-long)

Copied to clipboard

Challenge: Pretraining has been shown to scale well with compute, data size and data diversity.
Approach: They propose a method that provides benefits of multitask learning but leverages distributed computation . they propose 'coldfusion' can create synergistic loop where finetuned models can be "recycled"
Outcome: The proposed method outperforms RoBERTa and previous multitask models on 35 datasets.
oLMpics-On What Language Model Pre-training Captures (2020.tacl-1)

Copied to clipboard

Challenge: Recent success of pre-trained language models has spurred widespread interest in their capabilities.
Approach: They propose an evaluation protocol that includes zero-shot evaluation and no fine-tuning . they propose to compare the learning curve of a fine- tuned LM to the learning of multiple controls .
Outcome: The proposed evaluation protocol compares the learning curve of a fine-tuned LM to the learning of multiple controls.
Linguistically Grounded Analysis of Language Models using Shapley Head Values (2025.findings-naacl)

Copied to clipboard

Challenge: Existing methods for probing language models for morphosyntactic constructions are not well understood . language models gain knowledge of grammatical phenomena during pretraining, but exactly how this knowledge is encoded is not well established.
Approach: They propose a method for probing language models via Shapley Head Values . they use a BLiMP dataset to test their method on linguistic constructions based on a Shaply Head Value method .
Outcome: The proposed method can be used to investigate linguistic knowledge in language models . it shows that attention heads responsible for processing related linguistic phenomena cluster together .
Is ChatGPT a Financial Expert? Evaluating Language Models on Financial Natural Language Processing (2023.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have revolutionized general natural language preprocessing tasks, but their performance in financial domains is not evaluated comprehensively.
Approach: They propose a framework to evaluate financial language models on financial tasks . they compare performance of auto-encoding language models and ChatGPT .
Outcome: The proposed framework compares the performance of auto-encoding language models and the LLM ChatGPT on financial tasks.
Enhancing Job Evaluation with Data Augmentation and Text Classification (2026.acl-industry)

Copied to clipboard

Challenge: Recruiters rely on job titles, role descriptions, and responsibility levels to determine job grades and salary structures.
Approach: They propose to semi-automate job evaluation by fine-tuning a RoBERTa model for classification and using Gemini to generate synthetic job descriptions for rare job titles.
Outcome: The proposed method improves job evaluation by boosting consistency and speeding up workflows.
A Closer Look at Linguistic Knowledge in Masked Language Models: The Case of Relative Clauses in American English (2020.coling-main)

Copied to clipboard

Challenge: Despite the high performance of transformer-based language models, we still lack understanding of the kind of linguistic knowledge they learn and rely on.
Approach: They evaluate three transformer-based language models and test their grammatical and semantic knowledge by sentence-level probing, diagnostic cases, and masked prediction tasks.
Outcome: The models capture grammatical and semantic knowledge, but they lack model-specific weaknesses especially on semantic knowledge.
Probing Across Time: What Does RoBERTa Know and When? (2021.findings-emnlp)

Copied to clipboard

Challenge: Current approaches to natural language processing rely on fixed artifacts such as language models . current studies have focused on how these models acquire and demonstrate knowledge .
Approach: They apply probing techniques to examine how language models acquire knowledge . they aim to inform future work on more efficient pretraining and understanding dependencies .
Outcome: The proposed model learns linguistic abstractions, factual and commonsense knowledge, and reasoning abilities fast, stably, and robustly across domains.
Auto-Debias: Debiasing Masked Language Models with Automated Biased Prompts (2022.acl-long)

Copied to clipboard

Challenge: Existing methods to mitigate human-like biases in pretrained language models are based on external corpora and require a distribution alignment loss to mitigate them.
Approach: They propose an automatic method to mitigate biases in pretrained language models by searching for biased prompts such that cloze-style completions are the most different with respect to different demographic groups.
Outcome: The proposed method reduces biases in pretrained language models, including gender and racial bias, and improves fairness of the models.
Adaptive Structure Induction for Aspect-based Sentiment Analysis with Spectral Perspective (2023.findings-emnlp)

Copied to clipboard

Challenge: incorporating structure information can enhance the performance of aspect-based sentiment analysis.
Approach: They propose to use pre-trained language models to induct latent structures from a spectrum perspective.
Outcome: The proposed model shortens Aspects-sentiment Distance and improves structure induction ability.
From Hero to Zéroe: A Benchmark of Low-Level Adversarial Attacks (2020.aacl-main)

Copied to clipboard

Challenge: Adversarial attacks are label-preserving modifications to inputs of machine learning classifiers designed to fool machines but not humans.
Approach: They propose to use a dataset to test the robustness of future NLP models to identify low-level adversarial attacks that are less realistic in typical applications such as social media.
Outcome: The proposed dataset provides a benchmark for testing robustness of future more human-like NLP models.
Robust Explanations for User Trust in Enterprise NLP Systems (2026.acl-industry)

Copied to clipboard

Challenge: Existing studies on explanation stability under real user noise are limited . decoder LLMs produce significantly more stable explanations than encoder baselines .
Approach: They propose a black-box robustness evaluation framework for token-level explanations based on leave-one-out occlusion . they propose to operationalize explanation robustness with top-token flip rate under realistic perturbations at multiple severity levels .
Outcome: The proposed framework is compared with baseline models and encoder and decoder families.
When Do You Need Billions of Words of Pretraining Data? (2021.acl-long)

Copied to clipboard

Challenge: Pretrained language models (LMs) are dominated by models that can encode billions of words.
Approach: They use classifier probing, information-theoretic probing and unsupervised relative acceptability judgments to evaluate model ability.
Outcome: The proposed models require only about 10M to 100M words to learn to encode most syntactic and semantic features.
CoCoLM: Complex Commonsense Enhanced Language Model with Discourse Relations (2022.findings-acl)

Copied to clipboard

Challenge: Large-scale pre-trained language models have demonstrated strong knowledge representation ability, but struggle with complex commonsense knowledge that involves multiple eventualities.
Approach: They propose to help pre-trained language models better incorporate complex commonsense knowledge that involves multiple eventualities.
Outcome: The proposed model can learn to use the memorized knowledge for different tasks and achieve outstanding performance on many downstream natural language processing (NLP) tasks.
Winnowing Knowledge for Multi-choice Question Answering (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing reasoning models suffer from noises in retrieved knowledge . encoding methods that use commonsense knowledge are less effective .
Approach: They propose a method which conducts interception and soft filtering to reduce noise . they use commonsense knowledge from Wikipedia and ConceptNet to encode questions and options .
Outcome: The proposed method improves on commonsense question answering tasks compared to baselines . it is able to conduct interception and soft filtering to shield the encoder from noise .
DPTDR: Deep Prompt Tuning for Dense Passage Retrieval (2022.coling-1)

Copied to clipboard

Challenge: Recent studies show that prompt tuning is unfriendly for industrial deployment in dense retrieval tasks.
Approach: They propose to apply prompt tuning to dense retrieval tasks to reduce deployment cost . they propose to use retrieval-oriented intermediate pretraining and unified negative mining .
Outcome: The proposed method outperforms state-of-the-art models on MS-MARCO and Natural Questions.
Unsupervised Energy-based Adversarial Domain Adaptation for Cross-domain Text Classification (2021.findings-acl)

Copied to clipboard

Challenge: Extensive experiments on multidomain sentiment classification and yes/no question-answering classification are conducted.
Approach: They propose an unsupervised energy-based adversarial domain adaptation framework that maps the text sequences from both source and target domains to a feature space.
Outcome: The proposed framework improves on multidomain sentiment classification and Yes/No question-answering classification.
Mask More and Mask Later: Efficient Pre-training of Masked Language Models by Disentangling the [MASK] Token (2022.findings-emnlp)

Copied to clipboard

Challenge: Large-scale pre-trained MLMs can be used to generalize well to a wide range of tasks.
Approach: They propose to append [MASK]s at a later layer to reduce sequence length for earlier layers.
Outcome: The proposed method outperforms RoBERTa for 6 out of 8 GLUE tasks on average by 0.4%.
NLQuAD: A Non-Factoid Long Question Answering Data Set (2021.eacl-main)

Copied to clipboard

Challenge: Existing data sets for document-level question answering are limited in their ability to detect short text and require multiple-sentence descriptive answers and opinions.
Approach: They introduce a new data set with baseline methods for non-factoid long question answering . they compare BERT, RoBERTa, and Longformer models to establish baseline performances .
Outcome: Experimental results show that Longformer outperforms the other architectures but human evaluations show that it is far behind the human upper bound.
Always Keep your Target in Mind: Studying Semantics and Improving Performance of Neural Lexical Substitution (2020.coling-main)

Copied to clipboard

Challenge: Lexical substitution is a powerful technology used in various NLP applications . it generates plausible words that can replace a given word in a textual context .
Approach: They propose to use a large-scale comparative study to compare lexical substitution methods . they compare existing and new methods using word sense induction datasets .
Outcome: The proposed methods improve competitive results by incorporating information about the target word into the models.
Fast, Effective, and Self-Supervised: Transforming Masked Language Models into Universal Lexical and Sentence Encoders (2021.emnlp-main)

Copied to clipboard

Challenge: Existing studies have shown that pretrained Masked Language Models are not effective as universal lexical and sentence encoders off-the-shelf, i.e., without further task-specific fine-tuning on NLI, sentence similarity, or paraphrasing tasks using annotated task data.
Approach: They propose a contrastive learning technique which turns pretrained MLMs into effective universal lexical and sentence encoders without additional data.
Outcome: The proposed technique can turn MLMs into effective universal lexical and sentence encoders even without additional data.
Evaluating the Robustness of Neural Language Models to Input Perturbations (2021.emnlp-main)

Copied to clipboard

Challenge: High-performance neural language models have achieved state-of-the-art results on a wide range of NLP tasks, but results for common benchmark datasets often do not reflect model reliability and robustness when applied to noisy, real-world data.
Approach: They propose to implement character-level and word-level perturbation methods to simulate scenarios in which input texts may be slightly noisy or different from the data distribution on which NLP systems were trained.
Outcome: The proposed methods simulate scenarios in which input texts may be slightly noisy or different from the data distribution on which NLP systems were trained.
How much pretraining data do language models need to learn syntax? (2021.emnlp-main)

Copied to clipboard

Challenge: Pretraining methods are convenient, but expensive in terms of time and resources.
Approach: They investigate the impact of pretraining data size on the syntactic capabilities of RoBERTa by using syntaktic structural probes to determine whether models pretrained on more data encode a higher amount of syntastic information.
Outcome: The proposed models perform better on part-of-speech tagging, dependency parsing and paraphrase identification.
K-Adapter: Infusing Knowledge into Pre-Trained Models with Adapters (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods for injecting knowledge into pre-trained models are inconsistent and can flush out knowledge when multiple kinds of knowledge are injected.
Approach: They propose a framework that retains the original parameters of pre-trained models fixed and supports the development of versatile knowledge-infused models.
Outcome: The proposed framework retains the original parameters of the pre-trained model fixed and supports the development of versatile knowledge-infused models.
AutoMeTS: The Autocomplete for Medical Text Simplification (2020.coling-main)

Copied to clipboard

Challenge: Semi-automated text simplification approaches can be used to simplify text faster and at a higher quality.
Approach: They propose to use autocomplete to simplify medical texts using aligned English Wikipedia sentences and pretrained neural language models to analyze the additional context.
Outcome: The proposed model outperforms the best individual model by 2.1% and achieves a word prediction accuracy of 64.52%.
Leveraging a Bilingual Dictionary to Learn Wolastoqey Word Representations (2022.lrec-1)

Copied to clipboard

Challenge: Existing word embeddings for lowresource languages require large corpora of running text to learn high quality representations.
Approach: They leverage a bilingual dictionary to learn Wolastoqey word embeddings by encoding their corresponding English definitions into vector representations using pretrained English word and sequence representation models.
Outcome: The proposed model outperforms baseline models without language-specific training or fine-tuning.
Knowledge Augmentation Enhances Token Classification for Recipe Understanding (2026.eacl-long)

Copied to clipboard

Challenge: Using entity type-specific and knowledge-augmented token classification, we achieve state-of-the-art (SOTA) results on 5 out of 7 benchmark recipe datasets, significantly outperforming traditional token classification methods.
Approach: They propose an entity type-specific and knowledge-augmented token classification framework to improve encoder models’ performance on recipe texts.
Outcome: The proposed model outperforms traditional token classification methods on 5 out of 7 recipe datasets and is the largest annotated food-related dataset to date.
Discourse-Level Representations can Improve Prediction of Degree of Anxiety (2023.acl-short)

Copied to clipboard

Challenge: Anxiety disorders are the most common of mental illnesses, but little is known about how to detect them from language.
Approach: They propose to use discourse-level information in addition to lexical-level large language model embeddings to evaluate the utility of a lexico-discourse model.
Outcome: The proposed model outperforms models based on state-of-the-art contextual embeddings and uses discourse patterns of causal explanations significantly more than models derived from Sentence-BERT and DiscRE, and is comparable to psychological models.
ERNIE-Gram: Pre-Training with Explicitly N-Gram Masked Language Modeling for Natural Language Understanding (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods to model coarse-grained linguistic information do not integrate coarse-gram information into pre-training.
Approach: They propose an explicitly n-gram masking method to enhance integration of coarse-grained linguistic information into pre-training.
Outcome: The proposed method outperforms existing models on English and Chinese text corpora and fine-tunes on 19 downstream tasks.
Better Robustness by More Coverage: Adversarial and Mixup Data Augmentation for Robust Finetuning (2021.findings-acl)

Copied to clipboard

Challenge: Pretrained language models perform poorly under adversarial attacks due to the large search space.
Approach: They propose a method to cover a much larger proportion of the attack search space by adding textual adversarial examples during training.
Outcome: The proposed method covers a much larger proportion of the attack search space.
Enhancing Natural Language Representation with Large-Scale Out-of-Domain Commonsense (2022.findings-acl)

Copied to clipboard

Challenge: Using commonsense in text understanding tasks can cause catastrophic forgetting due to domain discrepancy . previous methods of using textual descriptions as extra input information cannot apply to large-scale commonsensing.
Approach: They propose to use out-of-domain commonsense to enhance text representation . they propose to integrate commonsensense descriptions into large-scale models .
Outcome: The proposed model can integrate commonsense descriptions and enhance them to the target text representation without pre-training on large-scale unsupervised corpora.
MelBERT: Metaphor Detection via Contextualized Late Interaction using Metaphorical Identification Theories (2021.naacl-main)

Copied to clipboard

Challenge: Existing studies have developed computational models to recognize metaphorical words in sentences.
Approach: They propose a model that leverages contextualized word representation and linguistic metaphor identification theories to detect whether the target word is metaphorical.
Outcome: The proposed model outperforms baseline models on four benchmark datasets . it leverages contextualized word representation and linguistic metaphor identification theories to detect whether the target word is metaphorical.
DIA-HARM: Dialectal Disparities in Harmful Content Detection Across 50 English Dialects (2026.acl-long)

Copied to clipboard

Challenge: Current disinformation detection systems are predominantly developed and evaluated on Standard American English (SAE) . however, their robustness to dialectal variation is unexplored.
Approach: They propose a benchmark for evaluating disinformation detection robustness across 50 English dialects . they use multi-value's linguistically-grounded transformations to introduce D-CUBE (Dialectal Disinformation Detection Corpus)
Outcome: The proposed model outperforms zero-shot LLMs in human-written dialects while AI-generated content remains stable.
Hence, Socrates is mortal: A Benchmark for Natural Language Syllogistic Reasoning (2023.findings-acl)

Copied to clipboard

Challenge: SylloBase is a benchmark for syllogistic reasoning, a critical capability widely required in natural language understanding tasks, such as text entailment and question answering.
Approach: They propose to use a benchmark to learn syllogistic reasoning on a set of templates and to use them to generate and understand slogisms.
Outcome: The proposed benchmark covers a complete taxonomy of syllogism reasoning patterns, and contains both automatically and manually constructed samples.
ScoNe: Benchmarking Negation Reasoning in Language Models With Fine-Tuning and In-Context Learning (2023.acl-short)

Copied to clipboard

Challenge: Negation is a ubiquitous but complex linguistic phenomenon that poses a significant challenge for NLP systems.
Approach: They propose a benchmark that measures how well models handle natural language negation . they extend ScoNe-NLI to embed negation reasoning in short narratives .
Outcome: The proposed model can reason about negation, but struggles to do so on NLI examples outside of its core pretraining regime.
Syntactically Aware Cross-Domain Aspect and Opinion Terms Extraction (2020.coling-main)

Copied to clipboard

Challenge: Supervised-learning approaches fail to scale across domains where labeled data is lacking.
Approach: They propose a method for incorporating external linguistic knowledge into a self-attention mechanism coupled with a transformer-based model.
Outcome: The proposed method enables leveraging syntactic knowledge from transformer-based models to bridge the gap between domains.
BERTAC: Enhancing Transformer-based Language Models with Adversarially Pretrained Convolutional Neural Networks (2021.acl-long)

Copied to clipboard

Challenge: Existing models of NLP are fading away, but new ones are needed to maintain their dominance.
Approach: They propose a method to pretrain a CNN using Wikipedia data and integrate it with standard TLMs.
Outcome: The proposed method outperforms the original ALBERT on GLUE tasks and achieves similar performance to SOTA on open-domain QA tasks.
Masking as an Efficient Alternative to Finetuning for Pretrained Language Models (2020.emnlp-main)

Copied to clipboard

Challenge: Extensive evaluations of masking BERT, RoBERTa, and DistilBERT on eleven diverse NLP tasks show that our binary masked language models encode information necessary for solving downstream tasks.
Approach: They propose an efficient method of utilizing pretrained language models where selective binary masks are learned instead of finetuning.
Outcome: Extensive evaluations of masking BERT, RoBERTa, and DistilBERT on eleven diverse NLP tasks show that the proposed method yields comparable performance to finetuning, but has a much smaller memory footprint when multiple tasks need to be solved.
What do tokens know about their characters and how do they know it? (2022.naacl-main)

Copied to clipboard

Challenge: Pre-trained language models that use subword tokenization schemes can succeed at a variety of language tasks that require character-level information.
Approach: They propose to use word tokenization schemes to probe what word pieces encode . they show that larger models can encode character-level information .
Outcome: The proposed models can encode character-level information and perform better on non-Latin alphabets.
Decoding Symbolism in Language Models (2023.acl-long)

Copied to clipboard

Challenge: Existing language models can be used to decode symbolism, but they are biased in pre-trained corpora.
Approach: They propose to use language models to decode symbols by re-ranking pre-trained models.
Outcome: The proposed framework shows that pre-trained models can mitigate the bias and improve performance to be on par with human models.
MiniLMv2: Multi-Head Self-Attention Relation Distillation for Compressing Pretrained Transformers (2021.findings-acl)

Copied to clipboard

Challenge: Existing work on deep self-attention distillation for natural language processing tasks is limited by computational resources and latency.
Approach: They generalize deep self-attention distillation in MINILM by using only self- attention relation distillation for taskagnostic compression of pretrained Transformers.
Outcome: The proposed model outperforms the state-of-the-art in a multilingual and multilingual teacher model.
Towards a Comprehensive Understanding and Accurate Evaluation of Societal Biases in Pre-Trained Transformers (2021.naacl-main)

Copied to clipboard

Challenge: Existing pre-trained language models are not fully considered for societal biases . pre-training models can be useful for many NLP tasks, but they can be harmful when used at scale.
Approach: They investigate gender and racial bias across pre-trained language models . they evaluate bias within pre-trainers using three metrics: WEAT, sequence likelihood, and pronoun ranking.
Outcome: The proposed model fails to detect gender and racial biases in pre-trained models . the model is ineffective when word embedding, demonstrating the need for more robust bias testing in transformers.
Neural Deepfake Detection with Factual Structure of Text (2020.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to deepfake detection typically represent documents with coarse-grained representations, but they struggle to capture factual structures of documents.
Approach: They propose a graph-based model that captures factual structures of documents for deepfake detection.
Outcome: The proposed model improves strong base models built with RoBERTa on two public deepfake datasets.
Knowledge Rumination for Pre-trained Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies have shown that pre-trained language models lack the capacity to handle knowledge-intensive tasks alone.
Approach: They propose a new paradigm to help pre-trained language models utilize latent knowledge without retrieving it from external corpus.
Outcome: The proposed paradigm can be applied to pre-trained language models without retrieving external knowledge from the corpus.
MTLS: Making Texts into Linguistic Symbols (2024.emnlp-main)

Copied to clipboard

Challenge: In linguistics, all languages can be considered as symbolic systems . most work overlooks the properties of languages as symbol systems - aaron et al., 1989).
Approach: They propose a method to make texts into linguistic symbols to improve multilingual capability . they use a pre-training method to replace pre-trained language models with a vocabulary map .
Outcome: The proposed method improves multilingual capabilities on multilingual tasks using BERT and RoBERTa as the backbone.
TCFLE-8: a Corpus of Learner Written Productions for French as a Foreign Language and its Application to Automated Essay Scoring (2023.emnlp-main)

Copied to clipboard

Challenge: Automated Essay Scoring (AES) aims to automatically assess the quality of essays.
Approach: They propose to use a corpus of 6.5k essays collected in the context of the Test de Connaissance du Français (TCF) certification exam to foster the development of AES for French.
Outcome: The proposed system can assess the quality of essays in a language certification exam using a corpus of 6.5k essays collected in the TCFLE-8 exam.
Real-Time Visual Feedback to Guide Benchmark Creation: A Human-and-Metric-in-the-Loop Workflow (2023.eacl-main)

Copied to clipboard

Challenge: Recent research has shown that language models exploit ‘artifacts’ in benchmarks to solve tasks, rather than learning them, leading to inflated model performance.
Approach: They propose a benchmark creation paradigm for NLP that focuses on guiding crowdworkers and provides realtime visual feedback to improve sample quality.
Outcome: The proposed paradigm decreases effort, frustration, mental, and temporal demands of crowdworkers and analysts, while increasing the performance of both user groups.
Probing the Category of Verbal Aspect in Transformer Language Models (2024.findings-naacl)

Copied to clipboard

Challenge: a particular challenge is posed by ”alternative contexts” where either the perfective or the imperfective aspect is suitable grammatically and semantically.
Approach: They investigate how pretrained language models encode the grammatical category of verbal aspect in Russian.
Outcome: The proposed model has high predictive uncertainty about aspect in alternative contexts, the authors show .
Learning to Ignore Adversarial Attacks (2023.eacl-main)

Copied to clipboard

Challenge: Despite the strong performance of current NLP models, they can be brittle against adversarial inputs.
Approach: They propose a rationale model that explicitly learns to ignore adversarial tokens . their approach leads to sizable improvements in robustness over baseline models .
Outcome: The proposed model outperforms data augmentation with adversarial examples and closes the gap between model performance and an attacked test set.
Retrofitting Light-weight Language Models for Emotions using Supervised Contrastive Learning (2023.emnlp-main)

Copied to clipboard

Challenge: a novel retrofitting method to induce emotion aspects into pre-trained language models is proposed . the models are computationally less expensive and open, but do not capture affective aspects of human communication well.
Approach: They propose a retrofitting method to induce emotion aspects into pre-trained language models . they retrofit text fragments exhibiting similar emotions into pretrained networks .
Outcome: The proposed method produces emotion-aware text representations for sentiment analysis and sarcasm detection tasks.
On the Interplay Between Fine-tuning and Sentence-level Probing for Linguistic Knowledge in Pre-trained Transformers (2020.findings-emnlp)

Copied to clipboard

Challenge: linguistic knowledge encoded in pre-trained contextual embeddings is poorly understood . fine-tuning can be used to investigate the representations of pre-train models .
Approach: They propose to investigate fine-tuning of contextualized embedding models through sentence-level probing.
Outcome: The proposed method improves probing accuracy for three pre-trained models.
Stress Test Evaluation of Transformer-based Models in Natural Language Understanding Tasks (2020.lrec-1)

Copied to clipboard

Challenge: Existing models are weak and take advantage of failures and errors in datasets to improve performance.
Approach: They evaluate three Transformer-based models in Natural Language Inference and Question Answering tasks to see if they are more robust or have the same flaws as their predecessors.
Outcome: The proposed models outperform recurrent neural network models to stress tests on both NLI and QA tasks.
Blockwise Self-Attention for Long Document Understanding (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in pre-training and fine-tuning methods have drastically reshaped the landscape of natural language processing research.
Approach: They propose a lightweight BERT model that introduces sparse block structures into the attention matrix to reduce memory consumption and training/inference time.
Outcome: The proposed model uses 18.7-36.1% less memory and 12.0-25.1% more time to learn compared to an advanced BERT-based model, RoBERTa.
Perturbations in the Wild: Leveraging Human-Written Text Perturbations for Realistic Adversarial Attack and Defense (2022.findings-acl)

Copied to clipboard

Challenge: ANTHRO extracts over 600K human-written text perturbations and leverages them for realistic adversarial attacks.
Approach: They propose an adversarial text manipulation algorithm that inductively extracts over 600K human-written text perturbations and leverages them for realistic adversarials.
Outcome: The proposed algorithm outperforms the TextBugger baseline with an increase of 50% and 40% in terms of semantic preservation and stealthiness when evaluated by layperson and professional human workers.
Linguistic Cues for LLM-based Implicit Discourse Relation Classification (2026.findings-eacl)

Copied to clipboard

Challenge: Large language models (LLMs) have been successful in many NLP tasks, but they struggle to capture subtle lexical relations between arguments.
Approach: They propose a strategy that enriches arguments with explicit lexical-level semantic cues before fine-tuning.
Outcome: The proposed approach improves F1 scores in cross-domain scenarios by more than 10 points compared to baselines.
DyLoRA: Parameter-Efficient Tuning of Pre-trained Models using Dynamic Search-Free Low-Rank Adaptation (2023.eacl-main)

Copied to clipboard

Challenge: Pre-training/fine-tuning of pre-training models has become more expensive and resource-hungry.
Approach: They propose a low-rank adaptation technique that trains LoRA blocks for a range of ranks instead of a single rank.
Outcome: The proposed method trains LoRA blocks for a range of ranks instead of a single rank . it can train dynamic search-free models with DyLoRA at least 4 to 7 times faster than LoRA .
Masked Language Model Scoring (2020.acl-main)

Copied to clipboard

Challenge: Pretrained masked language models require finetuning for most tasks.
Approach: They evaluate pretrained masked language models out of the box via their pseudo-log-likelihood scores (PLLs) they attribute this success to PLL’s unsupervised expression of linguistic acceptability without a left-to-right bias, greatly improving on scores from GPT-2 .
Outcome: The proposed model outperforms autoregressive language models in a variety of tasks.
Simple but Challenging: Natural Language Inference Models Fail on Simple Sentences (2022.findings-emnlp)

Copied to clipboard

Challenge: Natural language inference (NLI) tasks are difficult to perform on large datasets . a small number of simple sentences can improve model performance, authors say .
Approach: They propose to use syntactically simple sentences to test the inference ability of NLI models.
Outcome: The proposed set of simple sentences shows that the models fine-tuned on MNLI and SNLI perform poorly on Simple Pair.
Chapter Ordering in Novels (2022.emnlp-main)

Copied to clipboard

Challenge: a major challenge in research on long-form narrative texts is the cost of annotation . authors propose a new task that reconstructs the original order of chapters in novels without the need for human annotation.
Approach: They propose a task that reconstructs the original order of chapters in novels given a random permutation of the text.
Outcome: The proposed task yields a Spearman correlation of 0.59 on the novel and challenging task, substantially above baseline.
ERICA: Improving Entity and Relation Understanding for Pre-trained Language Models via Contrastive Learning (2021.acl-long)

Copied to clipboard

Challenge: Existing pre-training objectives do not explicitly model relational facts in text . Experimental results show that ERICA can improve typical PLMs on several language understanding tasks, including relation extraction, entity typing and question answering.
Approach: They propose a contrastive learning framework ERICA to obtain a deep understanding of entities and relations in text.
Outcome: The proposed framework can improve PLMs on several language understanding tasks, especially under low-resource settings.
Syntax-BERT: Improving Pre-trained Transformers with Syntax Trees (2021.eacl-main)

Copied to clipboard

Challenge: Pre-trained language models like BERT achieve superior performances in various NLP tasks without explicit consideration of syntactic information.
Approach: They propose a plug-and-play framework that incorporates syntax trees into pre-trained Transformers.
Outcome: The proposed framework improves on pre-trained models on natural language understanding datasets and shows that it can be used to train pre-structured neural networks.
Social Commonsense Reasoning with Multi-Head Knowledge Attention (2020.findings-emnlp)

Copied to clipboard

Challenge: Social Commonsense Reasoning requires understanding of text, knowledge about social events and their pragmatic implications, as well as commonsense reasoning skills.
Approach: They propose a multi-head knowledge attention model that encodes semi-structured commonsense inference rules and learns to incorporate them in a transformer-based reasoning cell.
Outcome: The proposed model improves performance on two reasoning tasks that require different reasoning skills.
Enhancing Contextual Word Representations Using Embedding of Neighboring Entities in Knowledge Graphs (2022.coling-1)

Copied to clipboard

Challenge: Existing methods for pre-trained language models lack explicit grounding in real-world entities.
Approach: They propose a mechanism that integrates the structure of a KG into recent PLM architectures by generalizing the embeddings of neighboring entities.
Outcome: The proposed method improves a classification task, entity typing task and language comprehension tasks.
Efficient Transformer-based Large Scale Language Representations using Hardware-friendly Block Structured Pruning (2020.findings-emnlp)

Copied to clipboard

Challenge: Pretrained large-scale language models have been criticized for their limited weight storage and computational speed on hardware platforms.
Approach: They propose an efficient transformer-based large-scale language representation using hardware-friendly block structure pruning.
Outcome: The proposed model achieves 5.0x accuracy on GLUE benchmarks and 1.79x compression rate on DistilBERT.
NSP-BERT: A Prompt-based Few-Shot Learner through an Original Pre-training Task —— Next Sentence Prediction (2022.coling-1)

Copied to clipboard

Challenge: Recent studies have shown that using prompts to utilize language models to perform downstream tasks is more effective than using token-level methods such as PET.
Approach: They propose to use a BERT original pre-training task abandoned by RoBERTa and other models to construct a sentence-level prompt-based method that does not need to fix the length of the prompt or the position to be predicted.
Outcome: The proposed method performs better than PET and EFL on a BERT pre-training task and is comparable to other prompt-based methods.
Classifying Social Media Users before and after Depression Diagnosis via Their Language Usage: A Dataset and Study (2024.lrec-main)

Copied to clipboard

Challenge: Mental illness can negatively impact individuals’ quality of life as it is considered one of the causes of years lived with disability and it is related to high suicide rates.
Approach: They collect first dataset of textual posts by same users before and after being diagnosed with depression and build multiple predictive models based on Transformers and BERT.
Outcome: The proposed model can be used to detect depression and suicidal thoughts in users who are not diagnosed with depression or suicide.
RobBERT: a Dutch RoBERTa-based Language Model (2020.findings-emnlp)

Copied to clipboard

Challenge: Pre-trained language models have been dominating the field of natural language processing in recent years, and have led to significant performance gains for various complex natural language tasks.
Approach: They used a robustly optimized BERT approach to train a Dutch language model called RobBERT.
Outcome: The proposed model outperforms models trained on a single language on dozens of tasks and is available for further downstream NLP applications.
Modelling Context and Syntactical Features for Aspect-based Sentiment Analysis (2020.acl-main)

Copied to clipboard

Challenge: Existing approaches to aspect-based sentiment analysis do not fully leverage syntactical information.
Approach: They propose an end-to-end aspect-based sentiment analysis solution that integrates syntactical information with part-of-speech embeddings and dependency-based embeddables to enhance the performance of the aspect extractor.
Outcome: The proposed solution outperforms the state-of-the-art models on SemEval-2014 dataset in both subtasks.
Diagnosing the First-Order Logical Reasoning Ability Through LogicNLI (2021.emnlp-main)

Copied to clipboard

Challenge: Existing studies have focused on diagnosing LMs' reasoning abilities in natural language understanding tasks.
Approach: They propose a diagnostic method for first-order logic reasoning with a proposed benchmark, LogicNLI.
Outcome: The proposed method disentangles the target FOL reasoning from commonsense inference and can be used to diagnose LMs from four perspectives: accuracy, robustness, generalization, and interpretability.
On the Robustness of Language Encoders against Grammatical Errors (2020.acl-main)

Copied to clipboard

Challenge: Pre-trained language encoders are effective in facilitating downstream natural language processing tasks, but they often assume training and test corpora are clean and it is unclear how the models behave when confronted with noisy input.
Approach: They conduct adversarial attacks to simulate grammatical errors on clean text data.
Outcome: The proposed model performs better when confronted with natural grammatical errors than when faced with noisy input.
Mining Knowledge for Natural Language Inference from Wikipedia Categories (2020.findings-emnlp)

Copied to clipboard

Challenge: Accurate lexical entailment (LE) and natural language inference (NLI) tasks require expensive annotations.
Approach: They propose to pretrain Wikipedia categories for lexical entailment and natural language inference by pretraining them on WikiNLI and transferring them to other knowledge bases.
Outcome: The proposed model can improve strong baselines such as BERT and RoBERTa by pretraining on WikiNLI and transferring the models on downstream tasks.
Interpreting the Robustness of Neural NLP Models to Textual Perturbations (2022.findings-acl)

Copied to clipboard

Challenge: Modern Natural Language Processing models are sensitive to input perturbations and their performance can decrease when applied to noisy data.
Approach: They propose to explain the extent to which a model is affected by an unseen textual perturbation by the learnability of the perturbation.
Outcome: The proposed model is better at identifying a perturbation (higher learnability) but worse at ignoring it (lower robustness).
Comprehensive Study on German Language Models for Clinical and Biomedical Text Understanding (2024.lrec-main)

Copied to clipboard

Challenge: Pre-trained language models can struggle in specialized domains such as medicine . existing generalpurpose pre-tried models can be used and refined through further pre-training on domainspecific unlabeled data.
Approach: They pre-trained German medical language models on 2.4B tokens from translated public data and 3B token of German clinical data.
Outcome: The proposed models outperform clinical models on various downstream tasks in germany . the authors show that continuous pre-training can match or exceed clinical models trained from scratch .
How is BERT surprised? Layerwise detection of linguistic anomalies (2021.acl-long)

Copied to clipboard

Challenge: a number of studies have shown that transformer-based language models detect when a word is anomalous in context, but likelihood scores do not tell the cause of the anomaly.
Approach: They propose to use Gaussian models for density estimation at intermediate layers of three language models to evaluate grammaticality.
Outcome: The proposed method on BLiMP shows that language models employ different mechanisms to detect different types of linguistic anomalies.
Considering Nested Tree Structure in Sentence Extractive Summarization with Pre-trained Transformer (2021.emnlp-main)

Copied to clipboard

Challenge: Sentence extractive summarization shortens a document by selecting sentences for a summary while preserving its important contents.
Approach: They propose a nested tree-based extractive summarization model on RoBERTa that uses syntactic and discourse trees to represent sentences in a given document.
Outcome: The proposed model outperforms baseline models on the CNN/DailyMail dataset and achieves significantly better scores than the baseline models in terms of coherence and comparable scores to the state-of-the-art models.
Assessing the Syntactic Capabilities of Transformer-based Multilingual Language Models (2021.findings-acl)

Copied to clipboard

Challenge: Multilingual Transformer-based language models have been shown to be excellent learners in crosslingual transfer tasks.
Approach: They evaluate the syntactic generalization capabilities of BERT and RoBERTa models on English and Spanish tests.
Outcome: The proposed models perform well on English and Spanish tests, and the proposed tests are compared against models on the same language and models on two different languages.
Evaluating Pretrained Transformer-based Models on the Task of Fine-Grained Named Entity Recognition (2020.coling-main)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a fundamental task in Natural Language Processing (NLP).
Approach: They compare three transformer-based names to two non-transformer-based ones . they find transformer-derived models incrementally outperform non-tranformer models .
Outcome: The proposed models outperform the studied models in most domains with respect to the F1 score.
Merely Judging Metaphor is Not Enough: Research on Reasonable Metaphor Detection (2024.findings-emnlp)

Copied to clipboard

Challenge: Current metaphor detection tasks only provide labels without interpreting how to understand them.
Approach: They propose to improve the current metaphor detection task by using mainstream Large Language Models.
Outcome: The proposed model is based on the original sentence, target word, and usage . the model is then evaluated using manual evaluation .
Learning from Child-directed Speech in Two-language Scenarios: A French-English Case-Study (2026.findings-eacl)

Copied to clipboard

Challenge: a systematic study of compact language models with limited computational resources is challenging for many research contexts and real-world applications.
Approach: They extend BabyBERTa to English-French scenarios under strictly sizematched data conditions.
Outcome: The proposed model extends to English-French scenarios under sizematched data conditions . the results show context-dependent effects of multilingual training .
Empirical Evaluation of Pre-trained Transformers for Human-Level NLP: The Role of Sample Size and Dimensionality (2021.naacl-main)

Copied to clipboard

Challenge: In human-level NLP tasks, the number of observations is often smaller than the standard 768+ hidden state sizes of each layer within transformer-based language models.
Approach: They propose to use dimension reduction methods to fine-tune large models with limited data and to use pre-trained dimension reduction regimes to improve model performance.
Outcome: The proposed model outperforms other models in human-level NLP tasks with a pre-trained dimension reduction regime.
The Topic Confusion Task: A Novel Evaluation Scenario for Authorship Attribution (2021.findings-emnlp)

Copied to clipboard

Challenge: Autorship attribution is the problem of identifying the most plausible author of an anonymous text from a set of candidate authors.
Approach: They propose a topic confusion task where they switch the author-topic configuration between training and testing sets and propose attribution errors that are caused by the topic shift and by the features’ inability to capture the writing styles.
Outcome: The proposed task combines author-topic configuration with other features to lower topic confusion and higher attribution accuracy.
Multi-granularity Textual Adversarial Attack with Behavior Cloning (2021.emnlp-main)

Copied to clipboard

Challenge: Existing adversarial attack models are vulnerable to adversarials crafted by human-imperceptible perturbations.
Approach: They propose a multi-granularity adversarial attack model that generates high-quality adversarials with fewer queries to victim models.
Outcome: The proposed model generates high-quality adversarial samples with fewer queries to victim models compared to baseline models . the proposed model also reduces query times for black-box models that only output labels without confidence scores .
FLamE: Few-shot Learning from Natural Language Explanations (2023.acl-long)

Copied to clipboard

Challenge: Recent work has shown limited utility of natural language explanations in improving classification.
Approach: They propose a two-stage few-shot learning framework that generates explanations and fine-tunes a smaller model with generated explanations.
Outcome: The proposed framework increases inference accuracy over strong baselines, but human evaluation reveals that the majority of generated explanations does not adequately justify classification decisions.
Models In a Spelling Bee: Language Models Implicitly Learn the Character Composition of Tokens (2022.naacl-main)

Copied to clipboard

Challenge: Standard pre-trained language models do not see the characters that compose each token's string representation.
Approach: They probe the embedding layer of pretrained language models and show that models learn the internal character composition of whole word and subword tokens without seeing the characters coupled with the tokens.
Outcome: The embedding layers of RoBERTa and GPT2 hold enough information to accurately spell up to a third of the vocabulary and reach high character ngram overlap across all token types.
Self-supervised Cross-modal Pretraining for Speech Emotion Recognition and Sentiment Analysis (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to multimodal speech emotion recognition and sentiment analysis have not improved results due to their relatively simple fusion mechanisms and lack of proper cross-modal pretraining.
Approach: They propose a deep-fused audio-text bi-modal transformer with carefully designed cross-modal fusion mechanism and stage-wise cross-mod pretraining scheme to facilitate cross-modulation.
Outcome: The proposed method exceeds benchmarks on public IEMOCAP emotion and CMU-MOSEI sentiment datasets by a large margin.
Unlocking Continual Learning Abilities in Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to learning models (LMs) incorporate old task data or task-wise inductive bias into LMs, but old data and accurate task information are often unavailable or costly to collect.
Approach: They propose a rehearsal-free method that updates model parameters with large magnitudes . they found that the L1-normalized magnitude distribution is different when different task data is used .
Outcome: The proposed method improves accuracy and performance on four CL benchmarks.
Learning Semantic Structure through First-Order-Logic Translation (2024.findings-emnlp)

Copied to clipboard

Challenge: a recent study shows that transformer-based language models can confuse which predicates apply to which objects . a this is a crucial building block of semantic structure, but if an LM mixes up which objects have which property, it makes errors in reasoning .
Approach: They propose to use transformer-based language models to learn predicate argument structure from simple sentences.
Outcome: The proposed model can learn predicate argument structure from simple sentences.
How Effective is Task-Agnostic Data Augmentation for Pretrained Transformers? (2020.findings-emnlp)

Copied to clipboard

Challenge: Task-agnostic data augmentations have proven widely effective in computer vision, even on pretrained models.
Approach: They examine the effects of two types of task-agnostic data augmentation on pretrained transformers using 5 classification tasks and 6 datasets.
Outcome: The proposed techniques improve performance on 5 classification tasks, 6 datasets, and 3 variants of modern pretrained transformers.
Fight Fire with Fire: Fine-tuning Hate Detectors using Large Samples of Generated Hate Speech (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for hate speech detection are limited in size and lack of labeled datasets.
Approach: They employ pretrained language models to generate large amounts of hate speech sequences from available labeled examples.
Outcome: The proposed model improves generalization significantly and consistently within and across data distributions.
Conspiracy Theories and Where to Find Them on TikTok (2025.acl-long)

Copied to clipboard

Challenge: Existing studies on TikTok's potential to promote and amplify harmful content have not been conducted.
Approach: They analyze a longitudinal dataset of 1.5M videos shared in the U.S. over three years and evaluate the effects of TikTok’s Creativity Program for monetization.
Outcome: The proposed model achieves high precision in detecting harmful content, but its overall performance is comparable to fine-tuned traditional models such as RoBERTa.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks (D19-1)

Copied to clipboard

Challenge: Existing methods for finding similar sentences require multiple inferences . a modern GPU requires 65 hours to find the most similar pair in 10,000 sentences .
Approach: They propose a modification of the pretrained BERT network that uses siamese and triplet networks to derive semantically meaningful sentence embeddings.
Outcome: The proposed method outperforms existing methods on sentence-pair regression tasks.
Making Language Models Robust Against Negation (2025.naacl-long)

Copied to clipboard

Challenge: Negation is a semantic phenomenon that alters an expression to convey the opposite meaning.
Approach: They propose a self-supervised method to make language models more robust against negation by pre-training models.
Outcome: The proposed task outperforms the off-the-shelf versions on nine negation-related benchmarks.
Detecting Urgency Status of Crisis Tweets: A Transfer Learning Approach for Low Resource Languages (2020.coling-main)

Copied to clipboard

Challenge: We train monolingual and cross-lingual classifiers on the extracted features of tweets . we use a few state-of-the-art contextual embeddings to extract features of the tweets.
Approach: They propose to use tweets to train a dataset of English and two low-resource languages to train zero-shot transfer models.
Outcome: The proposed model performs well in English and in low-resource languages . the proposed model is based on state-of-the-art embeddings and semi-supervised methods .
Leveraging pre-trained language models for linguistic analysis: A case of argument structure constructions (2024.emnlp-main)

Copied to clipboard

Challenge: Argument structure constructions (ASCs) are lexicogrammatical patterns at the clausal level.
Approach: They evaluate the effectiveness of pre-trained language models in identifying argument structure constructions . they use supervised training with RoBERTa and prompt-guided annotation with GPT-4 .
Outcome: The proposed model outperforms the gold-standard model on three methods . the results show that the model performs better on gold-standardized data .
StereoSet: Measuring stereotypical bias in pretrained language models (2021.acl-long)

Copied to clipboard

Challenge: Existing literature on stereotypical biases in language models is limited . current evaluations focus on measuring bias without considering language modeling ability .
Approach: They propose to measure stereotypical biases in four domains: gender, profession, race, and religion . they compare stereotypical and language modeling ability of popular models like BERT, GPT-2, RoBERTa and XLnet .
Outcome: The proposed model shows strong stereotypical biases in gender, profession, race, and religion domains.
Hyperbolic Relevance Matching for Neural Keyphrase Extraction (2022.naacl-main)

Copied to clipboard

Challenge: Keyphrase extraction is a fundamental task in natural language processing that aims to extract a set of phrases with important information from a source document.
Approach: They propose a hyperbolic matching model to explore keyphrase extraction in hyperbolical space using word embeddings from RoBERTa to capture hierarchical syntactic and semantic structures.
Outcome: The proposed model outperforms the state-of-the-art models on six benchmark datasets and outperformed previous models.
Syntax-Enhanced Pre-trained Model (2021.acl-long)

Copied to clipboard

Challenge: Existing methods that use syntax of text in pre-training and fine-tuning suffer from discrepancy between the two stages.
Approach: They propose a model that utilizes the syntactic structure of text in pre-training and fine-tuning stages.
Outcome: The proposed model achieves state-of-the-art on six public benchmark datasets.
Compressing Transformer-Based Semantic Parsing Models using Compositional Code Embeddings (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing task-oriented semantic parsing models use BERT or RoBERTa as pretrained encoders.
Approach: They propose to learn compositional code embeddings to greatly reduce the sizes of BERT and RoBERTa encoders.
Outcome: The proposed model reduces the size of BERT and RoBERTa encoders while maintaining performance.
Rather a Nurse than a Physician - Contrastive Explanations under Investigation (2023.emnlp-main)

Copied to clipboard

Challenge: a recent study suggests that contrastive explanations are closer to how humans explain a decision than non-contrastive explanations.
Approach: They analyze four English text-classification datasets to determine whether humans explain in contrast to alternatives.
Outcome: The proposed explanations are closer to how humans explain a decision than non-contrastive explanations.
How does BERT’s attention change when you fine-tune? An analysis methodology and a case study in negation scope (2020.acl-main)

Copied to clipboard

Challenge: Recent work probing pre-trained language models for downstream tasks is difficult to explain . a growing body of research is devoted to understanding what linguistic properties these language models have acquired.
Approach: They propose a procedure and analysis method that takes a hypothesis of how a transformer-based model might encode a linguistic phenomenon and tests its validity.
Outcome: The proposed method tests a hypothesis that some attention heads will consistently attend from a word in negation scope to the negation cue.
Severing the Edge Between Before and After: Neural Architectures for Temporal Ordering of Events (2020.emnlp-main)

Copied to clipboard

Challenge: Existing models for temporal ordering of events rely on pretrained representations, transfer and multitask learning, and self-training techniques.
Approach: They propose a neural architecture and a set of training methods for ordering events by predicting temporal relations by pre-training models.
Outcome: The proposed models can predict temporal relations between two pairs of events within a span of text and identify temporal relationships between them.
How transfer learning impacts linguistic knowledge in deep NLP models? (2021.findings-acl)

Copied to clipboard

Challenge: Several researchers have shown that deep NLP models learn non-trivial amount of linguistic knowledge, captured at different layers of the model.
Approach: They propose to fine-tune pre-trained models towards downstream NLP tasks to capture linguistic knowledge.
Outcome: The proposed model is adapted to GLUE tasks and retains linguistic information in the network while forgetting it.
Acquiring Clean Language Models from Backdoor Poisoned Datasets by Downscaling Frequency Space (2024.acl-long)

Copied to clipboard

Challenge: Prior work attempts to mitigate backdoor learning during training LMs on poisoned datasets . backdoor attack poisons a small portion of training data by implanting specific text patterns .
Approach: They propose a multi-scale low-rank adaptive model that prioritizes learning of clean mapping . they propose radial scalings to reduce the success rate of diverse backdoor attacks .
Outcome: The proposed model outperforms baselines significantly in the frequency space . it reduces the success rate of diverse backdoor attacks to below 15% across datasets .
Contextualized and Generalized Sentence Representations by Contrastive Self-Supervised Learning: A Case Study on Discourse Relation Analysis (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods to learn contextualized and generalized sentence representations are limited by the size of manually annotated data.
Approach: They propose a method to learn contextualized and generalized sentence representations using contrastive self-supervised learning.
Outcome: The proposed method outperforms baseline methods based on BERT, XLNet, and RoBERTa in English and Japanese and outperformed strong baseline methods.
DSEE: Dually Sparsity-embedded Efficient Tuning of Pre-trained Language Models (2023.acl-long)

Copied to clipboard

Challenge: Conventional fine-tuning works through updating all of the parameters in the pre-trained model, but as the size of pre-train models grows, it can be time-consuming and computationally expensive.
Approach: They propose a framework for resource- and parameter-efficient fine-tuning by leveraging the sparsity prior in both weight updates and the final model weights.
Outcome: The proposed framework saves 25% inference FLOPs while maintaining competitive downstream performance.
Do CoNLL-2003 Named Entity Taggers Still Work Well in 2023? (2023.acl-long)

Copied to clipboard

Challenge: NER models trained on 20-year-old test set may not perform well on modern data.
Approach: They evaluate the generalization of over 20 different models trained on the CoNLL-2003 dataset . they find no evidence of performance degradation in pre-trained Transformers .
Outcome: The proposed model generalizations show that some models generalize well on new data while others do not.
SHIELD: Defending Textual Neural Networks against Multiple Black-Box Adversarial Attacks with Stochastic Multi-Expert Patcher (2022.acl-long)

Copied to clipboard

Challenge: Existing methods to defend textual neural network models against adversarial attacks often require retraining and retrain . e.g., BERT, RoBERTa require great time and computation resources.
Approach: They propose an algorithm that modifies and re-trains only the last layer of a textual NN and transforms it into a stochastic weighted ensemble of multi-expert prediction heads.
Outcome: The proposed algorithm outperforms existing models against black-box attacks by 15%–70% . the proposed algorithm is based on a novel algorithm from software engineering .
CICLe: Conformal In-Context Learning for Largescale Multi-Class Food Risk Classification (2024.findings-acl)

Copied to clipboard

Challenge: Contaminated or adulterated food poses a substantial risk to human health.
Approach: They present a dataset of 7,546 text messages describing public food recalls.
Outcome: The proposed model outperforms RoBERTa and XLM-R on classes with low support while reducing energy consumption.
Enabling Natural Zero-Shot Prompting on Encoder Models via Statement-Tuning (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) exhibit remarkable capabilities in zero-shot and few-shot settings, but they struggle with extending to few- shot and zero- shot settings due to their architectural design.
Approach: They propose a technique that models discriminative tasks as a set of finite statements and trains an encoder model to discriminate between the potential statements to determine the label.
Outcome: The proposed method achieves competitive performance compared to state-of-the-art LLMs with significantly fewer parameters.
Muppet: Massive Multi-task Representations with Pre-Finetuning (2021.emnlp-main)

Copied to clipboard

Challenge: Recent work shows gains from pre-training and fine-tuning that are multi-task . but it can be difficult to know which intermediate tasks will best transfer .
Approach: They propose a large-scale learning stage for pre-finetuning between pre-training and fine-tun.
Outcome: The proposed model improves performance on pretrained discriminators and generation models on a wide range of tasks while improving sample efficiency during fine-tuning.
Do Neural Language Models Inferentially Compose Concepts the Way Humans Can? (2024.lrec-main)

Copied to clipboard

Challenge: a new study shows that language models and humans may rely on different approaches to represent and compose lexical items across sentence structure.
Approach: They propose to use a dataset to test the performance of neural language models and humans on inferentially driven conceptual compositions.
Outcome: The proposed model elicits probability estimates for a noun in a minimally composed phrase . RoBERTa, BERT-large, and GPT-2 exhibited the closest resemblance to human responses .
Contrastive Code Representation Learning (2021.emnlp-main)

Copied to clipboard

Challenge: Recent work learns contextual representations of source code by reconstructing tokens from their context.
Approach: They propose a contrastive pre-training task that learns code functionality, not form . they propose scalable compilers that can generate variants of a program .
Outcome: The proposed task outperforms RoBERTa on an adversarial code clone detection benchmark by 39% AUROC.
EEE-QA: Exploring Effective and Efficient Question-Answer Representations (2024.lrec-main)

Copied to clipboard

Challenge: Current approaches to question answering rely on pre-trained language models like RoBERTa.
Approach: They propose a pooling approach that embeds all answer candidates with the question . they also propose enabling cross-reference between answer choices .
Outcome: The proposed methods improve throughput and memory efficiency with little sacrifice in performance.
Probing for Constituency Structure in Neural Language Models (2022.findings-emnlp)

Copied to clipboard

Challenge: Using standard probing techniques, we examine whether contextual neural language models implicitly learn syntactic structure.
Approach: They investigate to which extent contextual neural language models implicitly learn syntactic structure.
Outcome: The proposed model is able to represent constituents of different categories within the neuron activations of a LM such as RoBERTa with high performance even on manipulated data.
Transformers to Learn Hierarchical Contexts in Multiparty Dialogue for Span-based Question Answering (2020.acl-main)

Copied to clipboard

Challenge: Existing approaches to embedding in multiparty dialogues are poor for span-based question answering (QA)
Approach: They propose a novel approach to transformers that learns hierarchical representations in multiparty dialogue.
Outcome: The proposed model improves on the FriendsQA dataset by 3.8% and 1.4% over the two state-of-the-art models.
GhostBERT: Generate More Features with Cheap Operations for BERT (2021.acl-long)

Copied to clipboard

Challenge: Existing studies show that some parameters in pre-trained language models can be pruned away without severe accuracy degradation.
Approach: They propose a method which generates more features with very cheap operations from the remaining features and can be applied to unpruned BERT models to enhance their performance.
Outcome: Empirical results on the GLUE benchmark on three backbone models (i.e., BERT, RoBERTa and ELECTRA) verify the efficacy of the proposed method.
PreQuant: A Task-agnostic Quantization Approach for Pre-trained Language Models (2023.findings-acl)

Copied to clipboard

Challenge: Quantization is a viable solution for pre-trained language models, but most existing methods are task-specific and require customized training and quantization with a large number of trainable parameters.
Approach: They propose a "quantize before fine-tuning" framework that allows for quantization with a large number of trainable parameters on each individual task.
Outcome: The proposed framework is compatible with quantization-aware training and post-training quantization and corrects quantization errors.
Hire Me or Not? Examining Language Model’s Behavior with Occupation Attributes (2025.coling-main)

Copied to clipboard

Challenge: Large language models (LLMs) have been widely integrated into production pipelines due to their impressive performance across multiple tasks.
Approach: They construct a dataset using a standard occupation classification knowledge base and tested it on three families of LLMs.
Outcome: The proposed framework analyzes LLMs’ behavior with respect to gender stereotypes in the context of occupation decision making.
Aspect-based Document Similarity for Research Papers (2020.coling-main)

Copied to clipboard

Challenge: Traditional document similarity measures do not consider in what aspects two documents are similar.
Approach: They extend document similarity with aspect information by performing a pairwise document classification task.
Outcome: The proposed approach is best performing on 172,073 research paper pairs from the ACL Anthology and CORD-19 corpus.
ADEPT: An Adjective-Dependent Plausibility Task (2021.acl-long)

Copied to clipboard

Challenge: ADEPT is a large-scale semantic plausibility task that requires a significant degree of world knowledge and common-sense reasoning.
Approach: They propose a large-scale semantic plausibility task that pairs 16 thousand sentences with slightly modified versions obtained by adding an adjective to a noun.
Outcome: The proposed task is easier for humans (85% accuracy), but more difficult for transformer-based models (71% accuracy).
Evaluating Unsupervised Representation Learning for Detecting Stances of Fake News (2020.coling-main)

Copied to clipboard

Challenge: Using unsupervised representation learning, automated Fake News detection is a challenge for researchers.
Approach: They examine pre-trained language models with respect to their performance on two Fake News related data sets.
Outcome: The proposed models outperform the autoregression-based models on two Fake News related data sets.
FASTMATCH: Accelerating the Inference of BERT-based Text Matching (2020.coling-main)

Copied to clipboard

Challenge: Recent pre-trained language models have shown state-of-the-art accuracies in text matching.
Approach: They propose a BERT-based text matching model where representations and interactions are decoupled . they propose generating final matching scores using a lightweight attention network .
Outcome: Experiments show that the proposed model can achieve up to 100X speed-up to BERT and RoBERTa while keeping more up to 98.7% of the performance.
SIR-ABSC: Incorporating Syntax into RoBERTa-based Sentiment Analysis Models with a Special Aggregator Token (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to integrate syntactic dependency information into language models capture syntax . aspect-based sentiment classification tasks require a complex model to handle different aspects of a sentence .
Approach: They propose a method to incorporate syntactic dependency information directly into transformer-based language models for Aspect-Based Sentiment Classification.
Outcome: The proposed model outperforms existing models for aspect-based sentiment analysis tasks.
Leveraging Mental Health Forums for User-level Depression Detection on Social Media (2022.lrec-1)

Copied to clipboard

Challenge: Existing methods to detect depression on social media platforms are limited due to the vastness of social media content and the lack of linguistic features.
Approach: They propose to optimize the performance of user-level depression classification to lessen the burden on computational resources.
Outcome: The proposed system outperforms baselines across standard metrics for the task of depression detection in text.
Probing Pretrained Language Models for Lexical Semantics (2020.emnlp-main)

Copied to clipboard

Challenge: Existing studies have focused on morphosyntactic, semantic, and world knowledge, but it remains unclear to what extent LMs derive lexical type-level knowledge from words in context.
Approach: They propose to use multilingual and monolingual LMs to extract lexical type-level knowledge from words in context.
Outcome: The proposed models perform well across six typologically diverse languages and five lexical tasks.
Word-Level Coreference Resolution (2021.emnlp-main)

Copied to clipboard

Challenge: Recent coreference resolution models rely heavily on span representations to find coreference links between word spans.
Approach: They propose to consider coreference links between individual words rather than word spans and reconstruct the word span.
Outcome: The proposed model outperforms existing models on the OntoNotes benchmark while being highly efficient.
RobustLR: A Diagnostic Benchmark for Evaluating Logical Robustness of Deductive Reasoners (2022.emnlp-main)

Copied to clipboard

Challenge: Existing models that perform deductive reasoning on inputs containing rules and statements in the English natural language do not perform consistently on the RobustLR test set.
Approach: They propose a diagnostic benchmark that evaluates the robustness of language models to minimal logical edits in inputs and different logical equivalence conditions.
Outcome: The proposed models do not perform consistently on the RobustLR test set.
ConjNLI: Natural Language Inference Over Conjunctive Sentences (2020.emnlp-main)

Copied to clipboard

Challenge: Existing stress tests do not consider non-boolean usages of conjunctions and use templates . large-scale pre-trained models do not understand conjunctive semantics well, we find .
Approach: They propose a stress-test for natural language inference over conjunctive sentences where the premise differs from the hypothesis by conjunctions removed, added, or replaced.
Outcome: The proposed stress-test for natural language inference over conjunctive sentences is challenging . it finds that pre-trained models do not understand conjunction semantics well .
Advancing Semantic Textual Similarity Modeling: A Regression Framework with Translated ReLU and Smooth K2 Loss (2024.emnlp-main)

Copied to clipboard

Challenge: despite its efficiency, Sentence-BERT ignores the progressive nature of semantic relationships, despite a promising approach . contrastive learning methods have improved performance on renowned STS benchmarks, but they fail to leverage fine-grained information.
Approach: They propose a regression framework that categorizes text pairs as either semantically similar or dissimilar . they propose two loss functions: Translated ReLU and Smooth K2 Loss to bridge this gap .
Outcome: The proposed method achieves convincing performance across seven established STS benchmarks.
Contextualized Embeddings based Transformer Encoder for Sentence Similarity Modeling in Answer Selection Task (2020.lrec-1)

Copied to clipboard

Challenge: Word embeddings that consider context have attracted great attention for natural language processing tasks in recent years.
Approach: They propose two different approaches to integrate contextualized word embeddings with transformer encoders for sentence similarity modeling.
Outcome: The proposed model outperforms the feature-based approach on six datasets.
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension (2020.acl-main)

Copied to clipboard

Challenge: Recent work has shown gains by improving the distribution of masked tokens and the order in which mucked tokens are predicted.
Approach: They propose a denoising autoencoder for pretraining sequence-to-sequence models that corrupts text with an arbitrary noising function and learns a model to reconstruct the original text.
Outcome: The proposed model outperforms RoBERTa on GLUE and SQUAD and provides a 1.1 BLEU increase over a back-translation system for machine translation.
Attention-Enhancing Backdoor Attacks Against BERT-based Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing textual backdoor attacks focus on generating stealthy triggers or modifying model weights.
Approach: They propose a Trojan Attention Loss (TAL) which enhances the Trojan behavior by directly manipulating attention patterns.
Outcome: The proposed method improves the effectiveness of the backdoor attacks on different backbone models and tasks.
How Good Are LLMs at Out-of-Distribution Detection? (2024.lrec-main)

Copied to clipboard

Challenge: Out-of-distribution (OOD) detection is crucial for ensuring AI safety . large language models (LLMs) are becoming more prevalent due to their scale, pre-training objectives, and paradigms used for inference.
Approach: They propose to use large language models to investigate out-of-distribution (OOD) detection in machine learning.
Outcome: The proposed method outperforms other OOD detectors in zero-grad and fine-tuning scenarios.
Advancing Parameter Efficiency in Fine-tuning via Representation Editing (2024.acl-long)

Copied to clipboard

Challenge: Parameter Efficient Fine-Tuning (PEFT) has gained significant attention for its ability to achieve competitive results while updating only a small subset of trainable parameters.
Approach: They propose a new approach to fine-tuning neural models that scales and biases the representation produced at each layer.
Outcome: The proposed approach reduces the number of trainable parameters by a factor of 25,700 compared to full parameter fine-tuning and by . 32 compared with LoRA.
Structured Tuning for Semantic Role Labeling (2020.acl-main)

Copied to clipboard

Challenge: Recent neural network-driven semantic role labeling systems have shown impressive improvements in F1 scores.
Approach: They propose a framework to tune models using softened constraints only at training time.
Outcome: The proposed framework outperforms the baseline model with minimal training time and consistent improvements under low-resource scenarios.
The better your Syntax, the better your Semantics? Probing Pretrained Language Models for the English Comparative Correlative (2022.emnlp-main)

Copied to clipboard

Challenge: Construction Grammar posits constructions as the central building blocks of language . human-like performance of pretrained language models on many NLP tasks has been alleged .
Approach: They propose to use construction grammar to posit constructions as the central building blocks of language . they conduct experiments with three pretrained language models to examine their ability to classify and understand English comparative correlative .
Outcome: The proposed models are able to recognise the English comparative correlative (CC) but fail to use its meaning.
The Language of Trauma: Modeling Traumatic Event Descriptions Across Domains with Explainable AI (2024.findings-emnlp)

Copied to clipboard

Challenge: Psychological trauma can manifest following various distressing events, but studies focus on a single aspect of trauma, often neglecting the transferability of findings across different scenarios.
Approach: They propose a language model that fine-tunes a single aspect of trauma to better predict traumatic events across domains.
Outcome: The proposed model outperforms large language models on trauma-related datasets . it also outperformed models on court data, counseling conversations, and forum posts .
The Architectural Bottleneck Principle (2022.emnlp-main)

Copied to clipboard

Challenge: a recent study examined how much information a model's representations contain . a new approach to probing is to look exactly like the component .
Approach: They propose a new probing principle that aims to estimate how much information a model could extract from its representations.
Outcome: The proposed probes extract syntactic information from the representations of a neural network . the proposed probe is based on the architectural bottleneck principle .
A Curious Case of Searching for the Correlation between Training Data and Adversarial Robustness of Transformer Textual Models (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies show that fine-tuned textual transformer models are vulnerable to adversarial text perturbations.
Approach: They extract 13 different features representing a wide range of input fine-tuning corpora properties and use them to predict adversarial robustness of the fine- tuned models.
Outcome: The proposed framework can be used as an additional tool for robustness evaluation since it saves 30x-193x runtime compared to the traditional technique and can be easily used under adversarial training.
Attention-Focused Adversarial Training for Robust Temporal Reasoning (2022.lrec-1)

Copied to clipboard

Challenge: Current adversarial training approaches for NLP add adversarials to the embedding layer, ignoring other layers.
Approach: They propose an enhanced adversarial training algorithm for fine-tuning transformer-based language models . they add the adversarials to multiple hidden states or attention representations of the model layers .
Outcome: The proposed model improves performance on several temporal reasoning benchmarks and establishes new state-of-the-art results.
Pre-training Transformer Models with Sentence-Level Objectives for Answer Sentence Selection (2022.emnlp-main)

Copied to clipboard

Challenge: Existing models for answer sentence selection (AS2) are not yet available for AS2 .
Approach: They propose to incorporate paragraph-level semantics within and across documents to improve transformers for AS2 . they propose to use a dataset to predict whether two sentences are extracted from the same paragraph .
Outcome: The proposed model outperforms baseline models on public and industrial datasets on three public and one industrial dataset.
SparseGrad: A Selective Method for Efficient Fine-tuning of MLP Layers (2024.emnlp-main)

Copied to clipboard

Challenge: High-performance methods for parameter-efficient fine-tuning (PEFT) typically work with Attention blocks and overlook dense MLP blocks, which contain about half of the model parameters.
Approach: They propose a selective PEFT method that performs well on MLP blocks by converting layer gradients into a sparse structure and reducing the number of updated parameters.
Outcome: The proposed method outperforms LoRA and MeProp, robust state-of-the-art PEFT approaches.
Visually-augmented pretrained language models for NLP tasks without images (2023.acl-long)

Copied to clipboard

Challenge: Existing approaches to improve pre-trained language models lack visual commonsense and semantics.
Approach: They propose a visual-augmented approach to fine-tune pre-trained language models by using retrieved or generated images instead of relying on explicit images.
Outcome: The proposed approach outperforms baselines on ten tasks and consistently outperformed other approaches.
Pushing on Text Readability Assessment: A Transformer Meets Handcrafted Linguistic Features (2021.emnlp-main)

Copied to clipboard

Challenge: ML models with handcrafted features are linguistically explainable, expandable, and competent against the modern neural models.
Approach: They propose to combine traditional ML models with ML transformers to improve readability assessment by 99% accuracy.
Outcome: The proposed model achieves state-of-the-art (SOTA) accuracy on popular datasets.
Statement-Tuning Enables Efficient Cross-lingual Generalization in Encoder-only Models (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models excel in zero-shot and few-shot tasks, but their architecture makes them difficult to use.
Approach: They adapt Large Language Models (LLMs) for zero-shot generalization using Statement Tuning . they find encoders can achieve zero- shot cross-lingual generalization .
Outcome: The proposed model generalizes well across languages while being more efficient.
Exploring Large Language Models for Classical Philology (2023.acl-long)

Copied to clipboard

Challenge: Recent advances in NLP have led to the creation of powerful language models for many languages including Ancient Greek and Latin.
Approach: They propose to use encoder-only and encoder decoder architectures to create four models for Ancient Greek that vary along two dimensions for tasks of interest for Classical languages.
Outcome: The proposed models improve on existing models of Ancient Greek and Latin and provide a large pre-training corpus for Ancient Greek to support the creation of a larger, comparable model zoo for Classical Philology.
Misery Loves Complexity: Exploring Linguistic Complexity in the Context of Emotion Detection (2023.findings-emnlp)

Copied to clipboard

Challenge: a negative emotion is a cognitive bias that affects how we express thoughts and opinions online . a recent study shows that negative words generate more engagement and clicks than positive ones .
Approach: They propose to use readability and linguistic complexity metrics to better understand emotions . they propose to fine-tune three state-of-the-art transformers to detect emotions based on a dataset .
Outcome: The proposed model fails to predict emotions on complex texts, the authors show . they also show that more advanced models fail to predict complex texts .
CED: Comparing Embedding Differences for Detecting Out-of-Distribution and Hallucinated Text (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for detecting out-of-distribution (OOD) samples are limited due to their domain shift and computational limitations.
Approach: They propose a training-free method to detect out-of-distribution (OOD) samples . they theoretically validate that specific auxiliary and oracle samples improve this distinction .
Outcome: The proposed method improves the ability of pre-trained models to distinguish between ID and OOD samples in text classification and hallucination detection tasks.
Label-aware Hard Negative Sampling Strategies with Momentum Contrastive Learning for Implicit Hate Speech Detection (2024.findings-acl)

Copied to clipboard

Challenge: Existing models for implicit hate speech detection do not have significant advantage over cross-entropy loss-based learning.
Approach: They propose a label-aware hard negative sampling strategy that encourages the model to learn detailed features from hard negative samples instead of random batch.
Outcome: The proposed models outperform existing models for implicit hate speech detection both in- and cross-datasets.
Robust AI-Generated Text Detection by Restricted Embeddings (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches for artificial text detection are score-based and classifier-based . however, score-driven methods often rely on a score-derived score.
Approach: They investigate the ability of classifier-based detectors to transfer to unseen generators or semantic domains.
Outcome: The proposed methods improve the out-of-distribution classification score by up to 9% and 14%.
Stochastic Fine-Tuning of Language Models Using Masked Gradients (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are the dominant paradigm in Natural Language Processing but fine-tuning them for specific downstream tasks often requires updating a vast number of parameters.
Approach: They propose a method that selectively updates a small subset of parameters in each step of the tuning process.
Outcome: The proposed approach outperforms existing fine-tuning methods while updating merely **0.08**% of the model’s parameters.
On the Relationship between Skill Neurons and Robustness in Prompt Tuning (2024.lrec-main)

Copied to clipboard

Challenge: Prompt Tuning is a parameter-efficient finetuning method for pre-trained large language models (PLMs).
Approach: They propose to use RoBERTa to fine tune pre-trained large language models by finetuning only a small set of parameters to adjust for downstream tasks.
Outcome: The proposed method activates specific neurons in the transformer’s feed-forward networks that are highly predictive and selective for the given task.
Mining the uncertainty patterns of humans and models in the annotation of moral foundations and human values (2025.acl-long)

Copied to clipboard

Challenge: disagreement in annotation (HLV) is considered a constitutive feature of subjective tasks.
Approach: They investigate the relationship between disagreement in annotation and model uncertainty . they use linguistic features to calibrate models to HLV and uncertainty to analyze their impact on uncertainty.
Outcome: The proposed model uncertainty is calibrated to human label variation (HLV) the proposed model is calibrate to human labels, the authors show .
Small Encoders Can Rival Large Decoders in Detecting Groundedness (2025.findings-acl)

Copied to clipboard

Challenge: Large language models struggle to answer queries reliably when the provided context lacks information, often resorting to ungrounded speculation or internal knowledge.
Approach: They propose to detect whether a given query is grounded in a document provided in context before LLMs generate answers.
Outcome: The proposed model can generate answers that are grounded in the document provided in context while reducing inference latency by orders of magnitude.
Will LLMs Replace the Encoder-Only Models in Temporal Relation Classification? (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown promising performance in temporal reasoning tasks such as temporal question answering.
Approach: They propose to use large language models to detect temporal relations between events with in-context learning and lightweight fine-tuning approaches to assess their performance.
Outcome: The proposed models significantly underperform smaller encoder-only models based on RoBERTa in the Temporal Relation Classification task.
BiasWipe: Mitigating Unintended Bias in Text Classifiers through Model Interpretability (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods to mitigate unintended bias in social media platforms are re-training and adding extra parameters to the model.
Approach: They propose a technique to mitigate unintended bias in language models by pruning the neuron weights responsible for univ bias.
Outcome: The proposed technique achieves fairness by pruning the neuron weights responsible for unintended bias without loss of original performance.
GottBERT: a pure German Language Model (2024.emnlp-main)

Copied to clipboard

Challenge: Pre-trained language models have advanced natural language processing (NLP) despite the introduction of BERT, single-language models are still relevant.
Approach: They present a German singlelanguage RoBERT model pre-trained exclusively on the German portion of the OSCAR dataset.
Outcome: The GottBERT model outperforms the existing models on Named Entity Recognition and text classification tasks.
RoCoFT: Efficient Finetuning of Large Language Models with Row-Column Updates (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for fine-tuning large language models use full finetunation, but this is impractical as language models continue to scale up.
Approach: They propose a parameter-efficient fine-tuning method for large language models based on updating only a few rows and columns of the weight matrices in transformers.
Outcome: The proposed method gives comparable or better accuracies than state-of-the-art methods while being more memory and computation-efficient.
TacoERE: Cluster-aware Compression for Event Relation Extraction (2024.lrec-main)

Copied to clipboard

Challenge: Existing work on event relation extraction focuses on modeling the entire document . existing methods cannot handle long-range dependencies and information redundancy .
Approach: They propose a compression-then-extraction paradigm for event relation extraction . they propose document clustering for modeling event dependencies and then a cluster summarization method .
Outcome: The proposed method simplifies and highlights important text content of clusters for mitigating redundancy and event distance.
Detecting Legal Citations in United Kingdom Court Judgments (2025.emnlp-main)

Copied to clipboard

Challenge: citation detection in court judgments is challenging because of the complexity of legal language . citation analysis is critical for many legal applications, but the complexity is not always easy to solve.
Approach: They compare three different models for citation detection in court judgments using the Cambridge Law Corpus . they compare rulebased regular expressions, transformer-based encoders and large language models .
Outcome: The proposed model outperforms the existing models in the citation analysis and analysis of 190 court judgments.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations