Papers by Isabelle Augenstein

78 papers
Counterfactually Augmented Data and Unintended Bias: The Case of Sexism and Hate Speech Detection (2022.naacl-main)

Copied to clipboard

Challenge: sexism and hate speech detection models may be over-relying on core features . construct-driven CAD may induce models to ignore context in which core features are used .
Approach: They propose to use construct-driven and construct-agnostic CAD to reduce model bias . sexism and hate speech detection models are trained on counterfactually augmented data .
Outcome: Using a diverse set of CAD—construct-driven and construct-agnostic—reduces unintended bias.
Uncovering Probabilistic Implications in Typological Knowledge Bases (P19-1)

Copied to clipboard

Challenge: linguistic typology is concerned with mapping out the relationships between languages with structural and functional properties.
Approach: They propose a computational model which identifies known and new linguistic universals and uncovers them worthy of further linguistic investigation.
Outcome: The proposed model outperforms baselines and knowledge base baselines.
Fact Checking with Insufficient Evidence (2022.tacl-1)

Copied to clipboard

Challenge: Existing work on how to automate fact checking relies on information obtained from external sources.
Approach: They propose a fluency-preserving method for omitting information from the evidence at the constituent and sentence level and a diagnostic dataset for FC with omitted evidence.
Outcome: The proposed method improves evidence sufficiency prediction by 17.8 F1 score and 2.6 F1 scores.
DYNAMICQA: Tracing Internal Knowledge Conflicts in Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: LMs are useful in a variety of downstream applications from summarization to fact-checking, often relying on factual knowledge memorized during pre-training.
Approach: They use two knowledge conflict measures and a novel dataset DYNAMICQA to examine the effect of intra-memory conflict on LMs' ability to accept contextual knowledge.
Outcome: The proposed model can accept contextual knowledge with a higher degree of accuracy than models with fewer truth values.
What Can We Do to Improve Peer Review in NLP? (2020.findings-emnlp)

Copied to clipboard

Challenge: Traditionally, peer review is expected to act as a filter for high-quality, impactful work, but this does not hold in practice.
Approach: They argue that peer review is becoming increasingly spurious and that it is a problem for NLP . they propose a reproducibility checklist at EMNLP 2020 that could be used to ensure that papers are reproducible.
Outcome: The reproducibility checklist at EMNLP 2020 is the first step in that direction.
MultiFC: A Real-World Multi-Domain Dataset for Evidence-Based Fact Checking of Claims (D19-1)

Copied to clipboard

Challenge: Existing efforts to verify factual claims are limited by small datasets or artificially constructed datasets.
Approach: They propose to use the largest publicly available dataset of naturally occurring factual claims for automatic claim verification.
Outcome: The proposed model outperforms baseline models and evidence pages significantly.
Explaining Interactions Between Text Spans (2023.emnlp-main)

Copied to clipboard

Challenge: Existing highlight-based explanations focus on identifying individual important features or interactions only between adjacent tokens or tuples of tokens.
Approach: They propose a multi-annotator dataset of human span interaction explanations for NLU and FC.
Outcome: The proposed method compares human reasoning processes to those of a fine-tuned large language model.
From Phonology to Syntax: Unsupervised Linguistic Typology at Different Levels with Language Embeddings (N18-1)

Copied to clipboard

Challenge: linguistic typology is the classification of languages according to their linguistic properties.
Approach: They learn distributed language representations which can be used to predict typological properties on a massively multilingual scale.
Outcome: The proposed model can predict typological properties on a massively multilingual scale.
Multilingual Event Extraction from Historical Newspaper Adverts (2023.acl-long)

Copied to clipboard

Challenge: Developing NLP methods for historical corpora is difficult, as only domain experts can label them . off-the-shelf models are trained on modern language texts, rendering them weaker for historical documents .
Approach: They propose to use an annotated newspaper dataset to extract historical data from a novel domain of texts.
Outcome: The proposed method performs well on a multilingual dataset in English, French, and Dutch . it is possible to extract surprisingly good results even with scarce annotated data using existing models and datasets for modern languages .
People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language Detection (2023.emnlp-main)

Copied to clipboard

Challenge: Past work has shown that counterfactually augmented data (CADs) can improve models' performance on out-of-domain tests.
Approach: They use Polyjuice, ChatGPT, and Flan-T5 to automatically generate CADs . they find that CAD generates a model that flips the original label with minimal changes .
Outcome: The proposed model improves model robustness on out-of-domain test sets and individual data points.
Transductive Auxiliary Task Self-Training for Neural Multi-Task Models (D19-61)

Copied to clipboard

Challenge: Multi-task learning and self-training are two common ways to improve a machine learning model’s performance in settings with limited training data.
Approach: They propose a transductive auxiliary task self-training procedure that trains a model on auxiliary tasks and test instances with auxiliary labels generated by a single-task version of the model.
Outcome: The proposed method improves accuracy by 9.56% over the pure multi-task model for dependency relation tagging and 13.03% for semantic taging.
Generating Scientific Claims for Zero-Shot Scientific Fact Checking (2022.acl-long)

Copied to clipboard

Challenge: Existing methods for scientific fact checking require domain expertise and time consuming.
Approach: They propose a new supervised method for generating claims from scientific sentences and a novel method for negating claims.
Outcome: The proposed method improves on existing methods on biomedical claims and negations.
Mind the Style Gap: Meta-Evaluation of Style and Attribute Transfer Metrics (2025.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) make it easy to rewrite a text in any style, but they are not straightforward when evaluating content preservation.
Approach: They propose a large meta-evaluation of metrics for evaluating style and attribute transfer, focusing on content preservation.
Outcome: The proposed method achieves higher alignment with human judgements than prompting a model of a similar size as an autorater.
Issue Framing in Online Discussion Fora (N19-1)

Copied to clipboard

Challenge: In online discussion fora, speakers often make arguments by highlighting certain aspects of the topic.
Approach: They propose to use a newswire and social media annotated corpus to detect issue frames in online discussions.
Outcome: The proposed model can be applied to the domain of discussion fora using multi-task and adversarial training.
A Diagnostic Study of Explainability Techniques for Text Classification (2020.emnlp-main)

Copied to clipboard

Challenge: Existing explainability techniques that can be produced post-hoc with already trained models are lacking a definitive guide on how to choose one given a particular task and model architecture.
Approach: They propose to use a list of diagnostic properties to evaluate existing explainability techniques to compare them with human annotations of salient input regions.
Outcome: The proposed list compares a set of explainability techniques on downstream text classification tasks and neural network architectures.
Measuring Intersectional Biases in Historical Documents (2023.findings-acl)

Copied to clipboard

Challenge: digitised historical documents suffer from errors introduced by optical character recognition (OCR) and are written in an archaic language.
Approach: They investigate the continuities and transformations of bias in Caribbean historical newspapers during the colonial era . they use distributional semantics models and word embeddings to measure gender, race, and intersectional biases.
Outcome: The authors show that gender and racial biases are interdependent and their intersection triggers distinct effects.
Understanding Fine-grained Distortions in Reports of Scientific Findings (2024.findings-acl)

Copied to clipboard

Challenge: a fine-grained understanding of how scientific findings are reported is crucial, says a new study . a recent study found that tweets distort scientific findings more often than news reports .
Approach: They propose to annotate 1,600 scientific findings from academic papers paired with corresponding tweets . they also establish baselines for automatically detecting these characteristics .
Outcome: The proposed method outperforms few-shot prompting in detecting distortions in unpaired data.
Zero-Shot Cross-Lingual Transfer with Meta Learning (2020.emnlp-main)

Copied to clipboard

Challenge: There are more than 7,000 languages spoken in the world, over 90 of which have more than 10 million native speakers each.
Approach: They propose to use meta-learning to train a model on multiple languages at the same time . they use standard supervised, zero-shot cross-lingual, and few-shot crosses-lingual settings for different natural language understanding tasks.
Outcome: The proposed setup improves on the state-of-the-art for a total of 15 languages.
A Survey on Stance Detection for Mis- and Disinformation Identification (2022.findings-naacl)

Copied to clipboard

Challenge: Understanding attitudes expressed in texts plays an important role in systems for detecting false information online, be it misinformation (unintentionally false) or disinformation (intentional false information).
Approach: They examine the relationship between stance detection and mis- and disinformation detection online and examine the results of previous studies.
Outcome: The proposed task is a component of fact-checking, rumour detection, and detecting previously fact- checked claims, and is compared with other related tasks such as argumentation mining and sentiment analysis.
Modeling Information Change in Science Communication with Semantically Matched Paraphrases (2022.emnlp-main)

Copied to clipboard

Challenge: Whether the media faithfully communicate scientific information has long been a core issue to the science community.
Approach: They propose to use the SCIENTIFIC PARAPHRASE AND INFORMATION CHANGE DATASET to identify paraphrased scientific findings annotated for degree of information change to enable large-scale tracking and analysis of information changes in science communication.
Outcome: The proposed dataset contains 6,000 scientific finding pairs extracted from news stories, social media discussions, and full texts of original papers.
Does Typological Blinding Impede Cross-Lingual Sharing? (2021.eacl-main)

Copied to clipboard

Challenge: Existing work on bridging the performance gap between high- and low-resource languages has only found minor benefits from using typological information.
Approach: They propose to use typological features to train models in a cross-lingual setting to learn latent weights between languages.
Outcome: The proposed model overshadows the utility of explicitly using typological features by ignoring them, and shows that encouraging sharing according to typology improves performance.
SynDARin: Synthesising Datasets for Automated Reasoning in Low-Resource Languages (2025.coling-main)

Copied to clipboard

Challenge: Question Answering datasets are scarce for languages other than English due to the cost and difficulties of collection and manual annotation.
Approach: They propose a method for generating and validating QA datasets for low-resource languages . they use English data as context to generate synthetic multiple-choice (MC) question-answer pairs .
Outcome: The proposed method maintains quality, reduces likelihood of factual errors, and circumvents costly annotation.
Graph-Guided Textual Explanation Generation Framework (2025.emnlp-main)

Copied to clipboard

Challenge: Existing work has questioned their faithfulness, as they may not accurately reflect the model’s internal reasoning process regarding its predicted answer.
Approach: They propose a Graph-Guided Textual Explanation Generation framework that generates a graph neural network layer that guides the NLE generation and generates explanations with greater semantic and lexical similarity to human-written ones.
Outcome: The proposed framework improves NLE faithfulness by up to 12.12% compared to baseline methods on encoder-decoder and decoder-only models.
Specializing Large Language Models to Simulate Survey Response Distributions for Global Populations (2025.naacl-long)

Copied to clipboard

Challenge: Prior work has focused on using large language models to simulate human behaviors . but, LLMs are known to generate erroneous, stereotypical, or overconfident answers .
Approach: They propose to specialize large language models for simulating survey response distributions by first-token probabilities.
Outcome: The proposed model outperforms other methods and zero-shot classifiers on unseen questions, countries, and a completely unseened survey.
Jack the Reader – A Machine Reading Framework (P18-4)

Copied to clipboard

Challenge: Many Machine Reading and Natural Language Understanding tasks require reading supporting text in order to answer questions.
Approach: They propose a framework for Machine Reading that allows for quick prototyping by component reuse and evaluation of new models on existing datasets.
Outcome: The proposed framework supports question answering, natural language inference and link prediction tasks.
A strong baseline for question relevancy ranking (D18-1)

Copied to clipboard

Challenge: SemEval-16 and Semeval-17 community question answering shared tasks require complex pipelines and manual feature engineering to beat the IR baseline.
Approach: They train a multi-task feed forward network on a bag of 14 distance measures for the input question pair and train it using language-independent features.
Outcome: The proposed model outperforms the best shared task systems on the task of retrieving relevant previously asked questions.
Can Community Notes Replace Professional Fact-Checkers? (2025.acl-short)

Copied to clipboard

Challenge: Fact-checkers are crucial in combating misinformation on social media . however, community moderation is often employed in parallel due to the scale of misleading content shared online.
Approach: They use language models to annotate Twitter/X community notes with attributes such as topic, cited sources, and whether they refute misinformation claims.
Outcome: The results show that community notes cite fact-checking sources up to five times more than previously reported.
Investigating Human Values in Online Communities (2025.naacl-long)

Copied to clipboard

Challenge: Existing value frameworks struggle with sample sizes and rely on selfreported surveys to calculate values.
Approach: They propose a method to computationally analyse values on Reddit using in-domain and out-of-domain human annotations to train a value relevance and a polarity classifier.
Outcome: The proposed method can be used to analyse values on reddit using human annotations and human annotation.
Unstructured Evidence Attribution for Long Context Query Focused Summarization (2025.emnlp-main)

Copied to clipboard

Challenge: Existing systems struggle to copy and properly cite unstructured evidence, which also tends to be “lost-in-the-middle”.
Approach: They propose to extract unstructured evidence spans to improve the trustworthiness of large language models by citing unstructure . they propose to use this dataset as a training supervision for unstructure-based evidence summarization.
Outcome: The proposed pipeline generates more relevant and factually consistent evidence than baselines with no fine-tuning and fixed granularity evidence.
Can Transformers Learn n-gram Language Models? (2024.emnlp-main)

Copied to clipboard

Challenge: Existing work has tested transformers' ability to represent formal languages, but language models are not classifiers of strings but rather distributions over them.
Approach: They relate transformers' ability to learn random n-gram language models to ngram language model (LM) they find add- smoothing outperforms transformers on the former, while transformers perform better on the latter .
Outcome: The proposed models outperform classical methods designed to learn n-gram LMs, while transformers perform better on the latter.
FLARE: Faithful Logic-Aided Reasoning and Exploration (2025.emnlp-main)

Copied to clipboard

Challenge: Modern Question Answering (QA) and Reasoning approaches with Large Language Models (LLMs) use Chain-of-Thought (CoT) prompting but struggle with ambiguous tasks.
Approach: They propose a method that uses large language models to plan solutions and formalize queries without external solvers to generate outputs faithful to their intermediate reasoning chains.
Outcome: The proposed method achieves SOTA results on 7 out of 9 diverse reasoning benchmarks and 3 out of 3 logic inference benchmarks while enabling measurement of reasoning faithfulness.
Can Edge Probing Tests Reveal Linguistic Knowledge in QA Models? (2022.coling-1)

Copied to clipboard

Challenge: grammatical knowledge is encoded in large pre-trained language models (LMs) this is done through supervised classification tasks to predict the grammamatical properties of a span using only the token representations coming from the LM encoder.
Approach: They propose to use a supervised 'edge probing' task to detect grammatical knowledge in large pre-trained language models (LMs) this is done by encoding grammamatical properties using only token representations coming from the LM encoder.
Outcome: The proposed model performs well when fine-tuned or in adversarial situations where the model is forced to learn wrong correlations.
Factcheck-Bench: Fine-Grained Evaluation Benchmark for Automatic Fact-checkers (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models generate naturally sounding answers over a broad range of human inquiries, but they often generate answers that contradict real-world facts.
Approach: They propose a framework for annotating and evaluating the factuality of large language models . they propose 'factcheck-bench' which provides a multi-stage annotation scheme .
Outcome: The proposed framework outperforms several popular LLM fact-checkers in claim, sentence, and document levels.
PHD: Pixel-Based Language Modeling of Historical Documents (2023.emnlp-main)

Copied to clipboard

Challenge: Recent years have seen a boom in efforts to digitise historical documents in numerous languages and sources, leading to a transformation in the way historians work.
Approach: They propose a method for generating synthetic scans to resemble real historical documents by pre-training a model to reconstruct masked patches instead of predicting token distributions.
Outcome: The proposed model can reconstruct masked patches and understand language well.
Transformer Based Multi-Source Domain Adaptation (2020.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to improve machine learning performance are mixed experts and domain adversarial training.
Approach: They investigate the problem of unsupervised multi-source domain adaptation . they combine predictions of multiple domain experts and combine them to induce a domain agnostic representation space .
Outcome: The proposed methods improve models' performance while limiting learning time.
Generating Label Cohesive and Well-Formed Adversarial Claims (2020.emnlp-main)

Copied to clipboard

Challenge: Existing work on adversarial triggers for fact checking models reveals weaknesses and flaws of models . universal adversarials often inadvertently invert the meaning of instances they are inserted in .
Approach: They propose a method for automatically generating highly potent, well-formed, label cohesive claims for FC using universal adversarial triggers.
Outcome: The proposed method maintains the directionality and semantic validity of the claim better than previous work on the FEVER dataset.
How Does Counterfactually Augmented Data Impact Models for Social Computing Constructs? (2021.emnlp-main)

Copied to clipboard

Challenge: Recent studies have shown that models trained on CAD can learn cues in the dataset which are spuriously correlated with the construct.
Approach: They focus on sentiment, sexism, and hate speech as social constructs to investigate their effects on model performance.
Outcome: The proposed model generalizes better on out-of-domain datasets while relying less on spurious features.
Evaluating Input Feature Explanations through a Unified Diagnostic Evaluation Framework (2025.naacl-long)

Copied to clipboard

Challenge: Input feature explanations reveal how a model makes decisions based on a specific input.
Approach: They propose a framework that facilitates an automated comparison between highlight and interactive explanations comprised of four diagnostic properties.
Outcome: The proposed framework compares highlight and interactive explanations across two datasets and two models and shows that interactive span explanations outperform other explanation types across most diagnostic properties.
Topic-Guided Sampling For Data-Efficient Multi-Domain Stance Detection (2023.acl-long)

Copied to clipboard

Challenge: Stance Detection is a task that aims to identify the attitudes of an author towards a target of interest.
Approach: They propose a topic-guided diversity sampling technique and a contrastive objective to improve stance detection using the produced set.
Outcome: The proposed method outperforms the state-of-the-art on 16 datasets with in-domain and out-of domain evaluations and is more generalizable with an averaged 10.2 F1 on out-domain evaluation.
A Neighborhood Framework for Resource-Lean Content Flagging (2022.tacl-1)

Copied to clipboard

Challenge: Existing approaches to cross-lingual content flagging with limited target language data are lacking in many languages.
Approach: They propose a framework for cross-lingual content flagging with limited target- language data based on a nearest-neighbor architecture and a transformer representation in all its components.
Outcome: The proposed framework outperforms previous work in terms of predictive performance on eight languages from two different datasets.
Why Should This Article Be Deleted? Transparent Stance Detection in Multilingual Wikipedia Editor Discussions (2023.emnlp-main)

Copied to clipboard

Challenge: Currently, only 20% of the English comments explicitly mention content moderation policies, but as few as 2% of the German and Turkish comments.
Approach: They propose to use a multilingual dataset to predict stances with existing content moderation policies and to use them to explain moderation decisions.
Outcome: The proposed model predicts stances and corresponding reasons with high accuracy, adding transparency to the decision-making process.
Neighborhood Contrastive Learning for Scientific Document Representations with Citation Embeddings (2022.emnlp-main)

Copied to clipboard

Challenge: Prior work relies on discrete citation relations to generate contrast samples, but discrete ones enforce a hard cut-off to similarity.
Approach: They propose to use nearest neighbor sampling to learn continuous similarity and to sample hard-to-learn negatives and positives by controlling the sampling margin between them.
Outcome: The proposed method outperforms the state-of-the-art on the SciDocs benchmark and can train (or tune) language models sample-efficiently.
Reliable Evaluation Protocol for Low-Precision Retrieval (2026.acl-short)

Copied to clipboard

Challenge: Recent studies have shown that low-precision methods can improve performance, but they introduce high variability in the results based on tie resolution.
Approach: They propose a retrieval evaluation protocol designed to reduce tie variation . high-precision scoring and tie-aware retrieval metrics are proposed to reduce this variability .
Outcome: The proposed retrieval evaluation protocol reduces tie-induced instability and recovers expected scores and ranges on 12 retrieval datasets.
Stress Testing Factual Consistency Metrics for Long-Document Summarization (2026.acl-long)

Copied to clipboard

Challenge: Existing short-form summarization metrics struggle with input length limitations and long-range dependencies.
Approach: They propose to evaluate the reliability of six widely used reference-free factuality metrics in the long-document setting by applying seven factually-preserving perturbations to summaries.
Outcome: The proposed short-form summarization metrics struggle with long-range dependencies and input length limitations.
Inducing Language-Agnostic Multilingual Representations (2021.starsem-1)

Copied to clipboard

Challenge: Cross-lingual representations have the potential to make NLP techniques available to the vast majority of languages in the world, but they currently require large pretraining corpora or access to typologically similar languages.
Approach: They propose to remove language identity signals from multilingual embeddings by re-aligning vector spaces of target languages to a pivot source language and removing language-specific means and variances.
Outcome: The proposed approaches reduce cross-lingual transfer gap by 8.9 points (m-BERT) and 18.2 points (XLM-R) on average across all tasks and languages.
A Reality Check on Context Utilisation for Retrieval-Augmented Generation (2025.acl-long)

Copied to clipboard

Challenge: Existing studies on LM context utilisation of retrieved information have focused on synthetic text.
Approach: They propose a dataset of unreliable, insufficient and difficult-to-understand contexts with real-world queries and contexts manually annotated for stance to compare them to synthetic datasets.
Outcome: The proposed model outperforms synthetic datasets and exaggerates rare context characteristics, leading to inflated context utilisation results.
Parameter sharing between dependency parsers for related languages (D18-1)

Copied to clipboard

Challenge: Using parameter sharing between parsers of related languages can improve performance, but there is no consensus on what parameters to share.
Approach: They propose a model where transition classifier parameters are shared and word and character parameters are controlled by a parameter that can be tuned on validation data.
Outcome: The proposed model improves on a monolingually trained baseline.
Cross-Domain Label-Adaptive Stance Detection (2021.emnlp-main)

Copied to clipboard

Challenge: Stance detection is a task that focuses on the classification of a writer’s viewpoint towards a target.
Approach: They propose an end-to-end unsupervised framework for out-of-domain prediction of unseen, user-defined labels.
Outcome: The proposed framework shows that it can be used to predict unseen labels over strong baselines.
Semi-Supervised Exaggeration Detection of Health Science Press Releases (2021.emnlp-main)

Copied to clipboard

Challenge: Recent studies have shown that news media exaggerate scientific papers by exagging their findings.
Approach: They propose a method to detect when a news article has exaggerated a scientific finding . they use annotated press release/abstract pairs to compare machine learning models .
Outcome: The proposed method outperforms PET and supervised learning on a multi-task version of Pattern Exploiting Training.
Measuring and Benchmarking Large Language Models’ Capabilities to Generate Persuasive Language (2025.naacl-long)

Copied to clipboard

Challenge: Recent studies have focused on specific domains or types of persuasion, but a general study has focused on how LLMs produce persuasive text.
Approach: They construct a dataset to measure and benchmark the ability of Large Language Models (LLMs) to produce persuasive text.
Outcome: The proposed model can be used to generate persuasive text across domains and domains.
Thorny Roses: Investigating the Dual Use Dilemma in Natural Language Processing (2023.findings-emnlp)

Copied to clipboard

Challenge: Dual use is a problem in the context of natural language processing, says aaron eliotta . eelisa et al.: it is important to examine their rightful use and potential misuse .
Approach: They propose a definition and checklist for dual-use in natural language processing based on a survey of NLP researchers and practitioners.
Outcome: The proposed checklist focuses on dual-use in NLP based on a survey of NLP researchers and practitioners.
A Probabilistic Generative Model of Linguistic Typology (N19-1)

Copied to clipboard

Challenge: a generative model of languages based on principles-and-parameters posits that languages toggle on or off . linguistic typologists use a set of universal parameters to determine which languages toggle . we show that the correlation between parameters is significant, and that it is not enough to write down the set of parameters available to languages.
Approach: They propose a generative model of language based on exponential-family matrix factorisation.
Outcome: a linguistic model outperforms baseline models on predicting held-out features by exploiting similarities between languages and their features.
Investigating the Impact of Model Instability on Explanations and Uncertainty (2024.findings-acl)

Copied to clipboard

Challenge: Explainable AI methods are typically evaluated holistically, but small perturbations to inputs can vastly distort explanations.
Approach: They artificially simulate epistemic uncertainty in text input by introducing noise at inference time and measure the effect on the output of pre-trained language models.
Outcome: The proposed model can detect salient tokens when uncertain, but it is not reliable when small perturbations are exposed during training.
X-WikiRE: A Large, Multilingual Resource for Relation Extraction as Machine Comprehension (D19-61)

Copied to clipboard

Challenge: Existing knowledge bases are heavily biased towards English, but Wikipedias cover very different topics in different languages.
Approach: They propose a multilingual dataset that frams relation extraction as a machine reading problem.
Outcome: The proposed model can be used to transfer models cross-lingually and improves knowledge base completion across languages.
Analysing Differences in Persuasive Language in LLM-Generated Text: Uncovering Stereotypical Gender Patterns (2026.findings-acl)

Copied to clipboard

Challenge: Prior work has shown that large language models can successfully persuade humans and amplify persuasive language.
Approach: They propose a framework for evaluating how persuasive language generation is affected by recipient gender, sender intent, or output language.
Outcome: The proposed framework varies persuasive language when the recipient gender is specified or when the sender intent is specified.
2kenize: Tying Subword Sequences for Chinese Script Conversion (2020.acl-main)

Copied to clipboard

Challenge: Traditional Chinese character conversion is a common step in Chinese NLP but current methods do not take into account that a simplified Chinese character can correspond to multiple traditional characters.
Approach: They propose a model that can disambiguate between mappings and convert between the two scripts by using subword segmentation and two language models.
Outcome: The proposed model outperforms previous Chinese Character conversion approaches by 6 points in accuracy.
CiteWorth: Cite-Worthiness Detection for Improved Scientific Document Understanding (2021.findings-acl)

Copied to clipboard

Challenge: Scientific document understanding is challenging due to the highly domain specific nature of scientific language.
Approach: They propose a large, contextualized, rigorously cleaned labelled dataset for cite-worthiness detection built from extracted scientific documents.
Outcome: The proposed model improves on a paragraphlevel contextualized sentence labelling model based on Longformer . the model shows a 5 F1 point improvement over SciBERT which considers only individual sentences .
Generating Fact Checking Explanations (2020.acl-main)

Copied to clipboard

Challenge: Existing work on automated fact checking is concerned with predicting the veracity of claims based on metadata, social network spread, language used in claims, and, more recently, evidence supporting or denying claims.
Approach: They propose to combine the generation of justifications for verdicts on claims with the multi-task model to optimize both objectives at the same time rather than training them separately.
Outcome: The proposed model improves the informativeness, coverage and overall quality of the generated explanations, rather than training them separately.
Same Neurons, Different Languages: Probing Morphosyntax in Multilingual Pre-trained Models (2022.naacl-main)

Copied to clipboard

Challenge: Existing studies show that multilingual pre-trained models can learn to generalise across languages . however, it remains unclear how these models learn to learn multilingual representations .
Approach: They propose a hypothesis that multilingual pre-trained models can derive language-universal abstractions about grammar by aligning morphosyntactic markers that fulfil a similar grammatical function across languages.
Outcome: The proposed model can derive language-universal abstractions even without explicit supervision.
SubjQA: A Dataset for Subjectivity and Review Comprehension (2020.emnlp-main)

Copied to clipboard

Challenge: Subjectivity is the expression of internal opinions or beliefs which cannot be objectively observed or verified.
Approach: They develop a dataset which investigates subjectivity in question answering . they find that subjectivity is an important feature in the case of QA .
Outcome: The proposed dataset shows that subjectivity is an important feature in question answering (QA) it also shows that subjective questions and answers can have more complex interactions than previously thought.
Presumed Cultural Identity: How Names Shape LLM Responses (2025.findings-emnlp)

Copied to clipboard

Challenge: Names can be used as markers of individuality, cultural heritage, and personal history when interacting with chatbots.
Approach: They propose to use names as cultural bias in chatbots to adapt to user input and task contexts.
Outcome: The proposed method demonstrates that LLMs make cultural identity assumptions based on their users’ presumed backgrounds based upon their names .
Multi-Modal Framing Analysis of News (2025.emnlp-main)

Copied to clipboard

Challenge: Automated frame analysis of political communication has been limited by the use of predefined frames and the visual contexts in which they appear.
Approach: They propose a method for doing multi-modal, multi-label framing analysis at scale using large (vision-) language models.
Outcome: The proposed method provides a more complete picture for understanding media bias.
Claim Verification in the Age of Large Language Models: A Survey (2026.acl-srw)

Copied to clipboard

Challenge: Recent election cycles have seen a large number of false information spread across social media and news platforms.
Approach: They propose a framework for automated claim verification using Large Language Models and Retrieval Augmented Generation.
Outcome: The proposed frameworks are based on large-scale models and new methods such as Retrieval Augmented Generation (RAG).
Claim Check-Worthiness Detection as Positive Unlabelled Learning (2020.findings-emnlp)

Copied to clipboard

Challenge: a unified approach to claim check-worthiness detection is a critical component of fact checking systems.
Approach: They propose a unified approach which corrects for misinformation by positive unlabelled learning . they propose citation needed detection from Wikipedia and a ranking task which is a critical component of automatic fact checking systems.
Outcome: The proposed method outperforms the state of the art in two of the three tasks studied in English.
Explainability and Interpretability of Multilingual Large Language Models: A Survey (2025.emnlp-main)

Copied to clipboard

Challenge: Existing literature on multilingual large language models lacks transparency in their internal processes.
Approach: They propose to use multilingual large language models to examine their explainability and interpretability methods.
Outcome: The present study examines the explainability and interpretability of multilingual large language models.
Explaining Sources of Uncertainty in Automated Fact-Checking (2026.acl-long)

Copied to clipboard

Challenge: Existing methods to explain model uncertainty as numbers or hedges do not reveal which evidence conflicts cause the uncertainty, leaving users unable to resolve disagreements.
Approach: They propose a plug-and-play framework that generates natural-language explanations of model uncertainty grounded in conflicting/agreeing evidence.
Outcome: The proposed framework generates explanations that more faithfully track model uncertainty and better align with the model’s fact-checking decisions than span-agnostic explanation prompting.
Machine Reading, Fast and Slow: When Do Models “Understand” Language? (2022.coling-1)

Copied to clipboard

Challenge: Existing models of reading comprehension score highly on NLU benchmarks, but they are often 'read fast', i.e. rely on shallow patterns.
Approach: They propose a definition for the reasoning steps expected from a system that would be 'reading slowly' they compare that behavior with five models of the BERT family of various sizes, observed through saliency scores and counterfactual explanations.
Outcome: The proposed model is compared with five models of the BERT family of various sizes, and compared using saliency scores and counterfactual explanations.
Mapping (Dis-)Information Flow about the MH17 Plane Crash (D19-50)

Copied to clipboard

Challenge: Digital media enables fast sharing of information, but also disinformation . studies on the spread of disinformation on social media focused on small, manually annotated datasets or used proxys for data annotation.
Approach: They propose to use text classifiers to label Twitter content related to the MH17 crash to improve annotation accuracy.
Outcome: The proposed classifier improves over a hashtag-based baseline, but still remains a challenge in labelling pro-Russian and pro-Ukrainian content with high precision.
Multi-Task Learning of Pairwise Sequence Classification Tasks over Disparate Label Spaces (N18-1)

Copied to clipboard

Challenge: Multi-task learning and semi-supervised learning are successful paradigms for learning in scenarios with limited labelled data.
Approach: They propose to induce a joint embedding space between disparate label spaces and learning transfer functions between label embeddments to leverage unlabelled data and auxiliary, annotated datasets.
Outcome: The proposed approach outperforms strong single and multi-task baselines and achieves state of the art on aspect-based and topic-based sentiment analysis.
LLM Tropes: Revealing Fine-Grained Values and Opinions in Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to evaluate latent values and opinions in large language models suffer from three notable shortcomings.
Approach: They propose to analyze 156k LLM responses to 62 propositions of the Political Compass Test (PCT) generated by 6 LLMs using 420 prompt variations.
Outcome: The proposed analysis of 156k LLM responses to the Political Compass Test (PCT) generated by 6 LLMs shows that tropes are recurrent and consistent across prompts.
Faithfulness Tests for Natural Language Explanations (2023.acl-short)

Copied to clipboard

Challenge: Existing methods for explaining neural models are misleading as they often present reasons that are unfaithful to the model’s inner workings.
Approach: They propose a counterfactual input editor for inserting reasons that lead to counterfacts but are not reflected by the NLEs.
Outcome: The proposed model can evaluate emerging NLE models, proving a fundamental tool in the development of faithful explanations.
CUB: Benchmarking Context Utilisation Techniques for Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing language models (LMs) can be distracted by irrelevant contexts or ignore relevant information that contradicts outdated parametric memory.
Approach: They develop a benchmark to help diagnose CMTs under diverse noisy context conditions within retrieval-augmented generation (RAG) they find that most existing CMT struggle to handle the full spectrum of context types encountered in real-world RAG scenarios.
Outcome: The proposed benchmark compares seven state-of-the-art methods across three datasets and tasks, and shows that many lack the robustness needed to handle the full spectrum of context types encountered in real-world RAG scenarios.
Is Sparse Attention more Interpretable? (2021.acl-short)

Copied to clipboard

Challenge: Sparse attention has been claimed to increase model interpretability . however, the attention distribution is typically over representations internal to the model rather than the inputs themselves .
Approach: They conduct experiments to understand how sparsity affects our ability to use attention as an explainability tool.
Outcome: The proposed model does not map to a sparse set of influential inputs, but rather to fewer inputs.
Social Bias Probing: Fairness Benchmarking for Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for evaluating social biases in language models have been limited to binary association tests on small datasets.
Approach: They propose a framework for probing language models for social biases by assessing disparate treatment . they use a large-scale benchmark to examine the diversity of identities and stereotypes .
Outcome: The proposed framework expands the analysis beyond the binary comparison of stereotypical versus anti-stereotypical identities to include a diverse range of identities and stereotypes.
Unsupervised Discovery of Gendered Language through Latent-Variable Modeling (P19-1)

Copied to clipboard

Challenge: a recent study has focused on the ways in which language is gendered . positive adjectives used to describe women are more often related to their bodies .
Approach: They propose a model that models adjective choice and its sentiment given the natural gender of a head noun.
Outcome: The proposed model shows that positive adjectives used to describe women are more often related to their bodies than positive adjective words used to explain men.
Semantic Sensitivities and Inconsistent Predictions: Measuring the Fragility of NLI Models (2024.eacl-long)

Copied to clipboard

Challenge: Recent studies have indicated that NLI models have an understanding of lexical and compositional semantics.
Approach: They propose a framework to assess the extent of semantic sensitivity in NLI models . they use adversarially generated examples with minor semantics-preserving surface-form variations .
Outcome: The proposed framework shows that NLI models struggle with minor variations requiring knowledge of compositional semantics .
Combining Sentiment Lexica with a Multi-View Variational Autoencoder (N19-1)

Copied to clipboard

Challenge: a new model of sentiment lexica is being developed to combine disparate scales into a common representation.
Approach: They propose a model that unifies disparate scales into a common latent representation . they evaluate a text classification task using nine English-Language sentiment datasets .
Outcome: The proposed model outperforms six individual sentiment lexica and a simple combination thereof.
Revealing the Parametric Knowledge of Language Models: A Unified Framework for Attribution Methods (2024.acl-long)

Copied to clipboard

Challenge: Language Models acquire parametric knowledge from their training process, embedding it within their weights.
Approach: They propose a new evaluation framework to quantify and compare the knowledge revealed by Instance Attribution and Neuron Attributions.
Outcome: The proposed evaluation framework compares the knowledge revealed by IA and NA with that of neuron attribution methods.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations