Papers by Riza Batista-Navarro

24 papers
Not all quantifiers are equal: Probing Transformer-based language models’ understanding of generalised quantifiers (2023.emnlp-main)

Copied to clipboard

Challenge: Recent popularity of generalised quantifiers and role in linguistics and logic raises the question of how they affect transformer-based language models (TLMs)
Approach: They propose to use textual entailment to assess the ability of TLMs to learn the meanings of generalised quantifiers by using a textual model-checking problem defined in a purely logical sense.
Outcome: The proposed method allows the automatic construction of datasets with respect to which we can assess the ability of TLMs to learn the meanings of generalised quantifiers.
ParlVote: A Corpus for Sentiment Analysis of Political Debates (2020.lrec-1)

Copied to clipboard

Challenge: Debate transcripts from the UK Parliament contain information about the positions taken by politicians towards important topics, but are difficult for humans to process manually.
Approach: They propose to use a linear classifier and a transformer word embedding model to classify sentiment polarity in debate speeches to evaluate sentiment analysis systems for the political domain.
Outcome: The proposed method performs better on the largest dataset and is more robust to other datasets.
Beyond Static Synthetic Noise: Assessing the Robustness of Large Language Models to Natural Context Variation in the Real World (2026.findings-acl)

Copied to clipboard

Challenge: Current robustness evaluation methods rely on static synthetic perturbations to stress-test models.
Approach: They propose a framework for automatically evaluating QA models under naturally occurring textual perturbations by replacing context passages with revised Wikipedia edit histories.
Outcome: The proposed framework replaces context passages with revised Wikipedia edit histories to improve model performance.
Natural Context Drift Undermines the Natural Language Understanding of Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: generative Large Language Models (LLMs) are based on natural text evolution .
Approach: They propose a framework for curating naturally evolved variants of reading passages from contemporary QA benchmarks and for analysing LLM performance across a range of semantic similarity scores.
Outcome: The proposed framework evaluates QA datasets and LLMs with publicly available training data.
Incorporating Zoning Information into Argument Mining from Biomedical Literature (2022.lrec-1)

Copied to clipboard

Challenge: Argumentative zoning is a text zonation scheme that is used to segment text into zones that serve distinct functions.
Approach: They propose to use zoning information to incorporate into argument mining tasks . they add zonation labels predicted by an off-the-shelf model to the beginning of each sentence .
Outcome: The proposed models improve argument mining models without additional annotation cost.
Knowledge Augmentation Enhances Token Classification for Recipe Understanding (2026.eacl-long)

Copied to clipboard

Challenge: Using entity type-specific and knowledge-augmented token classification, we achieve state-of-the-art (SOTA) results on 5 out of 7 benchmark recipe datasets, significantly outperforming traditional token classification methods.
Approach: They propose an entity type-specific and knowledge-augmented token classification framework to improve encoder models’ performance on recipe texts.
Outcome: The proposed model outperforms traditional token classification methods on 5 out of 7 recipe datasets and is the largest annotated food-related dataset to date.
Do You Hear The People Sing? Key Point Analysis via Iterative Clustering and Abstractive Summarisation (2023.acl-long)

Copied to clipboard

Challenge: Argument summarisation is a promising but currently under-explored field.
Approach: They propose a framework to generate key points from short texts in a task known as Key Point Analysis.
Outcome: The proposed framework improves state-of-the-art in argument summarisation with performance improvement of 14 percentage points compared to ROUGE and human evaluation scores.
A Framework for Evaluation of Machine Reading Comprehension Gold Standards (2020.lrec-1)

Copied to clipboard

Challenge: Existing literature on machine reading comprehension (MRC) data is limited on the data design of gold standards.
Approach: They propose a framework to investigate linguistic features, lexical cues and ambiguity in MRC gold standards.
Outcome: The proposed framework investigates the present linguistic features, required reasoning and background knowledge and factual correctness on the one hand, and the presence of lexical cues as a lower bound for the requirement of understanding on the other.
BEDAA: Bayesian Enhanced DeBERTa for Uncertainty-Aware Authorship Attribution (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for authorship attribution struggle with trustworthiness and interpretability across domains, languages, and stylistic variations.
Approach: They propose a Bayesian-Enhanced DeBERTa framework that integrates Bayes' reasoning with transformer-based language models to enable uncertainty-aware authorship attribution.
Outcome: The proposed framework achieves 19.69% improvement in F1-score across multiple authorship attribution tasks, including binary, multiclass, and dynamic authorship detection.
Which Side Are You On? A Multi-task Dataset for End-to-End Argument Summarisation and Evaluation (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have made it difficult to build an automated debate system that helps people to synthesise persuasive arguments.
Approach: They propose to use an argument mining dataset to capture the end-to-end process of preparing an argumentative essay for a debate.
Outcome: The proposed dataset shows that it performs better on individual tasks than on human-centred evaluations.
Lost in Formatting: How Output Formats Skew LLM Performance on Information Extraction (2026.eacl-long)

Copied to clipboard

Challenge: Information extraction systems, powered by Large Language Models (LLMs), are increasingly deployed in high-stakes domains such as biomedicine.
Approach: They propose to use output formatting as a critical yet largely overlooked hyperparameter in information extraction tasks.
Outcome: The output formatting is a critical but largely overlooked hyperparameter in large language models on information extraction tasks.
Probing the Uniquely Identifiable Linguistic Patterns of Conversational AI Agents (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in deep learning and natural language processing have led to the proliferation of conversational AI agents.
Approach: They construct linguistic profiles for five CAAs and use authorship attribution techniques to identify uniquely identifieable linguistic patterns for each model.
Outcome: The proposed model identifies unique identifiers (UILPs) for each CAA using authorship attribution techniques.
RaFoLa: A Rationale-Annotated Corpus for Detecting Indicators of Forced Labour (2022.lrec-1)

Copied to clipboard

Challenge: Forced labour is the most common type of modern slavery, affecting at least 24.9 million people worldwide.
Approach: They propose to annotate an English corpus for multi-class and multi-label forced labour detection using specialised data from specialised sources.
Outcome: The proposed corpus consists of 989 news articles annotated according to risk indicators defined by the International Labour Organization (ILO).
Does Acceleration Cause Hidden Instability in Vision Language Models? Uncovering Instance-Level Divergence Through a Large-Scale Empirical Study (2025.emnlp-main)

Copied to clipboard

Challenge: Current acceleration evaluations focus on minimal overall performance degradation . however, accelerated models can exhibit significant changes in instance-level predictions .
Approach: They investigate whether accelerated vision-Language Models can still give the same answers as before . they found that accelerated models changed original answers up to 20% of the time .
Outcome: The results show that accelerated models changed their original answers up to 20% of the time.
TIMELINE: Exhaustive Annotation of Temporal Relations Supporting the Automatic Ordering of Events in News Articles (2023.emnlp-main)

Copied to clipboard

Challenge: Existing temporal relation extraction models have low inter-annotator agreement due to lack of specificity of annotation guidelines . authors propose a method for annotating all temporal relations, including long-distance ones, which automates the process .
Approach: They propose a new annotation scheme that defines criteria for temporal relations to be annotated . scheme includes events even if they are not expressed as verbs, they argue .
Outcome: The proposed method reduces time and manual effort on the part of annotators.
LVPruning: An Effective yet Simple Language-Guided Vision Token Pruning Approach for Multi-modal Large Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Multi-modal Large Language Models (MLLMs) incur significant computational overhead due to the large number of vision tokens processed, limiting their practicality in resource-constrained environments.
Approach: They propose a language-guided vision token pruning method that can be integrated into existing MLLMs with minimal architectural changes.
Outcome: The proposed method reduces vision tokens by 90% and preserves model performance.
Multi-Loss Fusion: Angular and Contrastive Integration for Machine-Generated Text Detection (2024.findings-emnlp)

Copied to clipboard

Challenge: Modern natural language generation systems have led to the development of synthetic human-like open-ended texts, posing concerns as to who the original author of a text is.
Approach: They propose a custom DeBERTa model with angular loss and contrastive loss functions for effective class separation in neural text classification tasks.
Outcome: The proposed model improves on binary machine-generated text detection and multi-class neural authorship attribution tasks on a number of benchmark datasets.
Unravelling the Logic: Investigating the Generalisation of Transformers in Numerical Satisfiability Problems (2025.acl-long)

Copied to clipboard

Challenge: Transformer models exhibit minimal scale and noise invariance, along with limited vocabulary and number invariancy.
Approach: They probe the generalisation prowess of Transformer models with respect to the hitherto unexplored domain of numerical satisfiability problems.
Outcome: The proposed models exhibit minimal scale and noise invariance, along with limited vocabulary and number invariancy.
‘Aye’ or ‘No’? Speech-level Sentiment Analysis of Hansard UK Parliamentary Debate Transcripts (L18-1)

Copied to clipboard

Challenge: Transcripts of UK parliamentary debates are difficult for human readers to process due to the large quantity of textual data and the specialised language used.
Approach: They propose to use annotated sentiment labels and labels derived from speakers' votes to classify the sentiment polarity of speakers as being either positive or negative towards motions proposed in the debates.
Outcome: The proposed model outperforms existing models on a dataset of parliamentary debate transcripts using textual and contextual features.
Arg-LLaDA: Argument Summarization via Large Language Diffusion Models and Sufficiency-Aware Refinement (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to argument summarization rely on single-pass generation, offering limited support for factual correction or structural refinement.
Approach: They propose a large language diffusion framework that iteratively improves argument summarization by sufficiency-guided remasking and regeneration.
Outcome: Empirical results show that Arg-LLaDA surpasses state-of-the-art baselines in 7 out of 10 evaluation metrics.
ReciFine: Finely Annotated Recipe Dataset for Controllable Recipe Generation (2026.findings-eacl)

Copied to clipboard

Challenge: Existing resources, such as RecipeNLG, extract food items only from ingredient lists, overlooking entities expressed in instructions, such tools, chef actions, food and tool states, and durations.
Approach: They extend RecipeNLG to extract 97 million entities from 2.2 million recipes.
Outcome: The proposed model outperforms existing models trained on ingredient-list data on both automatic and human evaluations.
Argument mining as a multi-hop generative machine reading comprehension task (2023.findings-emnlp)

Copied to clipboard

Challenge: Argument mining is a natural language processing task that aims to generate an argumentative graph given an unstructured argumentative text.
Approach: They propose a new approach which transfers the argument mining task into a multi-hop reading comprehension task by incorporating a "chain of thought" information into the model.
Outcome: The proposed approach surpasses SOTA results on two arguments mining benchmarks.
Is the Understanding of Explicit Discourse Relations Required in Machine Reading Comprehension? (2021.eacl-main)

Copied to clipboard

Challenge: Existing benchmarks for machine reading comprehension (MRC) are insufficient to assess models for their capabilities to read and comprehend .
Approach: They propose an ablation-based method to assess the extent to which MRC datasets evaluate the understanding of explicit discourse relations.
Outcome: The proposed method shows that the model's performance drops on three large-scale datasets . the results suggest that most of the answers do not require understanding the discourse structure of the text.
IDEM: The IDioms with EMotions Dataset for Emotion Recognition (2024.lrec-main)

Copied to clipboard

Challenge: idiomatic expressions are used in everyday language and typically convey affect, i.e., emotion.
Approach: They present a dataset of idiom-containing sentences that were generated and labelled with any one of 36 emotion types using a generative language model.
Outcome: The proposed method achieves an agreement rate of 62% on the IDioms with EMotions dataset, with human validation by two independent annotators.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations