Papers by Mario Giulianelli

25 papers
LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks (2025.acl-short)

Copied to clipboard

Challenge: Existing evaluations of NLP models with LLMs are based on human judgments . however, there are concerns about their validity and reproducibility in proprietary models .
Approach: They evaluate 11 current LLMs for their ability to replicate annotations. they show substantial variance across models and datasets.
Outcome: The proposed model can replicate human annotations on 20 NLP datasets and show substantial variance across models and datasets.
Speaking the Language of Your Listener: Audience-Aware Adaptation via Plug-and-Play Theory of Mind (2023.findings-acl)

Copied to clipboard

Challenge: Adaptation is a process in human communication by which a speaker tunes its language to that of a listener to achieve communicative success.
Approach: They propose a visual-based referential game between a knowledgeable speaker and a listener with limited visual and linguistic experience to model this adaptation mechanism.
Outcome: The proposed model improves on plug-and-play approaches to controlled language generation without finetuning the speaker’s underlying language model.
Construction Repetition Reduces Information Rate in Dialogue (2022.aacl-main)

Copied to clipboard

Challenge: We observe that construction usage lowers the information content of utterances.
Approach: They propose to use construction repetition to mitigate information rate in English open-domain spoken dialogues.
Outcome: The proposed method lowers the information content of utterances, while increasing the frequency and density of repetition.
Is Information Density Uniform when Utterances are Grounded on Perception and Discourse? (2026.eacl-long)

Copied to clipboard

Challenge: Existing studies on the distribution of information in visually grounded contexts have focused on text-only inputs.
Approach: They propose to use multilingual vision-and-language models to estimate surprisal . they find grounding on perception increases uniformity across typologically diverse languages .
Outcome: The proposed hypothesis is tested in visual-language models over 30 languages and 13 storytelling languages . the results show grounding on perception increases uniformity across languages compared to text-only settings .
Surprise! Uniform Information Density Isn’t the Whole Story: Predicting Surprisal Contours in Long-form Discourse (2024.emnlp-main)

Copied to clipboard

Challenge: Uniform Information Density (UID) hypothesis posits that speakers tend to distribute information evenly across linguistic units to achieve efficient communication.
Approach: They propose a functional pressure that speakers modulate information rate based on location within a hierarchically-structured model of discourse.
Outcome: The proposed hypothesis posits that speakers tend to distribute information evenly across linguistic units to achieve efficient communication.
Interpretable Word Sense Representations via Definition Generation: The Case of Semantic Change Analysis (2023.acl-long)

Copied to clipboard

Challenge: Existing approaches to semantic change analysis are limited in their interpretation power and lack of explanatory power.
Approach: They propose to use specialised Flan-T5 language models to generate a definition for each usage and a specialised word sense model to generate the most prototypical definition.
Outcome: The proposed representations outperform token or usage sentence embeddings in word-in-context semantic similarity judgements and are a promising type of lexical representation for NLP.
Information Locality as an Inductive Bias for Neural Language Models (2025.acl-long)

Copied to clipboard

Challenge: Inductive biases are inherent in every machine learning system, argues a new study . m-local entropy measures how well symbols disambiguate the next symbol .
Approach: They propose a framework that captures local uncertainty of a language by quantifying how effectively preceding symbols disambiguate the next symbol.
Outcome: The proposed framework captures the local uncertainty of a language by quantifying how effectively symbols disambiguate the next symbol.
AnaLog: Testing Analytical and Deductive Logic Learnability in Language Models (2022.starsem-1)

Copied to clipboard

Challenge: Existing approaches to NLP tasks rely on pre-trained language models, but some do not.
Approach: They propose a natural language inference task to test pre-trained language models for logical reasoning capabilities.
Outcome: The proposed language model performs better than other models across logical connectives and reasoning domains, but is sensitive to lexical and syntactic variations in the realisation of logical statements.
The Harmonic Structure of Information Contours (2025.acl-long)

Copied to clipboard

Challenge: Language typically does not maintain a uniform information rate, but it fluctuates around a global average . a new study suggests periodicity may be a factor in information rate oscillations .
Approach: They propose a hypothesis that language does not maintain a uniform information rate . they apply harmonic regression and introduce a new extension to detect periodicity .
Outcome: The proposed method reveals that language oscillates at periodic intervals across frequencies . it also offers a framework for uncovering structural pressures at various levels of linguistic granularity.
A Spatio-Temporal Point Process for Fine-Grained Modeling of Reading Behavior (2025.acl-long)

Copied to clipboard

Challenge: Standard models that focus on fixation durations ignore spatial dynamics of reading . authors propose a model that captures how long fixations last, where they land and when .
Approach: They propose a generative model that captures how long fixations last and where they land and when they occur.
Outcome: The proposed model exhibits higher likelihood on held-out reading data than baselines.
Triangulating LLM Progress through Benchmarks, Games, and Cognitive Tests (2025.findings-emnlp)

Copied to clipboard

Challenge: MMLU and BBH are three evaluation paradigms for language learning models . interactive games are superior to standard benchmarks in discriminating models based on human cognitive assessments .
Approach: They examine three evaluation paradigms: standard benchmarks, interactive games and cognitive tests . they examine whether interactive games are more effective at discriminating LLMs .
Outcome: The results show that interactive games are superior to standard benchmarks in discriminating models.
Surprisal Minimisation over Goal-directed Alternatives Predicts Production Choice in Dialogue (2026.acl-long)

Copied to clipboard

Challenge: a method to model utterance production is based on information-theoretic notions of cost . a technique to generate alternative sets of utterables is proposed .
Approach: They propose a procedure to generate both types of alternative sets using language models.
Outcome: The proposed procedure allows for speaker- and listener-oriented interpretations of different cost measures.
On the Proper Treatment of Tokenization in Psycholinguistics (2024.emnlp-main)

Copied to clipboard

Challenge: Language models are used in computational psycholinguistics to test theories that relate the surprisal of a region of interest to its cognitive cost experienced by readers.
Approach: They propose to marginalize token-level language models into character-level ones before they are used in psycholinguistic studies.
Outcome: The proposed model over token strings is better than character-level model, the authors show . the proposed model marginalizes token-level models into character-based models before they are used in psycholinguistic studies.
Efficiency and Effectiveness in Task-Oriented Dialogue: On Construction Repetition, Information Rate, and Task Success (2024.lrec-main)

Copied to clipboard

Challenge: Repetition of constructions in task-oriented dialogue can have negative and positive effects on information rate and delivery, but is also predictive of task success.
Approach: They investigate the role that efficiency and effectiveness play in speakers’ repetition of shared word sequences, or constructions, in task-oriented dialogue.
Outcome: The results show that repeating constructions has negative and positive effects on information rate and delivery and that information rate managing strategies are predictive of task success.
Generalized Measures of Anticipation and Responsivity in Online Language Processing (2024.findings-emnlp)

Copied to clipboard

Challenge: a generalization of classical information-theoretic measures of predictive uncertainty is proposed for online language processing . entropy and surprisal are two commonly deployed information- theoretic measure of predictive uncertainties in sentence processing based on the probability distribution of upcoming sequences of linguistic units .
Approach: They propose a generalization of classical information-theoretic measures of predictive uncertainty in online language processing based on simulations of incremental linguistic contexts.
Outcome: The proposed generalization of classical measures of predictive uncertainty in online language processing yields a positive effect on reading times and cloze completion probability.
Refer, Reuse, Reduce: Generating Subsequent References in Visual and Conversational Contexts (2020.emnlp-main)

Copied to clipboard

Challenge: Subsequent references exploit the common ground accumulated by the interlocutors and tend to be shorter and reuse expressions that were effective in previous mentions.
Approach: They propose a model that generates first and subsequent references in visually grounded dialogue . they also implement a reference resolution system to assess the referring effectiveness .
Outcome: The proposed model produces better, more effective referring utterances than one not grounded in the dialogue context.
Probing for Reading Times (2026.acl-long)

Copied to clipboard

Challenge: a large body of work on probing has demonstrated that language model representations encode a wealth of linguistic information, but it remains unclear whether they also capture cognitive signals about human processing.
Approach: They use regularized linear regression to compare language model representations against scalar predictors.
Outcome: The representations from early layers outperform surprisal in predicting early-pass measures such as first fixation and gaze duration.
Structure-Conditional Minimum Bayes Risk Decoding (2025.emnlp-main)

Copied to clipboard

Challenge: Minimum Bayes Risk (MBR) decoding has been used in machine translation for many years.
Approach: They propose three adaptations to the minimum bayes risk utility function to make it more sensitive to structural variability in the outcome space.
Outcome: The proposed adaptations significantly improve generation quality by up to 13.7 percentage points in win rate.
Towards Pragmatic Production Strategies for Natural Language Generation Tasks (2022.emnlp-main)

Copied to clipboard

Challenge: Using language to communicate successfully requires effort.
Approach: They propose a conceptual framework for the design of natural language generation systems that follow efficient and effective production strategies to achieve complex communicative goals.
Outcome: The proposed framework is applied to visually grounded referential games and abstractive text summarisation tasks with real-world applications.
Analysing Lexical Semantic Change with Contextualised Word Representations (2020.acl-main)

Copied to clipboard

Challenge: Existing studies on lexical semantic change have focused on detecting and characterising word meaning shifts using distributional semantic models.
Approach: They propose a method that exploits the BERT neural language model to obtain representations of word usages, clusters these representations into usage types, and measures change along time with three proposed metrics.
Outcome: The proposed method captures a variety of synchronic and diachronic linguistic phenomena and is highly reproducible and reproducible.
Towards a Similarity-adjusted Surprisal Theory (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies have shown that surprisal theory ignores the possibility of similarity between words and treats them as distinct entities.
Approach: They propose a new measure of comprehension effort called information value that accounts for communicative equivalences between possible continuations.
Outcome: The proposed measure of comprehension effort is based on the diversity index of the diversity of communicative units.
Information Value: Measuring Utterance Predictability as Distance from Plausible Alternatives (2023.emnlp-main)

Copied to clipboard

Challenge: 'information value' quantifies the predictability of an utterance relative to a set of plausible alternatives.
Approach: They propose a method to obtain interpretable estimates of information value using neural text generators and exploit their psychometric predictive power to investigate the dimensions of predictability that drive human comprehension behaviour.
Outcome: The proposed method is able to obtain interpretable estimates of information value using neural text generators and exploits their psychometric predictive power to investigate the dimensions of predictability that drive human comprehension behaviour.
Playpen: An Environment for Exploring Learning From Dialogue Game Feedback (2025.emnlp-main)

Copied to clipboard

Challenge: In this paper, we investigate whether Dialogue Games—goal-directed and rule-governed activities driven predominantly by verbal actions—can also serve as a source of feedback signals for learning.
Approach: They introduce Playpen, an environment for off- and online learning through Dialogue Game self-play, and investigate a representative set of post-training methods: supervised fine-tuning, direct alignment and reinforcement learning with Group Relative Policy Optimization.
Outcome: The proposed model improves performance on unseen instances, but negatively impacts other skills, while interactive learning shows balanced improvements without loss of skills.
What Comes Next? Evaluating Uncertainty in Neural Text Generators Against Human Production Variability (2023.emnlp-main)

Copied to clipboard

Challenge: In Natural Language Generation tasks, multiple communicative goals are plausible and any goal can be put into words, or produced, in multiple ways.
Approach: They characterise the extent to which human production varies lexically, syntactically, and semantically across four NLG tasks, connecting human production variability to aleatoric or data uncertainty.
Outcome: The proposed model can be calibrated to human production variability using multiple samples and, when possible, multiple references.
Is Information Density Uniform in Task-Oriented Dialogues? (2021.emnlp-main)

Copied to clipboard

Challenge: Evidence for the uniform information density principle has been found at many levels of language production.
Approach: They propose to use the Uniform Information Density principle to test whether and within which contextual units it holds in task-oriented dialogues.
Outcome: The proposed method is able to reduce fluctuations in the density of the information transmitted.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations