Papers by Alessandro Suglia
ACT-Thor: A Controlled Benchmark for Embodied Action Understanding in Simulated Environments (2022.coling-1)
Copied to clipboard
| Challenge: | embodied AI tasks require a strong understanding of verbs and their corresponding actions. |
| Approach: | They propose a controlled benchmark for embodied action understanding using a simulated environment and a visual feature extractor. |
| Outcome: | The proposed benchmark achieves 81.4% accuracy and high inter-annotator agreement . the proposed model falls behind human models in a zero-shot scenario . |
LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks (2025.acl-short)
Copied to clipboard
Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fernández, Albert Gatt, Esam Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, Andre Martins, Philipp Mondorf, Vera Neplenbroek, Sandro Pezzelle, Barbara Plank, David Schlangen, Alessandro Suglia, Aditya K Surikuchi, Ece Takmaz, Alberto Testoni
| Challenge: | Existing evaluations of NLP models with LLMs are based on human judgments . however, there are concerns about their validity and reproducibility in proprietary models . |
| Approach: | They evaluate 11 current LLMs for their ability to replicate annotations. they show substantial variance across models and datasets. |
| Outcome: | The proposed model can replicate human annotations on 20 NLP datasets and show substantial variance across models and datasets. |
Imagining Grounded Conceptual Representations from Perceptual Information in Situated Guessing Games (2020.coling-main)
Copied to clipboard
Alessandro Suglia, Antonio Vergari, Ioannis Konstas, Yonatan Bisk, Emanuele Bastianelli, Andrea Vanzo, Oliver Lemon
| Challenge: | Existing models fail to learn multi-modal representations, relying on category labels at inference time. |
| Approach: | They propose a "imagination" module that learns context-aware and category-awful latent embeddings without relying on category labels at inference time. |
| Outcome: | The imagination module outperforms state-of-the-art competitors by 8.26% gameplay accuracy in the CompGuessWhat?! benchmark. |
Repairs in a Block World: A New Benchmark for Handling User Corrections with Multi-Modal Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | In dialogue, the addressee may misunderstand the speaker and respond erroneously. |
| Approach: | They collect, analyse, and publicly release a dataset of multi-modal TPR sequences in dialogue . they evaluate several state-of-the-art Vision and Language Models across multiple settings . |
| Outcome: | The proposed model underperforms in a human-robot interaction task compared to humans . the proposed model can benefit from specialised losses targeting relevant tokens . |
Combine to Describe: Evaluating Compositional Generalization in Image Captioning (2022.acl-srw)
Copied to clipboard
| Challenge: | Recent work on compositionality has focused on the ability to combine simpler concepts to understand & generate arbitrarily more complex conceptual structures. |
| Approach: | They propose to use a set of image captioning models to benchmark their compositional generalization properties. |
| Outcome: | The proposed models do not generalize in terms of systematicity and productivity, but are robust to synonym substitutions. |
Multitask Multimodal Prompted Training for Interactive Embodied Task Completion (2023.emnlp-main)
Copied to clipboard
Georgios Pantazopoulos, Malvina Nikandrou, Amit Parekh, Bhathiya Hemanthage, Arash Eshghi, Ioannis Konstas, Verena Rieser, Oliver Lemon, Alessandro Suglia
| Challenge: | Embodied MultiModal Agent (EMMA) is a unified encoder-decoder model that reasons over images and trajectories and casts action prediction as multimodal text generation. |
| Approach: | They propose an Embodied MultiModal Agent (EMMA) that uses a unified encoder-decoder model that reasons over images and trajectories and casts action prediction as multimodal text. |
| Outcome: | The proposed model performs on par with similar models on several VL benchmarks and sets a new state-of-the-art success rate on the Dialog-guided Task Completion (DTC) benchmark. |
PIXAR: Auto-Regressive Language Modeling in Pixel Space (2024.findings-acl)
Copied to clipboard
| Challenge: | Recent work shows the possibility of building open-vocabulary large language models that operate on pixel representations. |
| Approach: | They propose a pixel-based autoregressive LLM that performs generative tasks . they propose an adversarial pretraining stage that improves the readability and accuracy of PIXAR by 8.1 on LAMBADA and 8.5 on bAbI. |
| Outcome: | The proposed model performs free-form generative tasks while keeping the number of parameters on par with previous encoder-decoder models. |
CROPE: Evaluating In-Context Adaptation of Vision and Language Models to Culture-Specific Concepts (2025.naacl-long)
Copied to clipboard
| Challenge: | Recent Vision and Language models have shown impressive performance across benchmarks . however, frontier models lack cultural awareness and can affect global cultural diversity . |
| Approach: | They propose a visual question answering benchmark to probe the knowledge of culture-specific concepts and evaluate the capacity for cultural adaptation through contextual information. |
| Outcome: | The proposed model shows large performance disparities between culture-specific and common concepts in the parametric setting. |
Triangulating LLM Progress through Benchmarks, Games, and Cognitive Tests (2025.findings-emnlp)
Copied to clipboard
Filippo Momentè, Alessandro Suglia, Mario Giulianelli, Ambra Ferrari, Alexander Koller, Oliver Lemon, David Schlangen, Raquel Fernández, Raffaella Bernardi
| Challenge: | MMLU and BBH are three evaluation paradigms for language learning models . interactive games are superior to standard benchmarks in discriminating models based on human cognitive assessments . |
| Approach: | They examine three evaluation paradigms: standard benchmarks, interactive games and cognitive tests . they examine whether interactive games are more effective at discriminating LLMs . |
| Outcome: | The results show that interactive games are superior to standard benchmarks in discriminating models. |
Investigating the Role of Instruction Variety and Task Difficulty in Robotic Manipulation Tasks (2024.emnlp-main)
Copied to clipboard
| Challenge: | Evaluating generalisation capabilities of multimodal models based solely on performance on out-of-distribution data fails to capture their true robustness . proposed framework examines the role of instructions and inputs in generalisation abilities of such models, considering architectural design, input perturbations across language and vision modalities, and increased task complexity. |
| Approach: | They propose a framework that examines the role of instructions and inputs in the generalisation abilities of multimodal models, considering architectural design, input perturbations across language and vision modalities, and increased task complexity. |
| Outcome: | The proposed framework examines the role of instructions and inputs in the generalisation abilities of multimodal models, considering architectural design, input perturbations across language and vision modalities, and increased task complexity. |
Shaking Up VLMs: Comparing Transformers and Structured State Space Models for Vision & Language Modeling (2024.emnlp-main)
Copied to clipboard
| Challenge: | a task-agnostic visual encoding yields minimal performance gains on grounding, but Transformers outperform Mamba at in-context multimodal retrieval. |
| Approach: | They propose to replace Transformers in Visual Language Models with Mamba, a structured state space model that demonstrates promising performance in sequence modeling. |
| Outcome: | The proposed model outperforms Transformers-based models in captioning, question answering, and reading comprehension. |
CompGuessWhat?!: A Multi-task Evaluation Framework for Grounded Language Learning (2020.acl-main)
Copied to clipboard
Alessandro Suglia, Ioannis Konstas, Andrea Vanzo, Emanuele Bastianelli, Desmond Elliott, Stella Frank, Oliver Lemon
| Challenge: | Approaches to Grounded Language Learning focus on a single task-based final performance measure which may not depend on desirable properties of the learned hidden representations. |
| Approach: | They propose an evaluation framework for Grounded Language Learning with Attributes based on three sub-tasks: 1) Goal-oriented evaluation; 2) Object attribute prediction evaluation; and 3) Zero-shot evaluation. |
| Outcome: | The proposed framework evaluates the quality of learned representations with respect to attribute grounding. |
Lost in Space: Probing Fine-grained Spatial Understanding in Vision and Language Resamplers (2024.naacl-short)
Copied to clipboard
| Challenge: | Recent approaches for developing vision and language models leverage existing vision and a language expert and try to learn a mapping between them. |
| Approach: | They propose to use a resampler module to create a ‘visual prompt’ which is then fed to the large language models (LLM) using a textual prompt. |
| Outcome: | The proposed method has been shown to be effective across coarse-grained tasks like image captioning and visual question answering, but more fine-grounded tasks that require spatial understanding have not been thoroughly examined. |
AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding (2024.findings-emnlp)
Copied to clipboard
Alessandro Suglia, Claudio Greco, Katie Baker, Jose Part, Ioannis Papaioannou, Arash Eshghi, Ioannis Konstas, Oliver Lemon
| Challenge: | Current Vision-Language Models (VLMs) focus on third-person view videos, neglecting the richness of egocentric perceptual experience. |
| Approach: | They propose to use the Egocentric Video Understanding Dataset (EVUD) to train VLMs on video captioning and question answering tasks specific to egocentric videos. |
| Outcome: | The proposed model outperforms open-source models including strong Socratic models using GPT-4 as a planner by 3.6% and outperformed Claude 3 and Gemini Pro Vision 1.0. |
Playpen: An Environment for Exploring Learning From Dialogue Game Feedback (2025.emnlp-main)
Copied to clipboard
Nicola Horst, Davide Mazzaccara, Antonia Schmidt, Michael Sullivan, Filippo Momentè, Luca Franceschetti, Philipp Sadler, Sherzod Hakimov, Alberto Testoni, Raffaella Bernardi, Raquel Fernández, Alexander Koller, Oliver Lemon, David Schlangen, Mario Giulianelli, Alessandro Suglia
| Challenge: | In this paper, we investigate whether Dialogue Games—goal-directed and rule-governed activities driven predominantly by verbal actions—can also serve as a source of feedback signals for learning. |
| Approach: | They introduce Playpen, an environment for off- and online learning through Dialogue Game self-play, and investigate a representative set of post-training methods: supervised fine-tuning, direct alignment and reinforcement learning with Group Relative Policy Optimization. |
| Outcome: | The proposed model improves performance on unseen instances, but negatively impacts other skills, while interactive learning shows balanced improvements without loss of skills. |
FOSSIL: Harnessing Feedback on Suboptimal Samples for Data-Efficient Generalisation with Imitation Learning for Embodied Vision-and-Language Tasks (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Current approaches to embodied AI tend to learn policies from expert demonstrations, but without a mechanism to evaluate the quality of demonstrated actions, they are limited to learning from optimal behaviour or risk replicating errors and inefficiencies. |
| Approach: | They propose to embed language feedback into a Transformer-based policy and optionally complement the traditional next action prediction objective with auxiliary self-supervised learning objectives for feedback prediction. |
| Outcome: | The proposed method improves agents’ compositional generalisation abilities and robustness on a range of embodied Vision-and-Language tasks in a custom babyAI-XGen environment. |
An Empirical Study on the Generalization Power of Neural Representations Learned via Visual Guessing Games (2021.eacl-main)
Copied to clipboard
Alessandro Suglia, Yonatan Bisk, Ioannis Konstas, Antonio Vergari, Emanuele Bastianelli, Andrea Vanzo, Oliver Lemon
| Challenge: | Using guessing games, an artificial agent can learn to perform on novel downstream tasks such as Visual Question Answering (VQA). |
| Approach: | They propose a supervised learning scenario in which an agent learns to mimic successful guessing games and a novel way for an agent to play by itself, called Self-play via Iterated Experience Learning. |
| Outcome: | The proposed model can be applied to a VQA dataset using a supervised learning scenario and a novel way for an agent to play by itself. |