Papers by Stefano Menini
Semantic Frame Extraction in Multilingual Olfactory Events (2024.lrec-main)
Copied to clipboard
| Challenge: | Despite the interest in studying this domain, little effort has been devoted to develop tools and models that can extract olfactory information from large amounts of text in a structured and scalable way. |
| Approach: | They propose a system for multilingual olfactory information extraction covering six European languages, namely English, French, Italian, Dutch, German and Slovene. |
| Outcome: | The proposed system detects olfactory related text adopting a FrameNet-like structure and identifies the lexical units triggering the smell event and a set of frame elements. |
CrisiText: A dataset of warning messages for LLM training in emergency communication (2026.findings-eacl)
Copied to clipboard
| Challenge: | Identifying threats and mitigating their potential damage during crisis situations is paramount for safeguarding endangered individuals. |
| Approach: | They present a large-scale dataset for the generation of warning messages across 13 different types of crisis scenarios. |
| Outcome: | The proposed dataset contains more than 400,000 warning messages (spanning almost 18,000 crisis situations) aimed at assisting civilians during and after such events. |
Building a Multilingual Taxonomy of Olfactory Terms with Timestamps (2022.lrec-1)
Copied to clipboard
| Challenge: | olfactory references play a crucial role in our memory and experiences . but only few works in NLP have attempted to capture this sensory dimension from a computational perspective. |
| Approach: | They describe a process that has led to the semi-automatic development of a taxonomy for olfactory information in four languages (English, French, German and Italian) |
| Outcome: | The proposed taxonomy can be extended using existing language models and n-grams to include olfactory terms in four languages. |
Hybrid Emoji-Based Masked Language Models for Zero-Shot Abusive Language Detection (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Recent studies have demonstrated the effectiveness of cross-lingual language model pre-training on NLP tasks. |
| Approach: | They propose a hybrid emoji-based Masked Language Model to leverage eojis across languages to improve the learning of short text messages. |
| Outcome: | The proposed model performs better on German, Italian and Spanish. |
Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators’ Disagreement (2021.emnlp-main)
Copied to clipboard
| Challenge: | supervised learning is a key component of offensive language detection, but there is little attention given to the quality of annotated data. |
| Approach: | They propose to examine the level of agreement among annotators while selecting data to create offensive language datasets, a task involving a high level of subjectivity. |
| Outcome: | The proposed datasets show that annotators' agreement has a strong effect on classifiers performance and robustness. |
Variationist: Exploring Multifaceted Variation and Bias in Written Language Data (2024.acl-demos)
Copied to clipboard
| Challenge: | Existing tools that inspect and visualize language data are limited in their capabilities. |
| Approach: | They propose a highly-modular, extensible, and task-agnostic tool that inspects language variation and bias across multiple variables, language units, and diverse metrics. |
| Outcome: | The proposed tool can inspect and visualize language variation and bias across variables, language units, and diverse metrics that go beyond descriptive statistics. |
EuroVerdict: A Multilingual Dataset for Verdict Generation Against Misinformation (2025.findings-acl)
Copied to clipboard
| Challenge: | a global issue that shapes public discourse shapes opinion and decision-making . many multilingual work has focused on claim verification rather than generating explanatory verdicts . |
| Approach: | They propose a multilingual dataset designed for verdict generation covering eight European languages. |
| Outcome: | The EuroVerdict dataset covers claims, manual verdicts, and supporting evidence . it is compared with other datasets in eight European languages . |
First-AID: the first Annotation Interface for grounded Dialogues (2025.acl-demo)
Copied to clipboard
| Challenge: | Existing tools to fine-tune Large Language Models for specific tasks are limited due to financial constraints and limited availability of human experts. |
| Approach: | They propose a human-in-the-loop framework for the knowledge-driven generation of synthetic dialogues using LLM prompting that implements different strategies of data collection that require different user intervention during dialogue generation. |
| Outcome: | The proposed framework reduces post-editing efforts and improves quality of generated dialogues. |