Papers by Stefano Menini

8 papers
Semantic Frame Extraction in Multilingual Olfactory Events (2024.lrec-main)

Copied to clipboard

Challenge: Despite the interest in studying this domain, little effort has been devoted to develop tools and models that can extract olfactory information from large amounts of text in a structured and scalable way.
Approach: They propose a system for multilingual olfactory information extraction covering six European languages, namely English, French, Italian, Dutch, German and Slovene.
Outcome: The proposed system detects olfactory related text adopting a FrameNet-like structure and identifies the lexical units triggering the smell event and a set of frame elements.
CrisiText: A dataset of warning messages for LLM training in emergency communication (2026.findings-eacl)

Copied to clipboard

Challenge: Identifying threats and mitigating their potential damage during crisis situations is paramount for safeguarding endangered individuals.
Approach: They present a large-scale dataset for the generation of warning messages across 13 different types of crisis scenarios.
Outcome: The proposed dataset contains more than 400,000 warning messages (spanning almost 18,000 crisis situations) aimed at assisting civilians during and after such events.
Building a Multilingual Taxonomy of Olfactory Terms with Timestamps (2022.lrec-1)

Copied to clipboard

Challenge: olfactory references play a crucial role in our memory and experiences . but only few works in NLP have attempted to capture this sensory dimension from a computational perspective.
Approach: They describe a process that has led to the semi-automatic development of a taxonomy for olfactory information in four languages (English, French, German and Italian)
Outcome: The proposed taxonomy can be extended using existing language models and n-grams to include olfactory terms in four languages.
Hybrid Emoji-Based Masked Language Models for Zero-Shot Abusive Language Detection (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent studies have demonstrated the effectiveness of cross-lingual language model pre-training on NLP tasks.
Approach: They propose a hybrid emoji-based Masked Language Model to leverage eojis across languages to improve the learning of short text messages.
Outcome: The proposed model performs better on German, Italian and Spanish.
Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators’ Disagreement (2021.emnlp-main)

Copied to clipboard

Challenge: supervised learning is a key component of offensive language detection, but there is little attention given to the quality of annotated data.
Approach: They propose to examine the level of agreement among annotators while selecting data to create offensive language datasets, a task involving a high level of subjectivity.
Outcome: The proposed datasets show that annotators' agreement has a strong effect on classifiers performance and robustness.
Variationist: Exploring Multifaceted Variation and Bias in Written Language Data (2024.acl-demos)

Copied to clipboard

Challenge: Existing tools that inspect and visualize language data are limited in their capabilities.
Approach: They propose a highly-modular, extensible, and task-agnostic tool that inspects language variation and bias across multiple variables, language units, and diverse metrics.
Outcome: The proposed tool can inspect and visualize language variation and bias across variables, language units, and diverse metrics that go beyond descriptive statistics.
EuroVerdict: A Multilingual Dataset for Verdict Generation Against Misinformation (2025.findings-acl)

Copied to clipboard

Challenge: a global issue that shapes public discourse shapes opinion and decision-making . many multilingual work has focused on claim verification rather than generating explanatory verdicts .
Approach: They propose a multilingual dataset designed for verdict generation covering eight European languages.
Outcome: The EuroVerdict dataset covers claims, manual verdicts, and supporting evidence . it is compared with other datasets in eight European languages .
First-AID: the first Annotation Interface for grounded Dialogues (2025.acl-demo)

Copied to clipboard

Challenge: Existing tools to fine-tune Large Language Models for specific tasks are limited due to financial constraints and limited availability of human experts.
Approach: They propose a human-in-the-loop framework for the knowledge-driven generation of synthetic dialogues using LLM prompting that implements different strategies of data collection that require different user intervention during dialogue generation.
Outcome: The proposed framework reduces post-editing efforts and improves quality of generated dialogues.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations