Papers by Mario Mezzanzanica

4 papers
SkiLLens: Recognising and Mapping Novel Skills from Millions of Job Ads Across Europe Using Language Models (2026.eacl-industry)

Copied to clipboard

Challenge: Online job ads (OJAs) provide a real-time view of changing demands but require first retrieving skill mentions from unstructured text and then solving the entity linking problem of connecting them to standardized skill taxonomies.
Approach: They propose a multilingual human-in-the-loop pipeline that extracts candidate skills from national OJA corpora using country-specific word embeddings.
Outcome: The proposed pipeline enables timely, multilingual monitoring of emerging skills, supporting agile policy-making and targeted training initiatives.
Contrastive Explanations of Text Classifiers as a Service (2022.naacl-demo)

Copied to clipboard

Challenge: ContrXT provides time contrastive explanations of black box text classifiers by manipulating binary decision diagrams.
Approach: They propose a system that provides time contrastive explanations of black box classifiers as a service by manipulating binary decision diagrams.
Outcome: The proposed system has a throughput of 2.55 users per second and is available as a python pip package.
RE-FIN: Retrieval-based Enrichment for Financial data (2025.coling-industry)

Copied to clipboard

Challenge: Financial sentiment analysis (FSA) is a powerful tool to support business decision-making and perform financial forecasting.
Approach: They propose a system that retrieves information from a knowledge base to enrich financial sentences, making them more knowledge-dense and explicit.
Outcome: The proposed system generates propositions from the knowledge base and employs Retrieval-Augmented Generation (RAG) to augment the original text with relevant information.
ITALIC: An Italian Culture-Aware Natural Language Benchmark (2025.naacl-long)

Copied to clipboard

Challenge: ITALIC is a large-scale benchmark dataset of 10,000 multiple-choice questions designed to evaluate the natural language understanding of the Italian language and culture.
Approach: They propose to use a large-scale benchmark dataset to evaluate the natural language understanding of the Italian language and culture.
Outcome: The ITALIC dataset spans 12 domains and uses 17 state-of-the-art LLMs to assess the natural language understanding of the italian language and culture.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations