Papers by Alberto Lavelli

4 papers
MedMT5: An Open-Source Multilingual Text-to-Text LLM for the Medical Domain (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies on large language models for medical applications have focused on a single language . medical mT5 outperforms both encoders and similar sized text-to-text models in English, French, and Italian benchmarks .
Approach: They propose to train Medical mT5, the first open-source text-to-text multilingual model for the medical domain.
Outcome: The proposed model outperforms encoders and similar sized models on the Spanish, French, and Italian benchmarks while being competitive with current state-of-the-art models in English.
PoSTWITA-UD: an Italian Twitter Treebank in Universal Dependencies (L18-1)

Copied to clipboard

Challenge: Various approaches and ad hoc resources are needed to provide proper coverage of specific linguistic phenomena.
Approach: They propose to annotate tweets using a well-known dependency-based annotation format . they propose to use the tweets for training NLP systems to improve their performance .
Outcome: The proposed resource can be used for training of NLP systems on social media texts.
Thesis Proposal: LLMs post-training for multilingual medical tasks. Instruction-Tuning, Continual-Pretraining or Reasoning? (2026.acl-srw)

Copied to clipboard

Challenge: Adapting Large Language Models to the medical domain remains an active area of research .
Approach: They propose to compare three common adaptation approaches to adapt large language models to the medical domain.
Outcome: The proposed models are built on top of foundational LLMs and rely on different post-training methodologies for domain and task performance.
Comparing Machine Learning and Deep Learning Approaches on NLP Tasks for the Italian Language (2020.lrec-1)

Copied to clipboard

Challenge: Using available datasets, we compare deep learning and traditional machine learning methods for various NLP tasks in Italian.
Approach: They compare deep learning and traditional machine learning methods for various NLP tasks in Italian.
Outcome: The proposed methods outperform traditional methods in sequence tagging tasks and classification tasks in Italian.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations