Papers by Maria Cassese

1 papers
The Invalsi Benchmarks: measuring the Linguistic and Mathematical understanding of Large Language Models in Italian (2025.coling-main)

Copied to clipboard

Challenge: Invalsi MATE is a high-resource language, but there are few benchmarks to evaluate generative Large Language Models in this language.
Approach: They propose three benchmarks to evaluate language models on mathematical understanding in italian . they use the Invalsi tests, which are administered to students aged 6 to 18 in the italian school system .
Outcome: The proposed benchmarks are based on the Invalsi tests and the Italian highschool math Olympics.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations