Papers by Maria Cassese
The Invalsi Benchmarks: measuring the Linguistic and Mathematical understanding of Large Language Models in Italian (2025.coling-main)
Copied to clipboard
| Challenge: | Invalsi MATE is a high-resource language, but there are few benchmarks to evaluate generative Large Language Models in this language. |
| Approach: | They propose three benchmarks to evaluate language models on mathematical understanding in italian . they use the Invalsi tests, which are administered to students aged 6 to 18 in the italian school system . |
| Outcome: | The proposed benchmarks are based on the Invalsi tests and the Italian highschool math Olympics. |