Papers by Marco Matassoni

2 papers
TLT-school: a Corpus of Non Native Children Speech (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of speech utterances collected in schools of northern italy is being used to assess the performance of students learning both English and German.
Approach: a corpus of speech utterances collected in schools of northern italy is described . the corpus is going to be freely distributed to scientific community .
Outcome: The corpus of speech utterances collected in schools of northern italy is a "Trentino Language Testing" in schools" the data are used to assess the performance of students learning English and German .
MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languages (2024.emnlp-main)

Copied to clipboard

Challenge: Existing speech FMs fall short of full compliance with open-source principles . existing models do not have model weights, code, and training data publicly available .
Approach: They propose to use a CC-BY license to create open-source speech FMs for EU languages . they collect suitable training data by surveying automatic speech recognition datasets .
Outcome: The proposed model can be used in the 24 official languages of the European Union.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations