Papers by Mohamed Nabih
MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languages (2024.emnlp-main)
Copied to clipboard
Marco Gaido, Sara Papi, Luisa Bentivogli, Alessio Brutti, Mauro Cettolo, Roberto Gretter, Marco Matassoni, Mohamed Nabih, Matteo Negri
| Challenge: | Existing speech FMs fall short of full compliance with open-source principles . existing models do not have model weights, code, and training data publicly available . |
| Approach: | They propose to use a CC-BY license to create open-source speech FMs for EU languages . they collect suitable training data by surveying automatic speech recognition datasets . |
| Outcome: | The proposed model can be used in the 24 official languages of the European Union. |