Papers by Soline Felice
Audiocite.net : A Large Spoken Read Dataset in French (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing self-supervised learning methods for speech processing have proved difficult to apply to French due to the scarcity of large speech datasets. |
| Approach: | They present a corpus of 6,682 hours of audiobooks from 130 readers . they describe the creation process and final statistics of the corpus . |
| Outcome: | The proposed model based on the audiocite.net corpus, which contains 6,682 hours of audiobooks, was able to perform in 14k version. |