Papers by Soline Felice

1 papers
Audiocite.net : A Large Spoken Read Dataset in French (2024.lrec-main)

Copied to clipboard

Challenge: Existing self-supervised learning methods for speech processing have proved difficult to apply to French due to the scarcity of large speech datasets.
Approach: They present a corpus of 6,682 hours of audiobooks from 130 readers . they describe the creation process and final statistics of the corpus .
Outcome: The proposed model based on the audiocite.net corpus, which contains 6,682 hours of audiobooks, was able to perform in 14k version.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations