Papers by Aghilas Sini
Investigating Inter- and Intra-speaker Voice Conversion using Audiobooks (2022.lrec-1)
Copied to clipboard
| Challenge: | Audiobook readers play with their voices to emphasize some text passages, highlight discourse changes or significant events, or in order to make listening easier and entertaining. |
| Approach: | They propose to modify the narrator’s voice to fit the context of the story, such as the character who is speaking, using voice conversion. |
| Outcome: | The proposed method improves the quality of the voice conversion system and the speaker similarity. |
SynPaFlex-Corpus: An Expressive French Audiobooks Corpus dedicated to expressive speech synthesis. (L18-1)
Copied to clipboard
| Challenge: | a French audiobooks corpus contains 87 hours of good audio quality speech . audiobooks provide mono-genre and multi-speaker speech whereas audiobooks usually provide a few hours of mono- and multispeakers . |
| Approach: | They present an expressive French audiobooks corpus containing eighty seven hours of speech . the corpus is annotated automatically and provides information as phone labels, phone boundaries, syllables, words or morpho-syntactic tagging. |
| Outcome: | The proposed corpus contains 87 hours of speech recorded by a single speaker . the data will allow developing models to better control expressiveness in speech synthesis . |
When depth is redundant: Efficient transformer-based speech anti-spoofing (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing anti-spoofing countermeasures exhibit limited generalization to unseen spoof attacks, especially in out-of-domain evaluation settings. |
| Approach: | They propose a training strategy that aligns shallow and intermediate representations with those of the final transformer layer for speech deepfake detection. |
| Outcome: | The proposed model improves robustness to unseen spoofing attacks and enhances out-of-domain generalization over strong baselines. |