Papers by Juliette Millet
Do self-supervised speech models develop human-like perception biases? (2022.acl-long)
Copied to clipboard
| Challenge: | Recent advances in speech recognition and representation learning show that self-supervised pretraining is an excellent way of improving performance while reducing the amount of labelled data needed for training. |
| Approach: | They compare the representational spaces of wav2vec, HuBERT and contrastive predictive coding (CPC) with the perceptual spaces of French-speaking and English-speaking human listeners. |
| Outcome: | The proposed models capture fine-grained perceptual phenomena while supervised models are better at capturing coarser, phone-level effects and effects of listeners’ native language on perception. |