Papers by Justin Spence
Investigating data partitioning strategies for crosslinguistic low-resource ASR evaluation (2023.eacl-main)
Copied to clipboard
| Challenge: | Automatic speech recognition data sets include a single pre-defined test set consisting of one or more speakers whose speech never appears in the training set. |
| Approach: | They propose to use hold-speaker(s)-out partitioning to partition data for five languages . utterance duration and intensity are more predictive factors of variability . |
| Outcome: | The proposed method can produce results that do not reflect model performance on unseen data or speakers. |