Papers by Thomas Niesler
A First South African Corpus of Multilingual Code-switched Soap Opera Speech (L18-1)
Copied to clipboard
| Challenge: | a corpus of code-switched speech from soap operas is compiled from soaps . the corpus contains 14.3 hours of annotated and segmented speech . |
| Approach: | They propose a speech corpus containing multilingual code-switching from soap operas . the corpus contains English, isiZulu, isisXhosa, Setswana and Sesotho speech . |
| Outcome: | The corpus contains 14.3 hours of annotated and segmented speech from soap operas . the speech rate is 1.22 to 1.83 times higher than prompted speech in the same languages . |
Automatic Partitioning of a Code-Switched Speech Corpus Using Mixed-Integer Programming (2024.lrec-main)
Copied to clipboard
| Challenge: | Currently, partitioning speech corpora is done by hand, but this is not feasible for the dataset under investigation. |
| Approach: | They propose to partition a 41.6-hour corpus of code-switched speech into training, development and testing partitions using mixed-integer linear programming. |
| Outcome: | The proposed method allows to partition a 41.6-hour corpus of code-switched speech into training, development and testing partitions while maintaining a fixed number of speakers and a specific amount of codeswitching speech in the development and test partitions. |
Semi-supervised Development of ASR Systems for Multilingual Code-switched Speech in Under-resourced Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing models for code-switching between languages are under-resourced and limited by text and acoustic data. |
| Approach: | They propose to construct four separate bilingual automatic speech recognisers corresponding to four different language pairs between which speakers switch frequently. |
| Outcome: | The proposed models are compared with a non-batch-wise approach and show that they perform better when used with sparse training data. |