Papers by Jesin James
Language Models for Code-switch Detection of te reo Māori and English in a Low-resource Setting (2022.findings-naacl)
Copied to clipboard
Jesin James, Vithya Yogarajan, Isabella Shields, Catherine Watson, Peter Keegan, Keoni Mahelona, Peter-Lucas Jones
| Challenge: | Te reo Mori is New Zealand’s only indigenous language spoken by 4.5% of the population of 5 million. |
| Approach: | They train bilingual sub-word embeddings to detect Mori-English code-switching points using a cloud-based multilingual system such as Google and Microsoft Azure. |
| Outcome: | The proposed model outperforms large-scale contextual models on down streaming tasks of detecting Mori language. |
Advocating Character Error Rate for Multilingual ASR Evaluation (2025.findings-naacl)
Copied to clipboard
| Challenge: | Word error rate (WER) has been used for automatic speech recognition (ASR) evaluations for English datasets for many years. |
| Approach: | They propose to use the character error rate as the primary metric in multilingual ASR evaluation to account for morphologically complex languages. |
| Outcome: | The character error rate (CER) is the primary evaluation metric in multilingual ASR evaluation. |
Development of Community-Oriented Text-to-Speech Models for Māori ‘Avaiki Nui (Cook Islands Māori) (2024.lrec-main)
Copied to clipboard
Jesin James, Rolando Coto-Solano, Sally Akevai Nicholas, Joshua Zhu, Bovey Yu, Fuki Babasaki, Jenny Tyler Wang, Nicholas Derby
| Challenge: | Text-to-speech synthesis is used to transform text into a synthesized voice for a specific language. |
| Approach: | They describe the development of a text-to-speech system for Mori ‘Avaiki Nui (Cook Islands Mi) they used two approaches to train the system, the HMM-system MaryTTS and the deep learning system FastSpeech2 . |
| Outcome: | The proposed system is based on the HMM-system MaryTTS and the deep learning system FastSpeech2 . the ground truth voice had the highest quality, but the fastspeech 2 voice had a significantly higher quality than the MaryTTs synthesized recordings. |