Papers by Isin Demirsahin
Processing South Asian Languages Written in the Latin Script: the Dakshina Dataset (2020.lrec-1)
Copied to clipboard
Brian Roark, Lawrence Wolf-Sonkin, Christo Kirov, Sabrina J. Mielke, Cibu Johny, Isin Demirsahin, Keith Hall
| Challenge: | a new resource is available for 12 South Asian languages that use the Latin script for text entry . the Latin-script system is not widely used in South Asian language writing, despite the Latin alphabet . |
| Approach: | They describe the Dakshina dataset, a new resource consisting of text in both the Latin and native scripts for 12 South Asian languages. |
| Outcome: | The Dakshina dataset includes text in both the Latin and native scripts for 12 languages . the authors provide baseline results on several tasks made possible by the dataset . |
Crowdsourcing Latin American Spanish for Low-Resource Text-to-Speech (2020.lrec-1)
Copied to clipboard
Adriana Guevara-Rukoz, Isin Demirsahin, Fei He, Shan-Hui Cathy Chu, Supheakmungkol Sarin, Knot Pipatsrisawat, Alexander Gutkin, Alena Butryna, Oddur Kjartansson
| Challenge: | Using crowd-sourced datasets, we build a text-to-speech voice for a new dialect in a language with existing resources. |
| Approach: | They propose a multidialectal corpus approach for building a text-to-speech voice for a new dialect in a language with existing resources using crowd-sourcing. |
| Outcome: | The proposed model outperforms baseline models in a “zero-resource” dialect scenario while holding out target dialect recordings from the training data. |
Open-source Multi-speaker Corpora of the English Accents in the British Isles (2020.lrec-1)
Copied to clipboard
| Challenge: | Using a dataset of high-quality audio, the authors examine the accents of 120 volunteers in the British Isles. |
| Approach: | They present a dataset of high-quality audio of English sentences recorded by volunteers with different accents of the British Isles. |
| Outcome: | The transcribed audio includes pronunciations of global locations, major airlines and common personal names in different accents. |
Open-source Multi-speaker Speech Corpora for Building Gujarati, Kannada, Malayalam, Marathi, Tamil and Telugu Speech Synthesis Systems (2020.lrec-1)
Copied to clipboard
Fei He, Shan-Hui Cathy Chu, Oddur Kjartansson, Clara Rivera, Anna Katanova, Alexander Gutkin, Isin Demirsahin, Cibu Johny, Martin Jansche, Supheakmungkol Sarin, Knot Pipatsrisawat
| Challenge: | We present free high quality multi-speaker speech corpora for Gujarati, Kannada, Malayalam, Marathi, Tamil and Telugu . the datasets are primarily intended for use in text-to-speech applications, such as constructing multilingual voices or language adaptation. |
| Approach: | They present a free high quality multi-speaker speech corpora for Gujarati, Kannada, Malayalam, Marathi, Tamil and Telugu . they use it to build a multilingual text-to-speech model that can be scaled to other languages of interest. |
| Outcome: | The proposed model produces good quality voices with MOS > 3.6 for all the languages tested. |
Criteria for Useful Automatic Romanization in South Asian Languages (2022.lrec-1)
Copied to clipboard
| Challenge: | a number of possible criteria for systems that transliterate South Asian languages are considered . romanization is the special case where the target script is the Latin script. |
| Approach: | They propose a set of criteria for systems that transliterate South Asian languages . criteria include fidelity to human linguistic behavior, processing utility for people, invertibility . they then propose several algorithms that address different criteria . |
| Outcome: | The proposed algorithms address linguistic considerations in the context of Brahmic scripts and languages that use them, such as Hindi and Malayalam. |