Papers by Alexander Gutkin
Improving Informally Romanized Language Identification (2025.emnlp-main)
Copied to clipboard
| Challenge: | Latin script is often used to informally write languages with non-Latin native scripts, resulting in high spelling variability. |
| Approach: | They propose to improve methods used to synthesize training sets to incorporate natural spelling variations into training sets. |
| Outcome: | The proposed method improves test F1 from the reported 74.7% (using a pretrained neural model) to 85.4% (using the linear classifier trained solely on synthetic data). |
FonBund: A Library for Combining Cross-lingual Phonological Segment Data (L18-1)
Copied to clipboard
| Challenge: | Speech and language technology is currently only available for a tiny fraction of the world's languages. |
| Approach: | They propose to map phonetic segments in International Phonetic Alphabet into multiple articulatory feature representations using a free-source library. |
| Outcome: | The proposed library can be easily modified to support new phonological segment inventories. |
Building Open Javanese and Sundanese Corpora for Multilingual Text-to-Speech (L18-1)
Copied to clipboard
Jaka Aris Eko Wibawa, Supheakmungkol Sarin, Chenfang Li, Knot Pipatsrisawat, Keshan Sodimana, Oddur Kjartansson, Alexander Gutkin, Martin Jansche, Linne Ha
| Challenge: | Using multi-speaker text-to-speech systems, we build systems for Javanese and Sundanese . progress in this direction is difficult because languages in the long tail of the distribution of the majority of the world's languages lack adequate linguistic resources . |
| Approach: | They present multi-speaker text-to-speech corpora for Javanese and Sundanese . they use mixed-gender recordings to build multi-language text-based systems . |
| Outcome: | The proposed multi-speaker text-to-speech systems outperform the systems constructed from a single language. |
Crowdsourcing Latin American Spanish for Low-Resource Text-to-Speech (2020.lrec-1)
Copied to clipboard
Adriana Guevara-Rukoz, Isin Demirsahin, Fei He, Shan-Hui Cathy Chu, Supheakmungkol Sarin, Knot Pipatsrisawat, Alexander Gutkin, Alena Butryna, Oddur Kjartansson
| Challenge: | Using crowd-sourced datasets, we build a text-to-speech voice for a new dialect in a language with existing resources. |
| Approach: | They propose a multidialectal corpus approach for building a text-to-speech voice for a new dialect in a language with existing resources using crowd-sourcing. |
| Outcome: | The proposed model outperforms baseline models in a “zero-resource” dialect scenario while holding out target dialect recordings from the training data. |
XTREME-UP: A User-Centric Scarce-Data Benchmark for Under-Represented Languages (2023.findings-emnlp)
Copied to clipboard
Sebastian Ruder, Jonathan Clark, Alexander Gutkin, Mihir Kale, Min Ma, Massimo Nicosia, Shruti Rijhwani, Parker Riley, Jean-Michel Sarr, Xinyi Wang, John Wieting, Nitish Gupta, Anna Katanova, Christo Kirov, Dana Dickinson, Brian Roark, Bidisha Samanta, Connie Tao, David Adelani, Vera Axelrod, Isaac Caswell, Colin Cherry, Dan Garrette, Reeve Ingle, Melvin Johnson, Dmitry Panteleev, Partha Talukdar
| Challenge: | Existing datasets are often informed by established research directions in the NLP community. |
| Approach: | They propose a benchmark to evaluate the capabilities of language models across 88 under-represented languages over 9 key user-centric technologies including ASR, OCR, MT, and information access tasks. |
| Outcome: | The proposed benchmark evaluates the capabilities of language models across 88 under-represented languages over 9 key user-centric technologies including ASR, OCR, MT, and information access tasks. |
Open-source Multi-speaker Corpora of the English Accents in the British Isles (2020.lrec-1)
Copied to clipboard
| Challenge: | Using a dataset of high-quality audio, the authors examine the accents of 120 volunteers in the British Isles. |
| Approach: | They present a dataset of high-quality audio of English sentences recorded by volunteers with different accents of the British Isles. |
| Outcome: | The transcribed audio includes pronunciations of global locations, major airlines and common personal names in different accents. |
Extensions to Brahmic script processing within the Nisaba library: new scripts, languages and utilities (2022.lrec-1)
Copied to clipboard
| Challenge: | a brahmic script is used to record endangered languages such as Dogri and Bengali for low-resource languages such that do not require visual normalization. |
| Approach: | They propose to extend Brahmic script functionality within the Nisaba library of finite-state script normalization and processing utilities. |
| Outcome: | The proposed extensions extend coverage from the original ten scripts to an additional ten of South Asia and beyond, including some used to record endangered languages such as Dogri. |
Burmese Speech Corpus, Finite-State Text Normalization and Pronunciation Grammars with an Application to Text-to-Speech (2020.lrec-1)
Copied to clipboard
Yin May Oo, Theeraphol Wattanavekin, Chenfang Li, Pasindu De Silva, Supheakmungkol Sarin, Knot Pipatsrisawat, Martin Jansche, Oddur Kjartansson, Alexander Gutkin
| Challenge: | Using crowd-sourced speech corpus and finite-state transducer grammars, we build a text-to-speech system for Burmese, a tonal Southeast Asian language from the Sino-Tibetan family. |
| Approach: | They propose an open-source crowd-sourced multi-speaker speech corpus and finite-state grammars for performing grapheme-to-phoneme conversion for Burmese. |
| Outcome: | The proposed system performs well for Burmese in a low-resource setting. |
Finite-state script normalization and processing utilities: The Nisaba Brahmic library (2021.eacl-demos)
Copied to clipboard
| Challenge: | a library for low-level processing of brahmic scripts is available for free. |
| Approach: | They propose an open-source library for efficient low-level processing of ten major South Asian Brahmic scripts. |
| Outcome: | The proposed library supports low-level processing of ten major south Asian Brahmic scripts. |
Open-source Multi-speaker Speech Corpora for Building Gujarati, Kannada, Malayalam, Marathi, Tamil and Telugu Speech Synthesis Systems (2020.lrec-1)
Copied to clipboard
Fei He, Shan-Hui Cathy Chu, Oddur Kjartansson, Clara Rivera, Anna Katanova, Alexander Gutkin, Isin Demirsahin, Cibu Johny, Martin Jansche, Supheakmungkol Sarin, Knot Pipatsrisawat
| Challenge: | We present free high quality multi-speaker speech corpora for Gujarati, Kannada, Malayalam, Marathi, Tamil and Telugu . the datasets are primarily intended for use in text-to-speech applications, such as constructing multilingual voices or language adaptation. |
| Approach: | They present a free high quality multi-speaker speech corpora for Gujarati, Kannada, Malayalam, Marathi, Tamil and Telugu . they use it to build a multilingual text-to-speech model that can be scaled to other languages of interest. |
| Outcome: | The proposed model produces good quality voices with MOS > 3.6 for all the languages tested. |
Helpful Neighbors: Leveraging Neighbors in Geographic Feature Pronunciation (2023.tacl-1)
Copied to clipboard
| Challenge: | a new architecture learns to use pronunciations of neighboring names to guess pronunciations . features cause not infrequent problems in the US, but become a serious issue in Japan . |
| Approach: | They propose an architecture that learns to use pronunciations of neighboring names to guess pronunciations . they propose corrections for errors in Google Maps and an application to a totally different task . |
| Outcome: | The proposed model can be applied to finding and proposing corrections for errors in Google Maps. |
Criteria for Useful Automatic Romanization in South Asian Languages (2022.lrec-1)
Copied to clipboard
| Challenge: | a number of possible criteria for systems that transliterate South Asian languages are considered . romanization is the special case where the target script is the Latin script. |
| Approach: | They propose a set of criteria for systems that transliterate South Asian languages . criteria include fidelity to human linguistic behavior, processing utility for people, invertibility . they then propose several algorithms that address different criteria . |
| Outcome: | The proposed algorithms address linguistic considerations in the context of Brahmic scripts and languages that use them, such as Hindi and Malayalam. |