Challenge: Using multi-speaker text-to-speech systems, we build systems for Javanese and Sundanese . progress in this direction is difficult because languages in the long tail of the distribution of the majority of the world's languages lack adequate linguistic resources .
Approach: They present multi-speaker text-to-speech corpora for Javanese and Sundanese . they use mixed-gender recordings to build multi-language text-based systems .
Outcome: The proposed multi-speaker text-to-speech systems outperform the systems constructed from a single language.

Similar Papers

Open-source Multi-speaker Speech Corpora for Building Gujarati, Kannada, Malayalam, Marathi, Tamil and Telugu Speech Synthesis Systems (2020.lrec-1)

Copied to clipboard

Challenge: We present free high quality multi-speaker speech corpora for Gujarati, Kannada, Malayalam, Marathi, Tamil and Telugu . the datasets are primarily intended for use in text-to-speech applications, such as constructing multilingual voices or language adaptation.
Approach: They present a free high quality multi-speaker speech corpora for Gujarati, Kannada, Malayalam, Marathi, Tamil and Telugu . they use it to build a multilingual text-to-speech model that can be scaled to other languages of interest.
Outcome: The proposed model produces good quality voices with MOS > 3.6 for all the languages tested.
NusaX: Multilingual Parallel Sentiment Dataset for 10 Indonesian Local Languages (2023.eacl-main)

Copied to clipboard

Challenge: In Indonesia, many languages are endangered and some are even extinct due to the unavailability of data resources and benchmarks.
Approach: They propose a high-quality multilingual parallel corpus that covers 10 local languages from Indonesia.
Outcome: The proposed resource includes sentiment and machine translation datasets, and bilingual lexicons.
Common Voice: A Massively-Multilingual Speech Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Common Voice is a massively-multilingual collection of transcribed speech intended for speech technology research and development.
Approach: They propose to use Mozilla’s DeepSpeech Speech-to-Text toolkit to perform multilingual automatic speech recognition experiments.
Outcome: The proposed corpus is the largest in the public domain for speech recognition, both in terms of hours and languages.
A New Massive Multilingual Dataset for High-Performance Language Technologies (2024.lrec-main)

Copied to clipboard

Challenge: a new massive multilingual dataset is available for language modeling and machine translation training.
Approach: They present a massive multilingual dataset using web crawls from the Internet Archive and CommonCrawl . they use open-source software tools and high-performance computing to acquire, manage and process large corpora .
Outcome: The HPLT language resources is a massive multilingual dataset . it includes monolingual and bilingual corpora extracted from CommonCrawl and the Internet Archive . the results are published online at the journal journal cense4 .
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
Praaline: An Open-Source System for Managing, Annotating, Visualising and Analysing Speech Corpora (P18-4)

Copied to clipboard

Challenge: Praaline is an open-source software system for constituting and managing spoken language and multimodal corpora.
Approach: They present the latest developments of Praaline, an open-source software system for constituting and managing spoken language and multimodal corpora.
Outcome: The proposed system can be used for creating, managing, visualising and analysing spoken language and multimodal corpora.
What Do Indonesians Really Need from Language Technology? A Nationwide Survey (2025.emnlp-main)

Copied to clipboard

Challenge: Despite efforts to develop NLP for Indonesia’s 700+ local languages, progress remains costly due to the need for direct engagement with native speakers.
Approach: They conduct a nationwide survey to assess the actual needs of native Indonesian speakers.
Outcome: The findings indicate that addressing language barriers is the most critical priority . concerns around privacy, bias, and the use of public data highlight the need for greater transparency and clear communication to support broader AI adoption.
PMIndiaSum: Multilingual and Cross-lingual Headline Summarization for Languages in India (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing datasets for Indian languages are limited in terms of coverage and size.
Approach: They propose a multilingual and massively parallel summarization corpus focused on languages in India that provides a training and testing ground for four language families, 14 languages, and the largest to date with 196 language pairs.
Outcome: The proposed dataset provides a training and testing ground for four language families, 14 languages, and the largest to date with 196 language pairs.
BalsuTalka.lv - Boosting the Common Voice Corpus for Low-Resource Languages (2024.lrec-main)

Copied to clipboard

Challenge: Latvian is a low-resource language for many NLP tasks, but most speech corpora are closed data . a crowdsourcing campaign to create a relatively large, diverse and open speech corpus for Latvian has been launched .
Approach: a crowdsourcing campaign is helping to create an open speech corpus for Latvian . the goal is to enlarge the datasets and make them more diverse . authors use the opensource Mozilla Common Voice platform to validate speech samples .
Outcome: a crowdsourcing initiative has increased the size and speaker diversity of the Latvian Common Voice 17.0 dataset by more than tenfold in less than a year.
Speech-to-Speech Translation for a Real-world Unwritten Language (2023.findings-acl)

Copied to clipboard

Challenge: a new study examines speech-to-speech translation (S2ST) that translates speech from one language into another . the research area for unwritten languages remains a research area with little exploration due to the lack of training data.
Approach: They propose a system that translates speech from one language into another . they use Taiwanese Hokkien as an example of an unwritten language .
Outcome: The proposed system can be used to train models in languages without standard writing systems.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations