Papers with G2P
Making a Point: Pointer-Generator Transformers for Disjoint Vocabularies (2020.aacl-srw)
Copied to clipboard
| Challenge: | Existing neural models rely on an overlap between source and target vocabularies to perform sequence-to-sequence tasks. |
| Approach: | They propose a pointer-generator transformer model for disjoint vocabularies that does not rely on an overlap between source and target vocs. |
| Outcome: | The proposed model outperforms a standard pointer-generator transformer by an average of 5.1 WER over 15 languages. |
Fast Bilingual Grapheme-To-Phoneme Conversion (2022.naacl-industry)
Copied to clipboard
| Challenge: | Autoregressive transformers (ART)-based grapheme-to-phoneme (G2P) models have been proposed for bi/multilingual text-to speech systems. |
| Approach: | They propose a bilingual grapheme-to-phoneme (G2P) model with autoregressive transformers for fast and exact decoding and data augmentation for predicting output length. |
| Outcome: | The proposed model achieves better performance than the previous model and 2700% faster inference speed. |
Zero-shot Learning for Grapheme to Phoneme Conversion with Language Ensemble (2022.findings-acl)
Copied to clipboard
| Challenge: | Existing work focuses on low-resource and endangered languages with limited training sets. |
| Approach: | They propose a hypothesis set for any unseen target language and combine it with a confusion network to propose 'the most likely hypothesis' they test the approach on over 600 unseened languages and demonstrate it significantly outperforms baselines. |
| Outcome: | The proposed model outperforms baselines on over 600 unseen languages. |
GE2PE: Persian End-to-End Grapheme-to-Phoneme Conversion (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing text-to-speech systems struggle to produce natural speech from grapheme sequences . Grapheme-to phoneme conversion (G2P) systems face limitations when dealing with Persian texts due to the complexity of Persian transcription. |
| Approach: | They propose to use phonetic information to enhance the input sequence for Persian translations. |
| Outcome: | The proposed model surpasses state-of-the-art models by 1.86% in word error rate and 3.42% in homograph disambiguation accuracy. |
DLM: A Decoupled Learning Model for Long-tailed Polyphone Disambiguation in Mandarin (2024.naacl-long)
Copied to clipboard
| Challenge: | Grapheme-to-phoneme conversion datasets suffer from the long-tail problem . context learning for polyphonic characters often stems from a single dimension . |
| Approach: | They propose a model for long-tailed polyphone disambiguation in Mandarin that decouples representation and classification learnings. |
| Outcome: | The proposed model can decouple representation and classification learnings . it achieves transition learning of context from local to global . |
Epitran: Precision G2P for Many Languages (L18-1)
Copied to clipboard
| Challenge: | Epitran is a multilingual, multi-back-end system for grapheme-to-phoneme transduction . it supports 61 languages and is open source under an MIT license . |
| Approach: | Epitran is a multilingual back-end system for grapheme-to-phoneme transduction . it takes word tokens in the orthography of a language and outputs a phonemic representation . Epitran's efficacy has been demonstrated in multiple research projects . |
| Outcome: | Epitran is a multilingual, multi-backend system for grapheme-to-phoneme transduction . it supports 61 languages and is open source under MIT license . |
Pronunciation Dictionaries for the Alsatian Dialects to Analyze Spelling and Phonetic Variation (L18-1)
Copied to clipboard
| Challenge: | a new study compares phonetic transcriptions of Alsatian, German and French with existing pronunciation dictionaries . Alsatic dialects do not have a standardized spelling system, despite literary history dating back to the 19th century . |
| Approach: | They propose new pronunciation dictionaries for the under-resourced Alsatian dialects . they compare them with existing phonetic transcriptions of Alsalian, German and French . |
| Outcome: | The proposed dictionaries are compared with existing phonetic transcriptions of Alsatian, German and French to examine the relationship between speech and writing. |
POWSM: A Phonetic Open Whisper-Style Speech Foundation Model (2026.acl-long)
Copied to clipboard
Chin-Jou Li, Kalvin Chang, Shikhar Bharadwaj, Eunjung Yeo, Kwanghee Choi, Jian Zhu, David R. Mortensen, Shinji Watanabe
| Challenge: | Phone-level modeling of speech is a common approach to speech recognition, but it relies on task-specific architectures and datasets. |
| Approach: | They propose a phonetic framework capable of performing multiple phone-related tasks . they propose 'Phonetic Open Whisper-style Speech Model' that can perform these tasks together . |
| Outcome: | The proposed model outperforms or matches specialized PR models of similar size while supporting G2P, P2G, and ASR. |
Would LLMs be Good Historical Linguists and Chinese Dialect Learners? (2026.acl-long)
Copied to clipboard
| Challenge: | Large language models struggle with low-resource Chinese dialects due to substantial phonological divergence. |
| Approach: | They propose to incorporate Middle Chinese, the common historical ancestor of modern Chinese dialects, into LLMs to improve dialectal pronunciation modeling. |
| Outcome: | The proposed approach improves on standard Chinese but struggles with low-resource Chinese dialects . the proposed model improves over baselines while revealing variation across dialects. |
NB Uttale: A Norwegian Pronunciation Lexicon with Dialect Variation (2024.lrec-main)
Copied to clipboard
| Challenge: | lexicon is based on the NST Bokml lexiconic for East Norwegian . lexica are an essential linguistic resource in speech recognition and speech synthesis systems . |
| Approach: | They propose to use Bokml orthographic word forms and up to eight alternate phonological transcriptions per word form to generate a Norwegian pronunciation lexicon. |
| Outcome: | The proposed model improves the accuracy of the proposed model and its outputs with word- and phoneme-error-rate metrics. |