Challenge: Besides standard German, Swiss German is spoken in about two thirds of Switzerland.
Approach: They propose a dictionary containing normalized forms of common Swiss German words paired with Swiss German phonetic transcriptions to alleviate the uncertainty associated with this diversity.
Outcome: The proposed dictionary is the first to combine spontaneous translation and phonetic transcriptions in large-scale, scalable phoneme to grapheme model that generates credible novel Swiss German writings.

Similar Papers

Dialect Transfer for Swiss German Speech Translation (2023.findings-emnlp)

Copied to clipboard

Challenge: a study of Swiss German speech translation systems focuses on dialect diversity and differences between Swiss German and Standard German.
Approach: They focus on the impact of dialect diversity and differences between Swiss German and Standard German . they first review the Swiss German dialect landscape and the differences to Standard German.
Outcome: The proposed model is based on the Swiss German dialect landscape and differences to Standard German.
Machine Translation of Low-Resource Spoken Dialects: Strategies for Normalizing Swiss German (L18-1)

Copied to clipboard

Challenge: Using character-based neural MT, we normalize Swiss German input to address regional diversity.
Approach: They propose to use character-based neural MT to normalize Swiss German input and phrase-based statistical MT for a low-resource family of dialects.
Outcome: The proposed system achieves 36% BLEU score when translating from the Bernese dialect.
SDS-200: A Swiss German Speech to Standard German Text Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Using a web recording tool, participants were asked to translate their Swiss German text to their own dialect before recording it.
Approach: They present a corpus of Swiss German dialectal speech with Standard German text translations . the dataset allows for training speech translation, dialect recognition, and speech synthesis systems .
Outcome: The dataset allows for training speech translation, dialect recognition, and speech synthesis systems.
Data-Driven Pronunciation Modeling of Swiss German Dialectal Speech for Automatic Speech Recognition (L18-1)

Copied to clipboard

Challenge: a Swiss German speech recognizer is trained using a standard German annotation model.
Approach: They propose to train a Swiss German speech recognition system using a standard German annotation model.
Outcome: The proposed system is based on a standard German annotation model and a grapheme-to-phoneme conversion model.
STT4SG-350: A Speech Corpus for All Swiss German Dialect Regions (2023.acl-short)

Copied to clipboard

Challenge: We present a corpus of Swiss German speech annotated with Standard German text at the sentence level.
Approach: They present a corpus of Swiss German speech annotated with Standard German sentences . they use a web app to show the speakers standard German sentences and record them .
Outcome: The corpus contains 343 hours of speech from all Swiss German dialect regions . it is the largest public speech corpus for Swiss German to date .
Standard German Subtitling of Swiss German TV content: the PASSAGE Project (2022.lrec-1)

Copied to clipboard

Challenge: In Switzerland, two thirds of the population speak Swiss German, a primarily spoken language with no standardised written form.
Approach: They propose to combine a speech recognition system with an intralingual machine translation system to automate the subtitling process.
Outcome: The proposed systems improve the quality of the standardized Swiss German subtitles but are not capable of producing correct Standard German.
Make Every Letter Count: Building Dialect Variation Dictionaries from Monolingual Corpora (2025.findings-emnlp)

Copied to clipboard

Challenge: Dialects exhibit a substantial degree of variation due to the lack of a standard orthography . however, the ability of Large Language Models (LLMs) to process dialects remains understudied .
Approach: They propose a framework for creating dialect variation dictionaries from monolingual data . they use a dataset to examine how well LLMs can judge Bavarian terms as dialect translations .
Outcome: The proposed framework can judge dialects as dialect translations, inflected variants or unrelated forms of a given German lemma.
Pronunciation Dictionaries for the Alsatian Dialects to Analyze Spelling and Phonetic Variation (L18-1)

Copied to clipboard

Challenge: a new study compares phonetic transcriptions of Alsatian, German and French with existing pronunciation dictionaries . Alsatic dialects do not have a standardized spelling system, despite literary history dating back to the 19th century .
Approach: They propose new pronunciation dictionaries for the under-resourced Alsatian dialects . they compare them with existing phonetic transcriptions of Alsalian, German and French .
Outcome: The proposed dictionaries are compared with existing phonetic transcriptions of Alsatian, German and French to examine the relationship between speech and writing.
LibriS2S: A German-English Speech-to-Speech Translation Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Recent advances in speech-to-text translation have led to significant improvements, but the availability of appropriate training data is limiting.
Approach: They propose a new text-to-speech and speech-tospech translation model that directly learns to generate the speech signal based on the pronunciation of the source language.
Outcome: The proposed model learns to generate speech signal based on pronunciation of source language.
NB Uttale: A Norwegian Pronunciation Lexicon with Dialect Variation (2024.lrec-main)

Copied to clipboard

Challenge: lexicon is based on the NST Bokml lexiconic for East Norwegian . lexica are an essential linguistic resource in speech recognition and speech synthesis systems .
Approach: They propose to use Bokml orthographic word forms and up to eight alternate phonological transcriptions per word form to generate a Norwegian pronunciation lexicon.
Outcome: The proposed model improves the accuracy of the proposed model and its outputs with word- and phoneme-error-rate metrics.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations