Larissa Schmidt, Lucy Linder, Sandra Djambazovska, Alexandros Lazaridis, Tanja Samardžić, Claudiu Musat
| Challenge: | Besides standard German, Swiss German is spoken in about two thirds of Switzerland. |
| Approach: | They propose a dictionary containing normalized forms of common Swiss German words paired with Swiss German phonetic transcriptions to alleviate the uncertainty associated with this diversity. |
| Outcome: | The proposed dictionary is the first to combine spontaneous translation and phonetic transcriptions in large-scale, scalable phoneme to grapheme model that generates credible novel Swiss German writings. |
Similar Papers
Dialect Transfer for Swiss German Speech Translation (2023.findings-emnlp)
Copied to clipboard
| Challenge: | a study of Swiss German speech translation systems focuses on dialect diversity and differences between Swiss German and Standard German. |
| Approach: | They focus on the impact of dialect diversity and differences between Swiss German and Standard German . they first review the Swiss German dialect landscape and the differences to Standard German. |
| Outcome: | The proposed model is based on the Swiss German dialect landscape and differences to Standard German. |
Machine Translation of Low-Resource Spoken Dialects: Strategies for Normalizing Swiss German (L18-1)
Copied to clipboard
| Challenge: | Using character-based neural MT, we normalize Swiss German input to address regional diversity. |
| Approach: | They propose to use character-based neural MT to normalize Swiss German input and phrase-based statistical MT for a low-resource family of dialects. |
| Outcome: | The proposed system achieves 36% BLEU score when translating from the Bernese dialect. |
SDS-200: A Swiss German Speech to Standard German Text Corpus (2022.lrec-1)
Copied to clipboard
Michel Plüss, Manuela Hürlimann, Marc Cuny, Alla Stöckli, Nikolaos Kapotis, Julia Hartmann, Malgorzata Anna Ulasik, Christian Scheller, Yanick Schraner, Amit Jain, Jan Deriu, Mark Cieliebak, Manfred Vogel
| Challenge: | Using a web recording tool, participants were asked to translate their Swiss German text to their own dialect before recording it. |
| Approach: | They present a corpus of Swiss German dialectal speech with Standard German text translations . the dataset allows for training speech translation, dialect recognition, and speech synthesis systems . |
| Outcome: | The dataset allows for training speech translation, dialect recognition, and speech synthesis systems. |
Data-Driven Pronunciation Modeling of Swiss German Dialectal Speech for Automatic Speech Recognition (L18-1)
Copied to clipboard
| Challenge: | a Swiss German speech recognizer is trained using a standard German annotation model. |
| Approach: | They propose to train a Swiss German speech recognition system using a standard German annotation model. |
| Outcome: | The proposed system is based on a standard German annotation model and a grapheme-to-phoneme conversion model. |
STT4SG-350: A Speech Corpus for All Swiss German Dialect Regions (2023.acl-short)
Copied to clipboard
Michel Plüss, Jan Deriu, Yanick Schraner, Claudio Paonessa, Julia Hartmann, Larissa Schmidt, Christian Scheller, Manuela Hürlimann, Tanja Samardžić, Manfred Vogel, Mark Cieliebak
| Challenge: | We present a corpus of Swiss German speech annotated with Standard German text at the sentence level. |
| Approach: | They present a corpus of Swiss German speech annotated with Standard German sentences . they use a web app to show the speakers standard German sentences and record them . |
| Outcome: | The corpus contains 343 hours of speech from all Swiss German dialect regions . it is the largest public speech corpus for Swiss German to date . |
Standard German Subtitling of Swiss German TV content: the PASSAGE Project (2022.lrec-1)
Copied to clipboard
| Challenge: | In Switzerland, two thirds of the population speak Swiss German, a primarily spoken language with no standardised written form. |
| Approach: | They propose to combine a speech recognition system with an intralingual machine translation system to automate the subtitling process. |
| Outcome: | The proposed systems improve the quality of the standardized Swiss German subtitles but are not capable of producing correct Standard German. |
Make Every Letter Count: Building Dialect Variation Dictionaries from Monolingual Corpora (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Dialects exhibit a substantial degree of variation due to the lack of a standard orthography . however, the ability of Large Language Models (LLMs) to process dialects remains understudied . |
| Approach: | They propose a framework for creating dialect variation dictionaries from monolingual data . they use a dataset to examine how well LLMs can judge Bavarian terms as dialect translations . |
| Outcome: | The proposed framework can judge dialects as dialect translations, inflected variants or unrelated forms of a given German lemma. |
Pronunciation Dictionaries for the Alsatian Dialects to Analyze Spelling and Phonetic Variation (L18-1)
Copied to clipboard
| Challenge: | a new study compares phonetic transcriptions of Alsatian, German and French with existing pronunciation dictionaries . Alsatic dialects do not have a standardized spelling system, despite literary history dating back to the 19th century . |
| Approach: | They propose new pronunciation dictionaries for the under-resourced Alsatian dialects . they compare them with existing phonetic transcriptions of Alsalian, German and French . |
| Outcome: | The proposed dictionaries are compared with existing phonetic transcriptions of Alsatian, German and French to examine the relationship between speech and writing. |
LibriS2S: A German-English Speech-to-Speech Translation Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Recent advances in speech-to-text translation have led to significant improvements, but the availability of appropriate training data is limiting. |
| Approach: | They propose a new text-to-speech and speech-tospech translation model that directly learns to generate the speech signal based on the pronunciation of the source language. |
| Outcome: | The proposed model learns to generate speech signal based on pronunciation of source language. |
NB Uttale: A Norwegian Pronunciation Lexicon with Dialect Variation (2024.lrec-main)
Copied to clipboard
| Challenge: | lexicon is based on the NST Bokml lexiconic for East Norwegian . lexica are an essential linguistic resource in speech recognition and speech synthesis systems . |
| Approach: | They propose to use Bokml orthographic word forms and up to eight alternate phonological transcriptions per word form to generate a Norwegian pronunciation lexicon. |
| Outcome: | The proposed model improves the accuracy of the proposed model and its outputs with word- and phoneme-error-rate metrics. |