Challenge: a study of Swiss German speech translation systems focuses on dialect diversity and differences between Swiss German and Standard German.
Approach: They focus on the impact of dialect diversity and differences between Swiss German and Standard German . they first review the Swiss German dialect landscape and the differences to Standard German.
Outcome: The proposed model is based on the Swiss German dialect landscape and differences to Standard German.

Similar Papers

Machine Translation of Low-Resource Spoken Dialects: Strategies for Normalizing Swiss German (L18-1)

Copied to clipboard

Challenge: Using character-based neural MT, we normalize Swiss German input to address regional diversity.
Approach: They propose to use character-based neural MT to normalize Swiss German input and phrase-based statistical MT for a low-resource family of dialects.
Outcome: The proposed system achieves 36% BLEU score when translating from the Bernese dialect.
A Swiss German Dictionary: Variation in Speech and Writing (2020.lrec-1)

Copied to clipboard

Challenge: Besides standard German, Swiss German is spoken in about two thirds of Switzerland.
Approach: They propose a dictionary containing normalized forms of common Swiss German words paired with Swiss German phonetic transcriptions to alleviate the uncertainty associated with this diversity.
Outcome: The proposed dictionary is the first to combine spontaneous translation and phonetic transcriptions in large-scale, scalable phoneme to grapheme model that generates credible novel Swiss German writings.
SDS-200: A Swiss German Speech to Standard German Text Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Using a web recording tool, participants were asked to translate their Swiss German text to their own dialect before recording it.
Approach: They present a corpus of Swiss German dialectal speech with Standard German text translations . the dataset allows for training speech translation, dialect recognition, and speech synthesis systems .
Outcome: The dataset allows for training speech translation, dialect recognition, and speech synthesis systems.
Data-Driven Pronunciation Modeling of Swiss German Dialectal Speech for Automatic Speech Recognition (L18-1)

Copied to clipboard

Challenge: a Swiss German speech recognizer is trained using a standard German annotation model.
Approach: They propose to train a Swiss German speech recognition system using a standard German annotation model.
Outcome: The proposed system is based on a standard German annotation model and a grapheme-to-phoneme conversion model.
STT4SG-350: A Speech Corpus for All Swiss German Dialect Regions (2023.acl-short)

Copied to clipboard

Challenge: We present a corpus of Swiss German speech annotated with Standard German text at the sentence level.
Approach: They present a corpus of Swiss German speech annotated with Standard German sentences . they use a web app to show the speakers standard German sentences and record them .
Outcome: The corpus contains 343 hours of speech from all Swiss German dialect regions . it is the largest public speech corpus for Swiss German to date .
Standard-to-Dialect Transfer Trends Differ across Text and Speech: A Case Study on Intent and Topic Classification in German Dialects (2026.acl-long)

Copied to clipboard

Challenge: Research on cross-dialectal transfer from a standard to a non-standard dialect variety has typically focused on text data.
Approach: They compare standard-to-dialect transfer in three settings: text models, speech models, and cascaded systems where speech first gets automatically transcribed and then further processed by a text model.
Outcome: The proposed model performs best on German dialect data while the text-only model perform best on the standard data.
LibriS2S: A German-English Speech-to-Speech Translation Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Recent advances in speech-to-text translation have led to significant improvements, but the availability of appropriate training data is limiting.
Approach: They propose a new text-to-speech and speech-tospech translation model that directly learns to generate the speech signal based on the pronunciation of the source language.
Outcome: The proposed model learns to generate speech signal based on pronunciation of source language.
Strategies and Challenges for Crowdsourcing Regional Dialect Perception Data for Swiss German and Swiss French (L18-1)

Copied to clipboard

Challenge: a crowdsourcing project in the field of Swiss German dialects and Swiss French accents collects linguistic data.
Approach: a gamified crowdsourcing platform was set up to collect linguistic data on Swiss German and Swiss French accents.
Outcome: a gamified crowdsourcing platform collects linguistic data on Swiss German and Swiss French accents . the platform has provided 470,000 localizations, with 7,500 registered users and 30,000 anonymous visitors .
Quantifying the Dialect Gap and its Correlates Across Languages (2023.findings-emnlp)

Copied to clipboard

Challenge: Historically, studies investigating minority variants of languages have been limited to a select few languages.
Approach: They evaluate state-of-the-art large language models for regional dialects of several high- and low-resource languages and analyze how regional dialect gap is correlated with economic, social, and linguistic factors.
Outcome: The proposed model is compared with two high-use applications and shows that it can solve the regional dialect gap.
Dialect-to-Standard Normalization: A Large-Scale Multilingual Evaluation (2023.findings-emnlp)

Copied to clipboard

Challenge: Text normalization is a range of tasks that consist in replacing non-standard spellings with their standard equivalents.
Approach: They introduce dialect-to-standard normalization as a sentence-level character transduction task and provide a large-scale analysis of these methods.
Outcome: The proposed model performs best for Finnish, Swiss German and Slovene while the pre-trained model using full sentences performs the best for Norwegian.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations