Challenge: a new study compares phonetic transcriptions of Alsatian, German and French with existing pronunciation dictionaries . Alsatic dialects do not have a standardized spelling system, despite literary history dating back to the 19th century .
Approach: They propose new pronunciation dictionaries for the under-resourced Alsatian dialects . they compare them with existing phonetic transcriptions of Alsalian, German and French .
Outcome: The proposed dictionaries are compared with existing phonetic transcriptions of Alsatian, German and French to examine the relationship between speech and writing.

Similar Papers

A Swiss German Dictionary: Variation in Speech and Writing (2020.lrec-1)

Copied to clipboard

Challenge: Besides standard German, Swiss German is spoken in about two thirds of Switzerland.
Approach: They propose a dictionary containing normalized forms of common Swiss German words paired with Swiss German phonetic transcriptions to alleviate the uncertainty associated with this diversity.
Outcome: The proposed dictionary is the first to combine spontaneous translation and phonetic transcriptions in large-scale, scalable phoneme to grapheme model that generates credible novel Swiss German writings.
Automatic Identification of Maghreb Dialects Using a Dictionary-Based Approach (L18-1)

Copied to clipboard

Challenge: Automatic identification of Arabic dialects in texts is difficult, especially for Maghreb languages and when they are written in Arabic or Latin characters (Arabizi).
Approach: They propose a dictionary-based approach to detect Arabic dialects in texts . they focus on transliteration of Arabicizi into Latin script and code-switching .
Outcome: The proposed approach shows that it is possible to detect dialects in Arabic and Latin scripts.
Massively Multilingual Pronunciation Modeling with WikiPron (2020.lrec-1)

Copied to clipboard

Challenge: WikiPron is an open-source command-line tool for extracting pronunciation data from Wiktionary . the tool generates a database of 1.7 million pronunciations from 165 languages .
Approach: They propose a command-line tool for extracting pronunciation data from Wiktionary . they use it to generate a database of 1.7 million pronunciations from 165 languages .
Outcome: The proposed software generates a database of pronunciations for 165 languages . the proposed model is then validated by a grapheme-to-phoneme model .
NB Uttale: A Norwegian Pronunciation Lexicon with Dialect Variation (2024.lrec-main)

Copied to clipboard

Challenge: lexicon is based on the NST Bokml lexiconic for East Norwegian . lexica are an essential linguistic resource in speech recognition and speech synthesis systems .
Approach: They propose to use Bokml orthographic word forms and up to eight alternate phonological transcriptions per word form to generate a Norwegian pronunciation lexicon.
Outcome: The proposed model improves the accuracy of the proposed model and its outputs with word- and phoneme-error-rate metrics.
ELAL: An Emotion Lexicon for the Analysis of Alsatian Theatre Plays (2022.lrec-1)

Copied to clipboard

Challenge: a novel and manually corrected emotion lexicon is presented for Alsatian dialects . the dialects are used mainly orally and lack a stable and consensual spelling convention .
Approach: They propose a novel and manually corrected emotion lexicon for Alsatian dialects . they use graphical variants of Alsalian lexical items to perform automatic emotion analysis .
Outcome: The novel and manually corrected emotion lexicon is used to perform automatic emotion analysis in Alsatian theatre plays.
Visualizing the “Dictionary of Regionalisms of France” (DRF) (L18-1)

Copied to clipboard

Challenge: a corpus of regionalisms, parts of speech and recognition rates is published in the Dictionnaire des Régionalismes de France.
Approach: They propose to curate and analyze the corpus of regionalisms published in the Dictionnaire des Régionalismes de France.
Outcome: The corpus contains all entries in the DRF for which recognition rates were recorded . the analysis compares with previous work on regionalalisms and atlas .
Make Every Letter Count: Building Dialect Variation Dictionaries from Monolingual Corpora (2025.findings-emnlp)

Copied to clipboard

Challenge: Dialects exhibit a substantial degree of variation due to the lack of a standard orthography . however, the ability of Large Language Models (LLMs) to process dialects remains understudied .
Approach: They propose a framework for creating dialect variation dictionaries from monolingual data . they use a dataset to examine how well LLMs can judge Bavarian terms as dialect translations .
Outcome: The proposed framework can judge dialects as dialect translations, inflected variants or unrelated forms of a given German lemma.
Corpora with Part-of-Speech Annotations for Three Regional Languages of France: Alsatian, Occitan and Picard (L18-1)

Copied to clipboard

Challenge: RESTAURE project aims to develop resources and tools for three regional languages of France: Alsatian, Occitan and Picard.
Approach: They describe the creation of corpora with part-of-speech annotations for Alsatian, Occitan and Picard.
Outcome: The authors describe the creation of annotated corpora for Alsatian, Occitan and Picard . the project is part of the RESTAURE project, which aims to develop resources and tools for these under-resourced French regional languages.
A Speaking Atlas of the Regional Languages of France (L18-1)

Copied to clipboard

Challenge: a website is presented to show and promote the linguistic diversity of France through field recordings, a computer program and an orthographic transcription.
Approach: They propose to map linguistic diversity in France using field recordings and a computer program.
Outcome: The aim is to show and promote the linguistic diversity of France, through field recordings, a computer program and an orthographic transcription.
Arab Voices: Mapping Standard and Dialectal Arabic Speech Technology (2026.findings-acl)

Copied to clipboard

Challenge: Dialectal Arabic datasets embody a range of domain, dialect, and quality.
Approach: They propose a framework for automatic speech recognition in dialectal Arabic to address the limited data availability encountered in dialects.
Outcome: The proposed framework provides access to 31 datasets covering 14 dialects to better address the limited data availability encountered in dialectal Arabic speech processing.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations