Challenge: a corpus of 250 hours of speech in both west frisian and Dutch is being developed . the corpus is a sociolinguistic corpus based on the recordings of four generations of bilingual speakers .
Approach: This paper describes the Boarnsterhim Corpus project which started in 2016 . it aims to make available 250 hours of speech in both west frisian and Dutch by same speakers .
Outcome: The Boarnsterhim Corpus is a sociolinguistic corpus of west frisian and Dutch speakers . it spans four generations and includes panel and trend data .

Similar Papers

Corpus Creation and Automatic Alignment of Historical Dutch Dialect Speech (2024.lrec-main)

Copied to clipboard

Challenge: The Dutch Dialect Database contains dialectal variations of Dutch recorded in the second half of the twentieth century.
Approach: They propose to create a corpus containing audio recordings and orthographic transcriptions of Dutch dialects recorded in the second half of the 20th century.
Outcome: The Dutch Dialect Database contains dialectal variations recorded all over the Netherlands in the second half of the twentieth century.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
Corpus REDEWIEDERGABE (2020.lrec-1)

Copied to clipboard

Challenge: The corpus REDEWIEDERGABE contains detailed annotations for speech, thought and writing representation (ST&WR) with approximately 490,000 tokens, it is the largest resource of its kind.
Approach: This paper presents corpus REDEWIEDERGABE, a German-language historical corpus with detailed annotations for speech, thought and writing representation (ST&WR).
Outcome: The corpus REDEWIEDERGABE contains 490,000 tokens and is the largest resource of its kind.
GPT-SW3: An Autoregressive Language Model for the Scandinavian Languages (2024.lrec-main)

Copied to clipboard

Challenge: a growing interest in building and applying large language models for languages other than English is fueling interest in developing LLMs for smaller languages.
Approach: They describe the development process for the first native large generative language model for the North Germanic languages, GPT-SW3.
Outcome: The proposed model is based on the generative language model for the North Germanic languages . it is a first-generation model with a high-quality data set and a low cost of implementation .
LibriVoxDeEn: A Corpus for German-to-English Speech Translation and German Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of sentence-aligned triples of German audio, German text, and English translation is available for speech recognition . a large corpus is available to date for end-to-end speech translation based on parallel data .
Approach: They present a corpus of sentence-aligned triples of German audio, German text, and English translation based on German audio books.
Outcome: The proposed corpus is the largest resource for German speech recognition and for end-to-end German-to English speech translation.
100,000 Podcasts: A Spoken English Document Corpus (2020.coling-main)

Copied to clipboard

Challenge: Podcasts are a large and growing repository of spoken audio.
Approach: They propose to use podcasts as a resource for speech processing and linguistics . they use a corpus of 100,000 podcasts to study the complexity of the domain .
Outcome: The Spotify Podcast Dataset is the largest corpus of transcribed speech data . the dataset contains 60,000 hours of podcasts, with a range of genres and styles .
Merkel Podcast Corpus: A Multimodal Dataset Compiled from 16 Years of Angela Merkel’s Weekly Video Podcasts (2022.lrec-1)

Copied to clipboard

Challenge: a dataset of 16 years of (almost) weekly Internet podcasts of former german chancellor Angela Merkel is presented.
Approach: They propose to curate a German podcast corpus from 16 years of podcasts of former german chancellor Angela Merkel using audio-visual-text methods.
Outcome: The proposed pipeline can be used to curate other datasets of similar nature, such as talk show contents.
BasqueParl: A Bilingual Corpus of Basque Parliamentary Transcriptions (2022.lrec-1)

Copied to clipboard

Challenge: a new corpus of Basque parliamentary transcripts is released to study political discourse in contrasting languages . a corpus containing political discourses from public institutions can be used for computational social science research .
Approach: They present a corpus from Basque parliamentary transcripts and enrich it with metadata related to relevant attributes of speakers and speeches.
Outcome: The proposed corpus is characterized by heavy Basque-Spanish code-switching . it provides interesting insights about language use of political representatives across time, parties and gender .
Extrinsic Evaluation of French Dependency Parsers on a Specialized Corpus: Comparison of Distributional Thesauri (2020.lrec-1)

Copied to clipboard

Challenge: Using a frequency-based method, we can identify subsets of the same word contexts without any reference data.
Approach: They compare 11 different French dependency parsers on a specialized corpus to generate distributional thesauri using a frequency-based method.
Outcome: The proposed method can identify relevant subsets without reference data and the similarity is confirmed on a restricted distributional benchmark.
Writing System and Speaker Metadata for 2,800+ Language Varieties (2022.lrec-1)

Copied to clipboard

Challenge: Currently, language technologies are easily available in only a small minority of the world's 7,000+ language varieties.
Approach: They propose to use an open-source dataset to provide the writing system(s) for each of the 2,800+ languages used in the world today and an estimated speaker count for each.
Outcome: The dataset provides the attested writing system(s) for each of these 2,800+ varieties, as well as an estimated speaker count for each variety.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations