The Boarnsterhim Corpus: A Bilingual Frisian-Dutch Panel and Trend Study (L18-1)
Copied to clipboard
| Challenge: | a corpus of 250 hours of speech in both west frisian and Dutch is being developed . the corpus is a sociolinguistic corpus based on the recordings of four generations of bilingual speakers . |
| Approach: | This paper describes the Boarnsterhim Corpus project which started in 2016 . it aims to make available 250 hours of speech in both west frisian and Dutch by same speakers . |
| Outcome: | The Boarnsterhim Corpus is a sociolinguistic corpus of west frisian and Dutch speakers . it spans four generations and includes panel and trend data . |
Similar Papers
Corpus Creation and Automatic Alignment of Historical Dutch Dialect Speech (2024.lrec-main)
Copied to clipboard
| Challenge: | The Dutch Dialect Database contains dialectal variations of Dutch recorded in the second half of the twentieth century. |
| Approach: | They propose to create a corpus containing audio recordings and orthographic transcriptions of Dutch dialects recorded in the second half of the 20th century. |
| Outcome: | The Dutch Dialect Database contains dialectal variations recorded all over the Netherlands in the second half of the twentieth century. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
Corpus REDEWIEDERGABE (2020.lrec-1)
Copied to clipboard
| Challenge: | The corpus REDEWIEDERGABE contains detailed annotations for speech, thought and writing representation (ST&WR) with approximately 490,000 tokens, it is the largest resource of its kind. |
| Approach: | This paper presents corpus REDEWIEDERGABE, a German-language historical corpus with detailed annotations for speech, thought and writing representation (ST&WR). |
| Outcome: | The corpus REDEWIEDERGABE contains 490,000 tokens and is the largest resource of its kind. |
GPT-SW3: An Autoregressive Language Model for the Scandinavian Languages (2024.lrec-main)
Copied to clipboard
Ariel Ekgren, Amaru Cuba Gyllensten, Felix Stollenwerk, Joey Öhman, Tim Isbister, Evangelia Gogoulou, Fredrik Carlsson, Judit Casademont, Magnus Sahlgren
| Challenge: | a growing interest in building and applying large language models for languages other than English is fueling interest in developing LLMs for smaller languages. |
| Approach: | They describe the development process for the first native large generative language model for the North Germanic languages, GPT-SW3. |
| Outcome: | The proposed model is based on the generative language model for the North Germanic languages . it is a first-generation model with a high-quality data set and a low cost of implementation . |
LibriVoxDeEn: A Corpus for German-to-English Speech Translation and German Speech Recognition (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of sentence-aligned triples of German audio, German text, and English translation is available for speech recognition . a large corpus is available to date for end-to-end speech translation based on parallel data . |
| Approach: | They present a corpus of sentence-aligned triples of German audio, German text, and English translation based on German audio books. |
| Outcome: | The proposed corpus is the largest resource for German speech recognition and for end-to-end German-to English speech translation. |
100,000 Podcasts: A Spoken English Document Corpus (2020.coling-main)
Copied to clipboard
Ann Clifton, Sravana Reddy, Yongze Yu, Aasish Pappu, Rezvaneh Rezapour, Hamed Bonab, Maria Eskevich, Gareth Jones, Jussi Karlgren, Ben Carterette, Rosie Jones
| Challenge: | Podcasts are a large and growing repository of spoken audio. |
| Approach: | They propose to use podcasts as a resource for speech processing and linguistics . they use a corpus of 100,000 podcasts to study the complexity of the domain . |
| Outcome: | The Spotify Podcast Dataset is the largest corpus of transcribed speech data . the dataset contains 60,000 hours of podcasts, with a range of genres and styles . |
Merkel Podcast Corpus: A Multimodal Dataset Compiled from 16 Years of Angela Merkel’s Weekly Video Podcasts (2022.lrec-1)
Copied to clipboard
| Challenge: | a dataset of 16 years of (almost) weekly Internet podcasts of former german chancellor Angela Merkel is presented. |
| Approach: | They propose to curate a German podcast corpus from 16 years of podcasts of former german chancellor Angela Merkel using audio-visual-text methods. |
| Outcome: | The proposed pipeline can be used to curate other datasets of similar nature, such as talk show contents. |
BasqueParl: A Bilingual Corpus of Basque Parliamentary Transcriptions (2022.lrec-1)
Copied to clipboard
Nayla Escribano, Jon Ander Gonzalez, Julen Orbegozo-Terradillos, Ainara Larrondo-Ureta, Simón Peña-Fernández, Olatz Perez-de-Viñaspre, Rodrigo Agerri
| Challenge: | a new corpus of Basque parliamentary transcripts is released to study political discourse in contrasting languages . a corpus containing political discourses from public institutions can be used for computational social science research . |
| Approach: | They present a corpus from Basque parliamentary transcripts and enrich it with metadata related to relevant attributes of speakers and speeches. |
| Outcome: | The proposed corpus is characterized by heavy Basque-Spanish code-switching . it provides interesting insights about language use of political representatives across time, parties and gender . |
Extrinsic Evaluation of French Dependency Parsers on a Specialized Corpus: Comparison of Distributional Thesauri (2020.lrec-1)
Copied to clipboard
| Challenge: | Using a frequency-based method, we can identify subsets of the same word contexts without any reference data. |
| Approach: | They compare 11 different French dependency parsers on a specialized corpus to generate distributional thesauri using a frequency-based method. |
| Outcome: | The proposed method can identify relevant subsets without reference data and the similarity is confirmed on a restricted distributional benchmark. |
Writing System and Speaker Metadata for 2,800+ Language Varieties (2022.lrec-1)
Copied to clipboard
| Challenge: | Currently, language technologies are easily available in only a small minority of the world's 7,000+ language varieties. |
| Approach: | They propose to use an open-source dataset to provide the writing system(s) for each of the 2,800+ languages used in the world today and an estimated speaker count for each. |
| Outcome: | The dataset provides the attested writing system(s) for each of these 2,800+ varieties, as well as an estimated speaker count for each variety. |