German Radio Interviews: The GRAIN Release of the SFB732 Silver Standard Collection (L18-1)
Copied to clipboard
Katrin Schweitzer, Kerstin Eckart, Markus Gärtner, Agnieszka Falenska, Arndt Riester, Ina Rösiger, Antje Schweitzer, Sabrina Stehwien, Jonas Kuhn
| Challenge: | GRAIN contains German radio interviews and is annotated on multiple linguistic layers. |
| Approach: | They present GRAIN as part of the SFB732 Silver Standard Collection . GRAIN contains German radio interviews and is annotated on multiple linguistic layers . |
| Outcome: | The GRAIN data set contains German radio interviews and is annotated on multiple linguistic layers. |
Similar Papers
GRAIN-S: Manually Annotated Syntax for German Interviews (2020.lrec-1)
Copied to clipboard
| Challenge: | GRAIN-S is a set of manually created syntactic annotations for radio interviews in germany. |
| Approach: | They propose to use GRAIN-S to create syntactic annotations for radio interviews in germany. |
| Outcome: | The proposed dataset extends an existing corpus GRAIN and comes with constituency and dependency trees for six interviews. |
Fine-grained Named Entity Annotations for German Biographic Interviews (2020.lrec-1)
Copied to clipboard
| Challenge: | a NER annotation scheme is adapted for a corpus of transcripts of biographic interviews with emigrants to German . a dataset of spoken data and teaser tweets from newspaper sites are used to test the NER inventory. |
| Approach: | They propose a fine-grained NER annotation scheme with 30 labels and apply it to German data. |
| Outcome: | The proposed NER annotations can be applied to spoken data and teaser tweets from newspaper sites and achieve good inter-annotator agreement. |
Introducing a Parsed Corpus of Historical High German (2024.lrec-main)
Copied to clipboard
| Challenge: | outlines the development of the Indiana Parsed Corpus of (Historical) High German . outlines selection of texts, decisions on part-of-speech tags and other labels . |
| Approach: | They propose to build a parsed German corpus that spans Germanic from 1050 to 1950 . they propose to use Penn-style treebanks to capture syntactic relationships between words . |
| Outcome: | The proposed corpus spans Germanic languages from 1050 to 1950 and illustrative annotation issues unique to the language. |
Merkel Podcast Corpus: A Multimodal Dataset Compiled from 16 Years of Angela Merkel’s Weekly Video Podcasts (2022.lrec-1)
Copied to clipboard
| Challenge: | a dataset of 16 years of (almost) weekly Internet podcasts of former german chancellor Angela Merkel is presented. |
| Approach: | They propose to curate a German podcast corpus from 16 years of podcasts of former german chancellor Angela Merkel using audio-visual-text methods. |
| Outcome: | The proposed pipeline can be used to curate other datasets of similar nature, such as talk show contents. |
Corpus REDEWIEDERGABE (2020.lrec-1)
Copied to clipboard
| Challenge: | The corpus REDEWIEDERGABE contains detailed annotations for speech, thought and writing representation (ST&WR) with approximately 490,000 tokens, it is the largest resource of its kind. |
| Approach: | This paper presents corpus REDEWIEDERGABE, a German-language historical corpus with detailed annotations for speech, thought and writing representation (ST&WR). |
| Outcome: | The corpus REDEWIEDERGABE contains 490,000 tokens and is the largest resource of its kind. |
A Penn-style Treebank of Middle Low German (2020.lrec-1)
Copied to clipboard
| Challenge: | attestation for Middle Low German is rich, but its syntax remains relatively understudied. |
| Approach: | They outline the issues involved in creating a Penn-style treebank of Middle Low German . they describe the background for the corpus and the process by which texts were selected . |
| Outcome: | The proposed corpus will be a syntactically annotated treebank of Middle Low German . the proposed corpuse will be part of the Corpus of Historical Low German (CHLG) the proposed method will be used to generate strong empirical evidence for the language . |
STT4SG-350: A Speech Corpus for All Swiss German Dialect Regions (2023.acl-short)
Copied to clipboard
Michel Plüss, Jan Deriu, Yanick Schraner, Claudio Paonessa, Julia Hartmann, Larissa Schmidt, Christian Scheller, Manuela Hürlimann, Tanja Samardžić, Manfred Vogel, Mark Cieliebak
| Challenge: | We present a corpus of Swiss German speech annotated with Standard German text at the sentence level. |
| Approach: | They present a corpus of Swiss German speech annotated with Standard German sentences . they use a web app to show the speakers standard German sentences and record them . |
| Outcome: | The corpus contains 343 hours of speech from all Swiss German dialect regions . it is the largest public speech corpus for Swiss German to date . |
SDS-200: A Swiss German Speech to Standard German Text Corpus (2022.lrec-1)
Copied to clipboard
Michel Plüss, Manuela Hürlimann, Marc Cuny, Alla Stöckli, Nikolaos Kapotis, Julia Hartmann, Malgorzata Anna Ulasik, Christian Scheller, Yanick Schraner, Amit Jain, Jan Deriu, Mark Cieliebak, Manfred Vogel
| Challenge: | Using a web recording tool, participants were asked to translate their Swiss German text to their own dialect before recording it. |
| Approach: | They present a corpus of Swiss German dialectal speech with Standard German text translations . the dataset allows for training speech translation, dialect recognition, and speech synthesis systems . |
| Outcome: | The dataset allows for training speech translation, dialect recognition, and speech synthesis systems. |
A Large Annotated Reference Corpus of New High German Poetry (2024.lrec-main)
Copied to clipboard
| Challenge: | a corpus of public domain German poetry covering the time period 1600 to the 1920s contains 65k unique poems and over 1.6M lines, each tokenized, syllabified, pos-tagged, and meter-tagged. |
| Approach: | They present a large annotated corpus of public domain German poetry covering the time period 1600 to the 1920s with 65k poems. |
| Outcome: | The corpus contains 65k unique poems and over 1.6M lines, each tokenized, syllabified, pos-tagged, and meter-tagged. |
A Corpus Linguistic Perspective on Contemporary German Pop Lyrics with the Multi-Layer Annotated “Songkorpus” (2020.lrec-1)
Copied to clipboard
| Challenge: | TEI-compliant song lyrics are used as primary data, linguistically and literary motivated annotations, and extralinguistic metadata. |
| Approach: | They propose to annotate a multiply annotated corpus of German lyrics as a publicly available basis for multidisciplinary research. |
| Outcome: | The proposed corpus of german lyrics is available for evaluation and analysis using TEI-compliant, linguistically and literary motivated annotations and extralinguistic metadata. |