Challenge: GRAIN contains German radio interviews and is annotated on multiple linguistic layers.
Approach: They present GRAIN as part of the SFB732 Silver Standard Collection . GRAIN contains German radio interviews and is annotated on multiple linguistic layers .
Outcome: The GRAIN data set contains German radio interviews and is annotated on multiple linguistic layers.

Similar Papers

GRAIN-S: Manually Annotated Syntax for German Interviews (2020.lrec-1)

Copied to clipboard

Challenge: GRAIN-S is a set of manually created syntactic annotations for radio interviews in germany.
Approach: They propose to use GRAIN-S to create syntactic annotations for radio interviews in germany.
Outcome: The proposed dataset extends an existing corpus GRAIN and comes with constituency and dependency trees for six interviews.
Fine-grained Named Entity Annotations for German Biographic Interviews (2020.lrec-1)

Copied to clipboard

Challenge: a NER annotation scheme is adapted for a corpus of transcripts of biographic interviews with emigrants to German . a dataset of spoken data and teaser tweets from newspaper sites are used to test the NER inventory.
Approach: They propose a fine-grained NER annotation scheme with 30 labels and apply it to German data.
Outcome: The proposed NER annotations can be applied to spoken data and teaser tweets from newspaper sites and achieve good inter-annotator agreement.
Introducing a Parsed Corpus of Historical High German (2024.lrec-main)

Copied to clipboard

Challenge: outlines the development of the Indiana Parsed Corpus of (Historical) High German . outlines selection of texts, decisions on part-of-speech tags and other labels .
Approach: They propose to build a parsed German corpus that spans Germanic from 1050 to 1950 . they propose to use Penn-style treebanks to capture syntactic relationships between words .
Outcome: The proposed corpus spans Germanic languages from 1050 to 1950 and illustrative annotation issues unique to the language.
Merkel Podcast Corpus: A Multimodal Dataset Compiled from 16 Years of Angela Merkel’s Weekly Video Podcasts (2022.lrec-1)

Copied to clipboard

Challenge: a dataset of 16 years of (almost) weekly Internet podcasts of former german chancellor Angela Merkel is presented.
Approach: They propose to curate a German podcast corpus from 16 years of podcasts of former german chancellor Angela Merkel using audio-visual-text methods.
Outcome: The proposed pipeline can be used to curate other datasets of similar nature, such as talk show contents.
Corpus REDEWIEDERGABE (2020.lrec-1)

Copied to clipboard

Challenge: The corpus REDEWIEDERGABE contains detailed annotations for speech, thought and writing representation (ST&WR) with approximately 490,000 tokens, it is the largest resource of its kind.
Approach: This paper presents corpus REDEWIEDERGABE, a German-language historical corpus with detailed annotations for speech, thought and writing representation (ST&WR).
Outcome: The corpus REDEWIEDERGABE contains 490,000 tokens and is the largest resource of its kind.
A Penn-style Treebank of Middle Low German (2020.lrec-1)

Copied to clipboard

Challenge: attestation for Middle Low German is rich, but its syntax remains relatively understudied.
Approach: They outline the issues involved in creating a Penn-style treebank of Middle Low German . they describe the background for the corpus and the process by which texts were selected .
Outcome: The proposed corpus will be a syntactically annotated treebank of Middle Low German . the proposed corpuse will be part of the Corpus of Historical Low German (CHLG) the proposed method will be used to generate strong empirical evidence for the language .
STT4SG-350: A Speech Corpus for All Swiss German Dialect Regions (2023.acl-short)

Copied to clipboard

Challenge: We present a corpus of Swiss German speech annotated with Standard German text at the sentence level.
Approach: They present a corpus of Swiss German speech annotated with Standard German sentences . they use a web app to show the speakers standard German sentences and record them .
Outcome: The corpus contains 343 hours of speech from all Swiss German dialect regions . it is the largest public speech corpus for Swiss German to date .
SDS-200: A Swiss German Speech to Standard German Text Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Using a web recording tool, participants were asked to translate their Swiss German text to their own dialect before recording it.
Approach: They present a corpus of Swiss German dialectal speech with Standard German text translations . the dataset allows for training speech translation, dialect recognition, and speech synthesis systems .
Outcome: The dataset allows for training speech translation, dialect recognition, and speech synthesis systems.
A Large Annotated Reference Corpus of New High German Poetry (2024.lrec-main)

Copied to clipboard

Challenge: a corpus of public domain German poetry covering the time period 1600 to the 1920s contains 65k unique poems and over 1.6M lines, each tokenized, syllabified, pos-tagged, and meter-tagged.
Approach: They present a large annotated corpus of public domain German poetry covering the time period 1600 to the 1920s with 65k poems.
Outcome: The corpus contains 65k unique poems and over 1.6M lines, each tokenized, syllabified, pos-tagged, and meter-tagged.
A Corpus Linguistic Perspective on Contemporary German Pop Lyrics with the Multi-Layer Annotated “Songkorpus” (2020.lrec-1)

Copied to clipboard

Challenge: TEI-compliant song lyrics are used as primary data, linguistically and literary motivated annotations, and extralinguistic metadata.
Approach: They propose to annotate a multiply annotated corpus of German lyrics as a publicly available basis for multidisciplinary research.
Outcome: The proposed corpus of german lyrics is available for evaluation and analysis using TEI-compliant, linguistically and literary motivated annotations and extralinguistic metadata.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations