| Challenge: | The corpus REDEWIEDERGABE contains detailed annotations for speech, thought and writing representation (ST&WR) with approximately 490,000 tokens, it is the largest resource of its kind. |
| Approach: | This paper presents corpus REDEWIEDERGABE, a German-language historical corpus with detailed annotations for speech, thought and writing representation (ST&WR). |
| Outcome: | The corpus REDEWIEDERGABE contains 490,000 tokens and is the largest resource of its kind. |
Similar Papers
A Large Annotated Reference Corpus of New High German Poetry (2024.lrec-main)
Copied to clipboard
| Challenge: | a corpus of public domain German poetry covering the time period 1600 to the 1920s contains 65k unique poems and over 1.6M lines, each tokenized, syllabified, pos-tagged, and meter-tagged. |
| Approach: | They present a large annotated corpus of public domain German poetry covering the time period 1600 to the 1920s with 65k poems. |
| Outcome: | The corpus contains 65k unique poems and over 1.6M lines, each tokenized, syllabified, pos-tagged, and meter-tagged. |
Introducing a Parsed Corpus of Historical High German (2024.lrec-main)
Copied to clipboard
| Challenge: | outlines the development of the Indiana Parsed Corpus of (Historical) High German . outlines selection of texts, decisions on part-of-speech tags and other labels . |
| Approach: | They propose to build a parsed German corpus that spans Germanic from 1050 to 1950 . they propose to use Penn-style treebanks to capture syntactic relationships between words . |
| Outcome: | The proposed corpus spans Germanic languages from 1050 to 1950 and illustrative annotation issues unique to the language. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
An Annotated Social Media Corpus for German (2020.lrec-1)
Copied to clipboard
| Challenge: | Hate Speech (HS) against ethnic, religious and national minorities is a growing concern in online discourse. |
| Approach: | They present the German Twitter section of a large (2 billion word) bilingual Social Media corpus for Hate Speech research. |
| Outcome: | The proposed parser achieved F-scores of 97% for morphology and 92% for syntax on a cross-section of tweets. |
Manually Annotated Corpus of Polish Texts Published between 1830 and 1918 (L18-1)
Copied to clipboard
| Challenge: | a paper presents a manually annotated corpus of 625,000 tokens of Polish texts . the corpus provides three layers: transliteration, transcription and morphosyntactic annotation. |
| Approach: | The paper presents a manually annotated large historical corpus of Polish . the corpus provides three layers: transliteration, transcription and morphosyntactic annotation. |
| Outcome: | The corpus provides three layers: transliteration, transcription and morphosyntactic annotation. |
STT4SG-350: A Speech Corpus for All Swiss German Dialect Regions (2023.acl-short)
Copied to clipboard
Michel Plüss, Jan Deriu, Yanick Schraner, Claudio Paonessa, Julia Hartmann, Larissa Schmidt, Christian Scheller, Manuela Hürlimann, Tanja Samardžić, Manfred Vogel, Mark Cieliebak
| Challenge: | We present a corpus of Swiss German speech annotated with Standard German text at the sentence level. |
| Approach: | They present a corpus of Swiss German speech annotated with Standard German sentences . they use a web app to show the speakers standard German sentences and record them . |
| Outcome: | The corpus contains 343 hours of speech from all Swiss German dialect regions . it is the largest public speech corpus for Swiss German to date . |
SB-CH: A Swiss German Corpus with Sentiment Annotations (L18-1)
Copied to clipboard
| Challenge: | Using sentiment annotations, we find no corpus for written Swiss German, which is considered low-resourced due to its non-official status and phonetic differences. |
| Approach: | They propose to annotate a Swiss German corpus with sentiment annotations for sentiment analysis using Facebook comments and online chats. |
| Outcome: | The proposed corpus consists of more than 200,000 phrases and 1843 phrases with labels positive, negative, or neutral. |
A Corpus of German Abstract Meaning Representation (DeAMR) (2024.lrec-main)
Copied to clipboard
| Challenge: | Abstract Meaning Representations (AMRs) are semantic graphs that abstract away from surface syntax and capture the meaning of who does what to whom in a sentence. |
| Approach: | They propose to use German Abstract Meaning Representation (Deutsche AMR) to represent the structure and semantics of German. |
| Outcome: | The proposed framework is based on an annotated corpus of 400 DeAMR in German and is validated through inter-annotator agreement. |
The WAW Corpus: The First Corpus of Interpreted Speeches and their Translations for English and Arabic (L18-1)
Copied to clipboard
| Challenge: | Using the corpus, we study the characteristics of interpreters' work and train machine translation systems. |
| Approach: | They propose to build an interpreting corpus for Arabic and an Arabic corpus to study interpreters' work. |
| Outcome: | The proposed corpus can be used for teaching interpreters and to train machine translation systems. |
The Boarnsterhim Corpus: A Bilingual Frisian-Dutch Panel and Trend Study (L18-1)
Copied to clipboard
| Challenge: | a corpus of 250 hours of speech in both west frisian and Dutch is being developed . the corpus is a sociolinguistic corpus based on the recordings of four generations of bilingual speakers . |
| Approach: | This paper describes the Boarnsterhim Corpus project which started in 2016 . it aims to make available 250 hours of speech in both west frisian and Dutch by same speakers . |
| Outcome: | The Boarnsterhim Corpus is a sociolinguistic corpus of west frisian and Dutch speakers . it spans four generations and includes panel and trend data . |