Corpus REDEWIEDERGABE (2020.lrec-1)

Copied to clipboard

Challenge: The corpus REDEWIEDERGABE contains detailed annotations for speech, thought and writing representation (ST&WR) with approximately 490,000 tokens, it is the largest resource of its kind.
Approach: This paper presents corpus REDEWIEDERGABE, a German-language historical corpus with detailed annotations for speech, thought and writing representation (ST&WR).
Outcome: The corpus REDEWIEDERGABE contains 490,000 tokens and is the largest resource of its kind.

Similar Papers

A Large Annotated Reference Corpus of New High German Poetry (2024.lrec-main)

Copied to clipboard

Challenge: a corpus of public domain German poetry covering the time period 1600 to the 1920s contains 65k unique poems and over 1.6M lines, each tokenized, syllabified, pos-tagged, and meter-tagged.
Approach: They present a large annotated corpus of public domain German poetry covering the time period 1600 to the 1920s with 65k poems.
Outcome: The corpus contains 65k unique poems and over 1.6M lines, each tokenized, syllabified, pos-tagged, and meter-tagged.
Introducing a Parsed Corpus of Historical High German (2024.lrec-main)

Copied to clipboard

Challenge: outlines the development of the Indiana Parsed Corpus of (Historical) High German . outlines selection of texts, decisions on part-of-speech tags and other labels .
Approach: They propose to build a parsed German corpus that spans Germanic from 1050 to 1950 . they propose to use Penn-style treebanks to capture syntactic relationships between words .
Outcome: The proposed corpus spans Germanic languages from 1050 to 1950 and illustrative annotation issues unique to the language.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
An Annotated Social Media Corpus for German (2020.lrec-1)

Copied to clipboard

Challenge: Hate Speech (HS) against ethnic, religious and national minorities is a growing concern in online discourse.
Approach: They present the German Twitter section of a large (2 billion word) bilingual Social Media corpus for Hate Speech research.
Outcome: The proposed parser achieved F-scores of 97% for morphology and 92% for syntax on a cross-section of tweets.
Manually Annotated Corpus of Polish Texts Published between 1830 and 1918 (L18-1)

Copied to clipboard

Challenge: a paper presents a manually annotated corpus of 625,000 tokens of Polish texts . the corpus provides three layers: transliteration, transcription and morphosyntactic annotation.
Approach: The paper presents a manually annotated large historical corpus of Polish . the corpus provides three layers: transliteration, transcription and morphosyntactic annotation.
Outcome: The corpus provides three layers: transliteration, transcription and morphosyntactic annotation.
STT4SG-350: A Speech Corpus for All Swiss German Dialect Regions (2023.acl-short)

Copied to clipboard

Challenge: We present a corpus of Swiss German speech annotated with Standard German text at the sentence level.
Approach: They present a corpus of Swiss German speech annotated with Standard German sentences . they use a web app to show the speakers standard German sentences and record them .
Outcome: The corpus contains 343 hours of speech from all Swiss German dialect regions . it is the largest public speech corpus for Swiss German to date .
SB-CH: A Swiss German Corpus with Sentiment Annotations (L18-1)

Copied to clipboard

Challenge: Using sentiment annotations, we find no corpus for written Swiss German, which is considered low-resourced due to its non-official status and phonetic differences.
Approach: They propose to annotate a Swiss German corpus with sentiment annotations for sentiment analysis using Facebook comments and online chats.
Outcome: The proposed corpus consists of more than 200,000 phrases and 1843 phrases with labels positive, negative, or neutral.
A Corpus of German Abstract Meaning Representation (DeAMR) (2024.lrec-main)

Copied to clipboard

Challenge: Abstract Meaning Representations (AMRs) are semantic graphs that abstract away from surface syntax and capture the meaning of who does what to whom in a sentence.
Approach: They propose to use German Abstract Meaning Representation (Deutsche AMR) to represent the structure and semantics of German.
Outcome: The proposed framework is based on an annotated corpus of 400 DeAMR in German and is validated through inter-annotator agreement.
The WAW Corpus: The First Corpus of Interpreted Speeches and their Translations for English and Arabic (L18-1)

Copied to clipboard

Challenge: Using the corpus, we study the characteristics of interpreters' work and train machine translation systems.
Approach: They propose to build an interpreting corpus for Arabic and an Arabic corpus to study interpreters' work.
Outcome: The proposed corpus can be used for teaching interpreters and to train machine translation systems.
The Boarnsterhim Corpus: A Bilingual Frisian-Dutch Panel and Trend Study (L18-1)

Copied to clipboard

Challenge: a corpus of 250 hours of speech in both west frisian and Dutch is being developed . the corpus is a sociolinguistic corpus based on the recordings of four generations of bilingual speakers .
Approach: This paper describes the Boarnsterhim Corpus project which started in 2016 . it aims to make available 250 hours of speech in both west frisian and Dutch by same speakers .
Outcome: The Boarnsterhim Corpus is a sociolinguistic corpus of west frisian and Dutch speakers . it spans four generations and includes panel and trend data .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations