| Challenge: | Using sentiment annotations, we find no corpus for written Swiss German, which is considered low-resourced due to its non-official status and phonetic differences. |
| Approach: | They propose to annotate a Swiss German corpus with sentiment annotations for sentiment analysis using Facebook comments and online chats. |
| Outcome: | The proposed corpus consists of more than 200,000 phrases and 1843 phrases with labels positive, negative, or neutral. |
Similar Papers
An Annotated Social Media Corpus for German (2020.lrec-1)
Copied to clipboard
| Challenge: | Hate Speech (HS) against ethnic, religious and national minorities is a growing concern in online discourse. |
| Approach: | They present the German Twitter section of a large (2 billion word) bilingual Social Media corpus for Hate Speech research. |
| Outcome: | The proposed parser achieved F-scores of 97% for morphology and 92% for syntax on a cross-section of tweets. |
EmpiriST Corpus 2.0: Adding Manual Normalization, Lemmatization and Semantic Tagging to a German Web and CMC Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | EmpiriST corpus is a manually tokenized and part-of-speech tagged corpus of German web and CMC data. |
| Approach: | They extend the corpus with manually tokenized annotation layers for word form normalization, lemmatization and lexical semantics. |
| Outcome: | The EmpiriST corpus is a manually tokenized and part-of-speech tagged corpus of German web and CMC data. |
STT4SG-350: A Speech Corpus for All Swiss German Dialect Regions (2023.acl-short)
Copied to clipboard
Michel Plüss, Jan Deriu, Yanick Schraner, Claudio Paonessa, Julia Hartmann, Larissa Schmidt, Christian Scheller, Manuela Hürlimann, Tanja Samardžić, Manfred Vogel, Mark Cieliebak
| Challenge: | We present a corpus of Swiss German speech annotated with Standard German text at the sentence level. |
| Approach: | They present a corpus of Swiss German speech annotated with Standard German sentences . they use a web app to show the speakers standard German sentences and record them . |
| Outcome: | The corpus contains 343 hours of speech from all Swiss German dialect regions . it is the largest public speech corpus for Swiss German to date . |
Corpus REDEWIEDERGABE (2020.lrec-1)
Copied to clipboard
| Challenge: | The corpus REDEWIEDERGABE contains detailed annotations for speech, thought and writing representation (ST&WR) with approximately 490,000 tokens, it is the largest resource of its kind. |
| Approach: | This paper presents corpus REDEWIEDERGABE, a German-language historical corpus with detailed annotations for speech, thought and writing representation (ST&WR). |
| Outcome: | The corpus REDEWIEDERGABE contains 490,000 tokens and is the largest resource of its kind. |
A Large Annotated Reference Corpus of New High German Poetry (2024.lrec-main)
Copied to clipboard
| Challenge: | a corpus of public domain German poetry covering the time period 1600 to the 1920s contains 65k unique poems and over 1.6M lines, each tokenized, syllabified, pos-tagged, and meter-tagged. |
| Approach: | They present a large annotated corpus of public domain German poetry covering the time period 1600 to the 1920s with 65k poems. |
| Outcome: | The corpus contains 65k unique poems and over 1.6M lines, each tokenized, syllabified, pos-tagged, and meter-tagged. |
A Framenet and Frame Annotator for German Social Media (2022.lrec-1)
Copied to clipboard
| Challenge: | In corpus linguistics, semantic annotation is a valuable addition to ordinary, morphosyntactic tagging, lemmatization and dependency relations. |
| Approach: | They propose a parsing- and annotation-oriented framenet for German with almost 15,000 frames . they propose valency, syntactic function and semantic noun class as input conditions for frame disambiguation . |
| Outcome: | The proposed resource is based on a Danish/German study on hate speech . it achieves an overall F-score for frame senses of 93.6% on twitter . |
A Corpus for Argumentative Writing Support in German (2020.coling-main)
Copied to clipboard
| Challenge: | In today's world most information is readily available. Consequently, the sole reproduction of information is losing attention. |
| Approach: | They propose an annotation approach to capture claims and premises of arguments and their relations in student-written peer reviews on business models in german language. |
| Outcome: | The proposed annotation scheme guides annotators to moderate agreement with the proposed scheme on 50 persuasive student-written peer reviews on business models. |
The SSIX Corpora: Three Gold Standard Corpora for Sentiment Analysis in English, Spanish and German Financial Microblogs (L18-1)
Copied to clipboard
| Challenge: | SSIX corpora provide annotated data for supervised learning methods . polarity annotation is performed on two financial microblog platforms . |
| Approach: | They propose three SSIX corpora for sentiment analysis which provide annotated data for supervised learning methods. |
| Outcome: | The proposed corpora are in English, German and Spanish. |
SDS-200: A Swiss German Speech to Standard German Text Corpus (2022.lrec-1)
Copied to clipboard
Michel Plüss, Manuela Hürlimann, Marc Cuny, Alla Stöckli, Nikolaos Kapotis, Julia Hartmann, Malgorzata Anna Ulasik, Christian Scheller, Yanick Schraner, Amit Jain, Jan Deriu, Mark Cieliebak, Manfred Vogel
| Challenge: | Using a web recording tool, participants were asked to translate their Swiss German text to their own dialect before recording it. |
| Approach: | They present a corpus of Swiss German dialectal speech with Standard German text translations . the dataset allows for training speech translation, dialect recognition, and speech synthesis systems . |
| Outcome: | The dataset allows for training speech translation, dialect recognition, and speech synthesis systems. |
Annotated Corpus for Sentiment Analysis in Odia Language (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing sentiment analysis models are not available for Odia 1 as it is a resource-poor language. |
| Approach: | They create an annotated Odia corpus and test its usability by training and testing on the corpus using various classifiers. |
| Outcome: | The created corpus contains 2045 Odia sentences from news domain annotated with sentiment labels using a well-defined annotation scheme. |