SB-CH: A Swiss German Corpus with Sentiment Annotations (L18-1)

Copied to clipboard

Challenge: Using sentiment annotations, we find no corpus for written Swiss German, which is considered low-resourced due to its non-official status and phonetic differences.
Approach: They propose to annotate a Swiss German corpus with sentiment annotations for sentiment analysis using Facebook comments and online chats.
Outcome: The proposed corpus consists of more than 200,000 phrases and 1843 phrases with labels positive, negative, or neutral.

Similar Papers

An Annotated Social Media Corpus for German (2020.lrec-1)

Copied to clipboard

Challenge: Hate Speech (HS) against ethnic, religious and national minorities is a growing concern in online discourse.
Approach: They present the German Twitter section of a large (2 billion word) bilingual Social Media corpus for Hate Speech research.
Outcome: The proposed parser achieved F-scores of 97% for morphology and 92% for syntax on a cross-section of tweets.
EmpiriST Corpus 2.0: Adding Manual Normalization, Lemmatization and Semantic Tagging to a German Web and CMC Corpus (2020.lrec-1)

Copied to clipboard

Challenge: EmpiriST corpus is a manually tokenized and part-of-speech tagged corpus of German web and CMC data.
Approach: They extend the corpus with manually tokenized annotation layers for word form normalization, lemmatization and lexical semantics.
Outcome: The EmpiriST corpus is a manually tokenized and part-of-speech tagged corpus of German web and CMC data.
STT4SG-350: A Speech Corpus for All Swiss German Dialect Regions (2023.acl-short)

Copied to clipboard

Challenge: We present a corpus of Swiss German speech annotated with Standard German text at the sentence level.
Approach: They present a corpus of Swiss German speech annotated with Standard German sentences . they use a web app to show the speakers standard German sentences and record them .
Outcome: The corpus contains 343 hours of speech from all Swiss German dialect regions . it is the largest public speech corpus for Swiss German to date .
Corpus REDEWIEDERGABE (2020.lrec-1)

Copied to clipboard

Challenge: The corpus REDEWIEDERGABE contains detailed annotations for speech, thought and writing representation (ST&WR) with approximately 490,000 tokens, it is the largest resource of its kind.
Approach: This paper presents corpus REDEWIEDERGABE, a German-language historical corpus with detailed annotations for speech, thought and writing representation (ST&WR).
Outcome: The corpus REDEWIEDERGABE contains 490,000 tokens and is the largest resource of its kind.
A Large Annotated Reference Corpus of New High German Poetry (2024.lrec-main)

Copied to clipboard

Challenge: a corpus of public domain German poetry covering the time period 1600 to the 1920s contains 65k unique poems and over 1.6M lines, each tokenized, syllabified, pos-tagged, and meter-tagged.
Approach: They present a large annotated corpus of public domain German poetry covering the time period 1600 to the 1920s with 65k poems.
Outcome: The corpus contains 65k unique poems and over 1.6M lines, each tokenized, syllabified, pos-tagged, and meter-tagged.
A Framenet and Frame Annotator for German Social Media (2022.lrec-1)

Copied to clipboard

Challenge: In corpus linguistics, semantic annotation is a valuable addition to ordinary, morphosyntactic tagging, lemmatization and dependency relations.
Approach: They propose a parsing- and annotation-oriented framenet for German with almost 15,000 frames . they propose valency, syntactic function and semantic noun class as input conditions for frame disambiguation .
Outcome: The proposed resource is based on a Danish/German study on hate speech . it achieves an overall F-score for frame senses of 93.6% on twitter .
A Corpus for Argumentative Writing Support in German (2020.coling-main)

Copied to clipboard

Challenge: In today's world most information is readily available. Consequently, the sole reproduction of information is losing attention.
Approach: They propose an annotation approach to capture claims and premises of arguments and their relations in student-written peer reviews on business models in german language.
Outcome: The proposed annotation scheme guides annotators to moderate agreement with the proposed scheme on 50 persuasive student-written peer reviews on business models.
The SSIX Corpora: Three Gold Standard Corpora for Sentiment Analysis in English, Spanish and German Financial Microblogs (L18-1)

Copied to clipboard

Challenge: SSIX corpora provide annotated data for supervised learning methods . polarity annotation is performed on two financial microblog platforms .
Approach: They propose three SSIX corpora for sentiment analysis which provide annotated data for supervised learning methods.
Outcome: The proposed corpora are in English, German and Spanish.
SDS-200: A Swiss German Speech to Standard German Text Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Using a web recording tool, participants were asked to translate their Swiss German text to their own dialect before recording it.
Approach: They present a corpus of Swiss German dialectal speech with Standard German text translations . the dataset allows for training speech translation, dialect recognition, and speech synthesis systems .
Outcome: The dataset allows for training speech translation, dialect recognition, and speech synthesis systems.
Annotated Corpus for Sentiment Analysis in Odia Language (2020.lrec-1)

Copied to clipboard

Challenge: Existing sentiment analysis models are not available for Odia 1 as it is a resource-poor language.
Approach: They create an annotated Odia corpus and test its usability by training and testing on the corpus using various classifiers.
Outcome: The created corpus contains 2045 Odia sentences from news domain annotated with sentiment labels using a well-defined annotation scheme.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations