KRAUTS: A German Temporally Annotated News Corpus (L18-1)

Copied to clipboard

Challenge: Temporal tagging is an important task towards improved natural language understanding.
Approach: They present a new German temporally annotated corpus with 192 documents with 1,140 annotations . they propose to make temporal tagging a viable research area .
Outcome: The proposed corpus contains 192 documents with 1,140 annotated temporal expressions.

Similar Papers

Uncovering Temporal Framing in the News (2026.acl-long)

Copied to clipboard

Challenge: Temporal language is used to structure meaning rather than report chronology in news discourse . a recent study focused on temporal expression extraction and temporal reasoning .
Approach: They propose a taxonomy of eight temporal frames grounded in prior work on time and framing . they analyze frame prevalence, co-occurrence patterns, and lexical cues from a news corpus .
Outcome: The proposed taxonomy outperforms zero-shot models at the sentence level . it shows that temporal framing is learnable at the sentences level compared to other methods .
An Annotated Social Media Corpus for German (2020.lrec-1)

Copied to clipboard

Challenge: Hate Speech (HS) against ethnic, religious and national minorities is a growing concern in online discourse.
Approach: They present the German Twitter section of a large (2 billion word) bilingual Social Media corpus for Hate Speech research.
Outcome: The proposed parser achieved F-scores of 97% for morphology and 92% for syntax on a cross-section of tweets.
NarrativeTime: Dense Temporal Annotation on a Timeline (2024.lrec-main)

Copied to clipboard

Challenge: e.g. TimeBank contains 1-5% of all possible tlinks, and this information is underspecified in the text.
Approach: They propose a timeline-based framework that achieves full coverage of all possible TLINKs.
Outcome: The proposed framework achieves full coverage of all possible TLINKs in a text.
Comprehensive Annotation of Various Types of Temporal Information on the Time Axis (L18-1)

Copied to clipboard

Challenge: Existing studies linking event and time information have been conducted to train and evaluate models.
Approach: They propose an annotation scheme that anchors expressions in text to the time axis comprehensively.
Outcome: The proposed scheme can be utilized for integrated information analysis of events, entities and time.
TIMELINE: Exhaustive Annotation of Temporal Relations Supporting the Automatic Ordering of Events in News Articles (2023.emnlp-main)

Copied to clipboard

Challenge: Existing temporal relation extraction models have low inter-annotator agreement due to lack of specificity of annotation guidelines . authors propose a method for annotating all temporal relations, including long-distance ones, which automates the process .
Approach: They propose a new annotation scheme that defines criteria for temporal relations to be annotated . scheme includes events even if they are not expressed as verbs, they argue .
Outcome: The proposed method reduces time and manual effort on the part of annotators.
Annotation and Automatic Classification of Aspectual Categories (P19-1)

Copied to clipboard

Challenge: Annotated resource for aspectual classification of German verb tokens in context.
Approach: They present a resource for aspectual classification of German verb tokens in their clausal context.
Outcome: The proposed resource is compared with previous work on German verb tokens using aspectual features compatible with the plurality of aspectual classifications.
A Large Annotated Reference Corpus of New High German Poetry (2024.lrec-main)

Copied to clipboard

Challenge: a corpus of public domain German poetry covering the time period 1600 to the 1920s contains 65k unique poems and over 1.6M lines, each tokenized, syllabified, pos-tagged, and meter-tagged.
Approach: They present a large annotated corpus of public domain German poetry covering the time period 1600 to the 1920s with 65k poems.
Outcome: The corpus contains 65k unique poems and over 1.6M lines, each tokenized, syllabified, pos-tagged, and meter-tagged.
Dataset of Quotation Attribution in German News Articles (2024.lrec-main)

Copied to clipboard

Challenge: Lack of annotated data for quotation attribution in news articles severely limits the quality and usability of possible systems.
Approach: They propose a dataset for quotation attribution in German news articles using WIKINEWS and manually annotated quotes from 1000 articles.
Outcome: The proposed dataset provides curated, high-quality annotations across 1000 documents (250,000 tokens) in a fine-grained annotation schema enabling various downstream uses for the dataset.
Introducing a Parsed Corpus of Historical High German (2024.lrec-main)

Copied to clipboard

Challenge: outlines the development of the Indiana Parsed Corpus of (Historical) High German . outlines selection of texts, decisions on part-of-speech tags and other labels .
Approach: They propose to build a parsed German corpus that spans Germanic from 1050 to 1950 . they propose to use Penn-style treebanks to capture syntactic relationships between words .
Outcome: The proposed corpus spans Germanic languages from 1050 to 1950 and illustrative annotation issues unique to the language.
A Dataset of German Legal Documents for Named Entity Recognition (2020.lrec-1)

Copied to clipboard

Challenge: a dataset developed for Named Entity Recognition in German federal court decisions is available under a CC-BY 4.0 license.
Approach: They describe a dataset developed for Named Entity Recognition in German federal court decisions.
Outcome: The proposed dataset was developed for training an NER service for German legal documents in the EU project Lynx.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations