GerDraCor-Coref: A Coreference Corpus for Dramatic Texts in German (2020.lrec-1)
Copied to clipboard
| Challenge: | Dramatic texts are highly structured literary text types with linguistic and literary properties. |
| Approach: | They present an annotated corpus of German dramatic texts and preliminary experiments on automatic coreference resolution. |
| Outcome: | The proposed system achieves a 28.8 CoNLL score in dramatic texts compared to other dialogical text types such as interviews . the proposed system is expected to be extended to include the (partial) information given in the dramatis personae . |
Similar Papers
Annotation and Evaluation of Coreference Resolution in Screenplays (2021.findings-acl)
Copied to clipboard
| Challenge: | Screenplays refer to characters using different names, pronouns, and nominal expressions. |
| Approach: | They develop an automatic screenplay parser to extract structural information and design coreference rules based upon the structure of screenplays. |
| Outcome: | The proposed model outperforms a benchmark model on the screenplay coreference resolution task. |
ezCoref: Towards Unifying Annotation Guidelines for Coreference Resolution (2023.findings-eacl)
Copied to clipboard
Ankita Gupta, Marzena Karpinska, Wenlong Zhao, Kalpesh Krishna, Jack Merullo, Luke Yeh, Mohit Iyyer, Brendan O’Connor
| Challenge: | Existing datasets vary in definition of coreferences and are curated for linguistic experts. |
| Approach: | They propose to use ezCoref to create a crowdsourcing-friendly coreference annotation methodology that teaches annotators only cases that are treated similarly across existing datasets. |
| Outcome: | The proposed method reannotates 240 passages from seven existing english coreference datasets while teaching annotators only cases that are treated similarly across them. |
Variation in Coreference Strategies across Genres and Production Media (2020.coling-main)
Copied to clipboard
| Challenge: | a lack of work on automatic coreference resolution on spoken and written language has led to inconclusive results. |
| Approach: | They propose to use Ontonotes, Switchboard and Twitter to investigate coreference . they find fairly clear patterns of "behavior" for the different genres/medias . |
| Outcome: | The results show that the choice of genre and the medium (spoken versus spoken) relates to the spokenwritten spectrum for coreference strategies. |
GGPONC 2.0 - The German Clinical Guideline Corpus for Oncology: Curation Workflow, Annotation Policy, Baseline NER Taggers (2022.lrec-1)
Copied to clipboard
Florian Borchert, Christina Lohr, Luise Modersohn, Jonas Witt, Thomas Langer, Markus Follmann, Matthias Gietzelt, Bert Arnrich, Udo Hahn, Matthieu-P. Schapranow
| Challenge: | despite advances in language resources, there is still a shortage of annotated corpora covering (German) medical language. |
| Approach: | They propose to build on clinical guidelines with an annotation scheme based on SNOMED CT . they also train named entity recognition models on the new data set . |
| Outcome: | The new corpus can be built upon clinical guidelines with reasonable coverage of medical terminology. |
ParCorFull: a Parallel Corpus Annotated with Full Coreference (L18-1)
Copied to clipboard
| Challenge: | Recent research in multilingual coreference and automatic pronoun translation has led to important insights into the problem and some promising results. |
| Approach: | They propose a corpus annotated with full coreference chains that addresses a problem that machine translation and other multilingual natural language processing (NLP) technologies face: translation of coreference across languages. |
| Outcome: | The proposed corpus contains parallel texts for the language pair English-German, two major European languages. |
A Framenet and Frame Annotator for German Social Media (2022.lrec-1)
Copied to clipboard
| Challenge: | In corpus linguistics, semantic annotation is a valuable addition to ordinary, morphosyntactic tagging, lemmatization and dependency relations. |
| Approach: | They propose a parsing- and annotation-oriented framenet for German with almost 15,000 frames . they propose valency, syntactic function and semantic noun class as input conditions for frame disambiguation . |
| Outcome: | The proposed resource is based on a Danish/German study on hate speech . it achieves an overall F-score for frame senses of 93.6% on twitter . |
Persian Discourse Treebank and coreference corpus (L18-1)
Copied to clipboard
| Challenge: | Currently, we are adding a new document-level discourse annotation to our new corpus. |
| Approach: | They propose to build a Persian discourse treebank and a comprehensive Persian coreference corpus based on discourse analysis and coreference resolution. |
| Outcome: | The proposed corpus includes 30000 individual sentences with morphological, syntactic and semantic labels and nearly half a million tokens. |
BOOKCOREF: Coreference Resolution at Book Scale (2025.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks for coreference resolution systems are limited in length and do not adequately assess system capabilities at the book scale. |
| Approach: | They propose a novel pipeline that produces high-quality coreference resolution annotations on full narrative texts and a book-scale benchmark, BOOKCOREF. |
| Outcome: | The proposed pipeline produces high-quality coreference resolution annotations on full texts with an average document length of more than 200,000 tokens. |
An Annotated Social Media Corpus for German (2020.lrec-1)
Copied to clipboard
| Challenge: | Hate Speech (HS) against ethnic, religious and national minorities is a growing concern in online discourse. |
| Approach: | They present the German Twitter section of a large (2 billion word) bilingual Social Media corpus for Hate Speech research. |
| Outcome: | The proposed parser achieved F-scores of 97% for morphology and 92% for syntax on a cross-section of tweets. |
A Corpus for Automatic Readability Assessment and Text Simplification of German (2020.lrec-1)
Copied to clipboard
| Challenge: | Using monolingual-only data, we can automate readability assessment and text simplification of simplified language. |
| Approach: | They present a corpus for automatic readability assessment and automatic text simplification for German using parallel and monolingual data. |
| Outcome: | The proposed corpus is compiled from web sources and contains information on text structure, typography, font style, and images. |