Challenge: Dramatic texts are highly structured literary text types with linguistic and literary properties.
Approach: They present an annotated corpus of German dramatic texts and preliminary experiments on automatic coreference resolution.
Outcome: The proposed system achieves a 28.8 CoNLL score in dramatic texts compared to other dialogical text types such as interviews . the proposed system is expected to be extended to include the (partial) information given in the dramatis personae .

Similar Papers

Annotation and Evaluation of Coreference Resolution in Screenplays (2021.findings-acl)

Copied to clipboard

Challenge: Screenplays refer to characters using different names, pronouns, and nominal expressions.
Approach: They develop an automatic screenplay parser to extract structural information and design coreference rules based upon the structure of screenplays.
Outcome: The proposed model outperforms a benchmark model on the screenplay coreference resolution task.
ezCoref: Towards Unifying Annotation Guidelines for Coreference Resolution (2023.findings-eacl)

Copied to clipboard

Challenge: Existing datasets vary in definition of coreferences and are curated for linguistic experts.
Approach: They propose to use ezCoref to create a crowdsourcing-friendly coreference annotation methodology that teaches annotators only cases that are treated similarly across existing datasets.
Outcome: The proposed method reannotates 240 passages from seven existing english coreference datasets while teaching annotators only cases that are treated similarly across them.
Variation in Coreference Strategies across Genres and Production Media (2020.coling-main)

Copied to clipboard

Challenge: a lack of work on automatic coreference resolution on spoken and written language has led to inconclusive results.
Approach: They propose to use Ontonotes, Switchboard and Twitter to investigate coreference . they find fairly clear patterns of "behavior" for the different genres/medias .
Outcome: The results show that the choice of genre and the medium (spoken versus spoken) relates to the spokenwritten spectrum for coreference strategies.
GGPONC 2.0 - The German Clinical Guideline Corpus for Oncology: Curation Workflow, Annotation Policy, Baseline NER Taggers (2022.lrec-1)

Copied to clipboard

Challenge: despite advances in language resources, there is still a shortage of annotated corpora covering (German) medical language.
Approach: They propose to build on clinical guidelines with an annotation scheme based on SNOMED CT . they also train named entity recognition models on the new data set .
Outcome: The new corpus can be built upon clinical guidelines with reasonable coverage of medical terminology.
ParCorFull: a Parallel Corpus Annotated with Full Coreference (L18-1)

Copied to clipboard

Challenge: Recent research in multilingual coreference and automatic pronoun translation has led to important insights into the problem and some promising results.
Approach: They propose a corpus annotated with full coreference chains that addresses a problem that machine translation and other multilingual natural language processing (NLP) technologies face: translation of coreference across languages.
Outcome: The proposed corpus contains parallel texts for the language pair English-German, two major European languages.
A Framenet and Frame Annotator for German Social Media (2022.lrec-1)

Copied to clipboard

Challenge: In corpus linguistics, semantic annotation is a valuable addition to ordinary, morphosyntactic tagging, lemmatization and dependency relations.
Approach: They propose a parsing- and annotation-oriented framenet for German with almost 15,000 frames . they propose valency, syntactic function and semantic noun class as input conditions for frame disambiguation .
Outcome: The proposed resource is based on a Danish/German study on hate speech . it achieves an overall F-score for frame senses of 93.6% on twitter .
Persian Discourse Treebank and coreference corpus (L18-1)

Copied to clipboard

Challenge: Currently, we are adding a new document-level discourse annotation to our new corpus.
Approach: They propose to build a Persian discourse treebank and a comprehensive Persian coreference corpus based on discourse analysis and coreference resolution.
Outcome: The proposed corpus includes 30000 individual sentences with morphological, syntactic and semantic labels and nearly half a million tokens.
BOOKCOREF: Coreference Resolution at Book Scale (2025.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for coreference resolution systems are limited in length and do not adequately assess system capabilities at the book scale.
Approach: They propose a novel pipeline that produces high-quality coreference resolution annotations on full narrative texts and a book-scale benchmark, BOOKCOREF.
Outcome: The proposed pipeline produces high-quality coreference resolution annotations on full texts with an average document length of more than 200,000 tokens.
An Annotated Social Media Corpus for German (2020.lrec-1)

Copied to clipboard

Challenge: Hate Speech (HS) against ethnic, religious and national minorities is a growing concern in online discourse.
Approach: They present the German Twitter section of a large (2 billion word) bilingual Social Media corpus for Hate Speech research.
Outcome: The proposed parser achieved F-scores of 97% for morphology and 92% for syntax on a cross-section of tweets.
A Corpus for Automatic Readability Assessment and Text Simplification of German (2020.lrec-1)

Copied to clipboard

Challenge: Using monolingual-only data, we can automate readability assessment and text simplification of simplified language.
Approach: They present a corpus for automatic readability assessment and automatic text simplification for German using parallel and monolingual data.
Outcome: The proposed corpus is compiled from web sources and contains information on text structure, typography, font style, and images.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations