The Project Dialogism Novel Corpus: A Dataset for Quotation Attribution in Literary Texts (2022.lrec-1)
Copied to clipboard
| Challenge: | Annotated quotations are used to model attribution and coreference of literary texts . authors present project Dialogism Novel Corpus, or PDNC, for 22 novels . |
| Approach: | They present an annotated dataset of quotations for English literary texts . they use natural language processing to model aspects of narrative, events, and characters . |
| Outcome: | The project Dialogism Novel Corpus contains annotations for 35,978 quotations across 22 novels . authors show that NLP can be used to model aspects of narrative, events, and characters . |
Similar Papers
Improving Automatic Quotation Attribution in Literary Novels (2023.acl-short)
Copied to clipboard
| Challenge: | Existing methods for quotation attribution in literary novels require varying levels of available information. |
| Approach: | They propose to train and evaluate models for character identification, coreference resolution, quotation identification and speaker attribution tasks using an annotated dataset. |
| Outcome: | The proposed model scores on speaker attribution task on the same scale as state-of-the-art models. |
RiQuA: A Corpus of Rich Quotation Annotation for English Literary Text (2020.lrec-1)
Copied to clipboard
| Challenge: | In literature, spoken interactions between characters are of central importance to the narrative. |
| Approach: | They propose to annotate quotations, including their interpersonal structure, for English literary text. |
| Outcome: | The proposed dataset provides a rich view of dialogue structures not available from other available corpora. |
An Annotated Dataset of Coreference in English Literature (2020.lrec-1)
Copied to clipboard
| Challenge: | Using OntoNotes, coreference resolution systems are typically evaluated on this data exclusively. |
| Approach: | They present a new dataset of coreference annotations for works of literature in English covering 29,103 mentions in 210,532 tokens from 100 works of fiction published between 1719 and 1922. |
| Outcome: | The proposed dataset covers 29,103 mentions in 210,532 tokens from 100 works of fiction published between 1719 and 1922. |
DirectQuote: A Dataset for Direct Quotation Extraction and Attribution in News Articles (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing methods to extract and attribute quotations from news data are difficult and require a lot of effort. |
| Approach: | They propose a corpus of 19,760 paragraphs and 10,279 direct quotations manually annotated from online news media. |
| Outcome: | The proposed corpus contains 19,760 paragraphs and 10,279 direct quotations manually annotated from online news media. |
Evaluating LLMs for Quotation Attribution in Literary Texts: A Case Study of LLaMa3 (2025.naacl-short)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown promising results in literary tasks . however, quotation attribution remains a challenging task and methods that generalize across writing styles are lacking analysis regarding book memorization and annotation contamination. |
| Approach: | They evaluate the ability of Llama-3 to attribute utterances of direct-speech to their speaker in novels by assessing the impact of book memorization and annotation contamination. |
| Outcome: | The proposed model outperforms existing models on a corpus of 28 novels and shows that book memorization and annotation contamination do not explain the performance gain. |
Improving Quotation Attribution with Fictional Character Embeddings (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Recent methods to attribute quotes to human logic lack character representations, which often leads to errors in more challenging examples of attribution: anaphoric and implicit quotes. |
| Approach: | They propose to augment a popular quotation attribution system, BookNLP, with character embeddings that encode global stylistic information of characters derived from an off-the-shelf stylometric model, Universal Authorship Representation (UAR). |
| Outcome: | The proposed system improves anaphoric and implicit quotes, reaching state-of-the-art. |
CLAUSE-ATLAS: A Corpus of Narrative Information to Scale up Computational Literary Analysis (2024.lrec-main)
Copied to clipboard
| Challenge: | XIX and XX century English novels annotated automatically contain 41,715 labeled clauses . a new approach to analyze novels based on clauses captures structural patterns within books, as well as qualitative differences between them. |
| Approach: | They propose to use a corpus of XIX and XX century English novels annotated automatically to study stories as sequences of eventive, subjective and contextual information. |
| Outcome: | The proposed method captures structural patterns within books, as well as qualitative differences between them. |
JESC: Japanese-English Subtitle Corpus (L18-1)
Copied to clipboard
| Challenge: | Existing data on Japanese-English subtitles are limited due to the high cost of manual construction. |
| Approach: | They describe the Japanese-English Subtitle Corpus by crawling and aligning subtitles found on the web. |
| Outcome: | The JESC dataset covers the underrepresented domain of conversational dialogue. |
A Short Survey on Sense-Annotated Corpora (2020.lrec-1)
Copied to clipboard
| Challenge: | Word Sense Disambiguation (WSD) is a key task in Natural Language Understanding. |
| Approach: | They propose to use sense-annotated corpora for supervised Word Sense Disambiguation. |
| Outcome: | The proposed methods have been compared with knowledge-based approaches and have shown to be more efficient when they are available. |
AMR Beyond the Sentence: the Multi-sentence AMR corpus (C18-1)
Copied to clipboard
| Challenge: | Abstract Meaning Representation (AMR) is limited to capturing the semantics of individual sentences. |
| Approach: | They propose a corpus that annotates coreference and similar phenomena on top of existing AMRs. |
| Outcome: | The proposed corpus is compared with existing corpora on sentence-level semantics . it shows that it can be used for information extraction and question answering . |