CLAUSE-ATLAS: A Corpus of Narrative Information to Scale up Computational Literary Analysis (2024.lrec-main)
Copied to clipboard
| Challenge: | XIX and XX century English novels annotated automatically contain 41,715 labeled clauses . a new approach to analyze novels based on clauses captures structural patterns within books, as well as qualitative differences between them. |
| Approach: | They propose to use a corpus of XIX and XX century English novels annotated automatically to study stories as sequences of eventive, subjective and contextual information. |
| Outcome: | The proposed method captures structural patterns within books, as well as qualitative differences between them. |
Similar Papers
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
Text Mining for History: first steps on building a large dataset (L18-1)
Copied to clipboard
| Challenge: | a new corpus on the history domain is being created to mine text in the domain . primary motivation for the project is the need to query the material in a non-linear way . |
| Approach: | They propose to use a Brazilian historical-biographical dictionary as a resource for text mining. |
| Outcome: | The proposed corpus is a reference work on the Brazilian history domain . it contains almost 12 millions tokens in about three hundred thousand sentences . the authors argue that the proposed corpu is linguistically motivated . |
To Boldly Query What No One Has Annotated Before? The Frontiers of Corpus Querying (2020.acl-main)
Copied to clipboard
| Challenge: | a systematic review of corpora and query tools focuses on the query side . annotated corporata are the backbone of many fields in linguistics . |
| Approach: | They propose a chronology of the major interplay between corpus progression and query tool evolution . they focus on the query side and hints at exciting directions for future development . |
| Outcome: | This paper provides a broad overview of the history of corpora and query tools . it focuses on the query side and hints at exciting directions for future development . |
An Annotated Dataset of Coreference in English Literature (2020.lrec-1)
Copied to clipboard
| Challenge: | Using OntoNotes, coreference resolution systems are typically evaluated on this data exclusively. |
| Approach: | They present a new dataset of coreference annotations for works of literature in English covering 29,103 mentions in 210,532 tokens from 100 works of fiction published between 1719 and 1922. |
| Outcome: | The proposed dataset covers 29,103 mentions in 210,532 tokens from 100 works of fiction published between 1719 and 1922. |
A Structured Clustering Approach for Inducing Media Narratives (2026.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to modeling media narratives miss subtle narrative patterns through coarse-grained analysis or require domain-specific taxonomies that limit scalability. |
| Approach: | They propose a framework for inducing rich narrative schemas by jointly modeling events and characters via structured clustering. |
| Outcome: | The proposed framework produces explainable narrative schemas that align with established framing theory while scaling to large corpora without exhaustive manual annotation. |
An annotated dataset of literary entities (N19-1)
Copied to clipboard
| Challenge: | Existing datasets built on news focus on non-named entities, but not literary texts. |
| Approach: | They propose to annotate 210,532 tokens from 100 different English-language literary texts for ACE entity categories (person, location, geo-political entity, facility, organization, and vehicle). |
| Outcome: | The proposed dataset includes 210,532 tokens drawn from 100 different English-language literary texts. |
MAGPIE: A Large Corpus of Potentially Idiomatic Expressions (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing corpora cover less than 5,000 instances of less than 100 different idiom types . large corpus allows for better evaluation of assumptions about idiomatic expressions . |
| Approach: | They propose to build the largest-to-date corpus of idioms for English using crowdsourcing methods. |
| Outcome: | The proposed corpus is larger than existing resources and contains rich metadata and is made publicly available. |
NarrativeXL: a Large-scale Dataset for Long-Term Memory Models (2023.findings-emnlp)
Copied to clipboard
| Challenge: | 990,595 questions are needed to solve ultra-long-context reading comprehension problems. |
| Approach: | They propose a large-scale reading comprehension dataset using 1,500 hand-curated fiction books and a set of reading comprehension questions based on these summaries. |
| Outcome: | The proposed reading comprehension dataset is larger than the closest alternatives and has more questions than the existing models. |
AMR Beyond the Sentence: the Multi-sentence AMR corpus (C18-1)
Copied to clipboard
| Challenge: | Abstract Meaning Representation (AMR) is limited to capturing the semantics of individual sentences. |
| Approach: | They propose a corpus that annotates coreference and similar phenomena on top of existing AMRs. |
| Outcome: | The proposed corpus is compared with existing corpora on sentence-level semantics . it shows that it can be used for information extraction and question answering . |
Transactions of the Association for Computational Linguistics, Volume 8 (2020.tacl-1)
Copied to clipboard
| Challenge: | null |
| Approach: | null |
| Outcome: | null |