Challenge: XIX and XX century English novels annotated automatically contain 41,715 labeled clauses . a new approach to analyze novels based on clauses captures structural patterns within books, as well as qualitative differences between them.
Approach: They propose to use a corpus of XIX and XX century English novels annotated automatically to study stories as sequences of eventive, subjective and contextual information.
Outcome: The proposed method captures structural patterns within books, as well as qualitative differences between them.

Similar Papers

The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
Text Mining for History: first steps on building a large dataset (L18-1)

Copied to clipboard

Challenge: a new corpus on the history domain is being created to mine text in the domain . primary motivation for the project is the need to query the material in a non-linear way .
Approach: They propose to use a Brazilian historical-biographical dictionary as a resource for text mining.
Outcome: The proposed corpus is a reference work on the Brazilian history domain . it contains almost 12 millions tokens in about three hundred thousand sentences . the authors argue that the proposed corpu is linguistically motivated .
To Boldly Query What No One Has Annotated Before? The Frontiers of Corpus Querying (2020.acl-main)

Copied to clipboard

Challenge: a systematic review of corpora and query tools focuses on the query side . annotated corporata are the backbone of many fields in linguistics .
Approach: They propose a chronology of the major interplay between corpus progression and query tool evolution . they focus on the query side and hints at exciting directions for future development .
Outcome: This paper provides a broad overview of the history of corpora and query tools . it focuses on the query side and hints at exciting directions for future development .
An Annotated Dataset of Coreference in English Literature (2020.lrec-1)

Copied to clipboard

Challenge: Using OntoNotes, coreference resolution systems are typically evaluated on this data exclusively.
Approach: They present a new dataset of coreference annotations for works of literature in English covering 29,103 mentions in 210,532 tokens from 100 works of fiction published between 1719 and 1922.
Outcome: The proposed dataset covers 29,103 mentions in 210,532 tokens from 100 works of fiction published between 1719 and 1922.
A Structured Clustering Approach for Inducing Media Narratives (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to modeling media narratives miss subtle narrative patterns through coarse-grained analysis or require domain-specific taxonomies that limit scalability.
Approach: They propose a framework for inducing rich narrative schemas by jointly modeling events and characters via structured clustering.
Outcome: The proposed framework produces explainable narrative schemas that align with established framing theory while scaling to large corpora without exhaustive manual annotation.
An annotated dataset of literary entities (N19-1)

Copied to clipboard

Challenge: Existing datasets built on news focus on non-named entities, but not literary texts.
Approach: They propose to annotate 210,532 tokens from 100 different English-language literary texts for ACE entity categories (person, location, geo-political entity, facility, organization, and vehicle).
Outcome: The proposed dataset includes 210,532 tokens drawn from 100 different English-language literary texts.
MAGPIE: A Large Corpus of Potentially Idiomatic Expressions (2020.lrec-1)

Copied to clipboard

Challenge: Existing corpora cover less than 5,000 instances of less than 100 different idiom types . large corpus allows for better evaluation of assumptions about idiomatic expressions .
Approach: They propose to build the largest-to-date corpus of idioms for English using crowdsourcing methods.
Outcome: The proposed corpus is larger than existing resources and contains rich metadata and is made publicly available.
NarrativeXL: a Large-scale Dataset for Long-Term Memory Models (2023.findings-emnlp)

Copied to clipboard

Challenge: 990,595 questions are needed to solve ultra-long-context reading comprehension problems.
Approach: They propose a large-scale reading comprehension dataset using 1,500 hand-curated fiction books and a set of reading comprehension questions based on these summaries.
Outcome: The proposed reading comprehension dataset is larger than the closest alternatives and has more questions than the existing models.
AMR Beyond the Sentence: the Multi-sentence AMR corpus (C18-1)

Copied to clipboard

Challenge: Abstract Meaning Representation (AMR) is limited to capturing the semantics of individual sentences.
Approach: They propose a corpus that annotates coreference and similar phenomena on top of existing AMRs.
Outcome: The proposed corpus is compared with existing corpora on sentence-level semantics . it shows that it can be used for information extraction and question answering .
Transactions of the Association for Computational Linguistics, Volume 8 (2020.tacl-1)

Copied to clipboard

Challenge: null
Approach: null
Outcome: null

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations