Building Literary Corpora for Computational Literary Analysis - A Prototype to Bridge the Gap between CL and DH (L18-1)
Copied to clipboard
| Challenge: | Literature analysis using corpus-based literary analysis is slow, says aaron s. e. . literary studies researchers should focus on the research practices of literary studies, he says . |
| Approach: | et al. show litText can extract text from a 20 million word corpus using SPARQL queries. |
| Outcome: | The proposed method uses a 20 million word corpus from English, German, Spanish, French and Italian texts and an example query to identify texts where animals behave like humans as it is the case in fables. |
Similar Papers
Using Bibliodata LODification to Create Metadata-Enriched Literary Corpora in Line with FAIR Principles (2024.lrec-main)
Copied to clipboard
| Challenge: | Literature corpus building is relatively nascent, and standardized procedures for curating literary corpora are not yet developed. |
| Approach: | They propose a workflow for the creation and reuse of literary corpora using a metadata-enriched Polish Novel Corpus from the 19th and 20th centuries. |
| Outcome: | The proposed workflow includes a multi-stage metadata enrichment and verification process and efficient data collection and data sharing according to the FAIR principles and 5- and 7-star data standards. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
The BDCamões Collection of Portuguese Literary Documents: a Research Resource for Digital Humanities and Language Technology (2020.lrec-1)
Copied to clipboard
| Challenge: | a new corpus of literary documents in Portuguese is presented . it includes close to 4 million words from over 200 complete documents . the corpus is suitable for research in language technology and digital humanities . |
| Approach: | They present the BDCames Collection of Portuguese Literary Documents, a new corpus of literary texts written in Portuguese. |
| Outcome: | The BDCames Collection of Portuguese Literary Documents is a new corpus of literary documents written in Portuguese . it includes close to 4 million words from over 200 complete documents from 83 authors in 14 genres . the corpus is suitable for research in language technology and language science and digital humanities . |
On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey (2024.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) provide a data-centric solution to alleviate limitations of real-world data with synthetic data generation. |
| Approach: | They propose a generic workflow for LLM-driven synthetic data generation. |
| Outcome: | The proposed workflows highlight gaps in existing research and outline avenues for future studies. |
Casting Light on Invisible Cities: Computationally Engaging with Literary Criticism (N19-1)
Copied to clipboard
| Challenge: | Literary critics often attempt to uncover meaning in a single work of literature through careful reading and analysis. |
| Approach: | They propose to use a literary theory to analyze Italo Calvino's novel Invisible Cities to leverage contextualized representations to embed each city's description and use unsupervised methods to cluster embeddings. |
| Outcome: | The proposed method can be applied to Italo Calvino’s novel Invisible Cities . authors compare results to similarity judgments generated by human readers . |
Know thy Corpus! Robust Methods for Digital Curation of Web corpora (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing methods for estimating the lexicon of Web corpora have not been used to train pre-trained models. |
| Approach: | They propose a framework for digital curation of Web corpora to provide robust estimation of their parameters. |
| Outcome: | The proposed framework provides robust estimation of Web corpora's composition and lexicon . the proposed framework is similar to the BNC and ELMO models, but lacks curated categories . |
Literary Evidence Retrieval via Long-Context Language Models (2025.acl-short)
Copied to clipboard
| Challenge: | a recent study shows that long-context language models can exceed human expert performance in literary analysis . despite their speed and apparent accuracy, even the strongest models struggle with nuanced literary signals and overgeneration. |
| Approach: | They propose a task where a model is given an entire text of a book and a literary criticism with a missing quotation from that work and asked to generate the missing quote. |
| Outcome: | The proposed model outperforms open-weight models in literary evidence retrieval tasks. |
Text Mining for History: first steps on building a large dataset (L18-1)
Copied to clipboard
| Challenge: | a new corpus on the history domain is being created to mine text in the domain . primary motivation for the project is the need to query the material in a non-linear way . |
| Approach: | They propose to use a Brazilian historical-biographical dictionary as a resource for text mining. |
| Outcome: | The proposed corpus is a reference work on the Brazilian history domain . it contains almost 12 millions tokens in about three hundred thousand sentences . the authors argue that the proposed corpu is linguistically motivated . |
From FreEM to D’AlemBERT: a Large Corpus and a Language Model for Early Modern French (2022.lrec-1)
Copied to clipboard
Simon Gabay, Pedro Ortiz Suarez, Alexandre Bartz, Alix Chagué, Rachel Bawden, Philippe Gambette, Benoît Sagot
| Challenge: | Anguage models for historical states of language are becoming more complex to process and more scarce in the corpora available. |
| Approach: | They propose to use a contextualised language model to analyse historical states of language in French. |
| Outcome: | The proposed model is based on a corpus of historical texts and is evaluated with an NLP task. |
The Royal Society Corpus 6.0: Providing 300+ Years of Scientific Writing for Humanistic Study (2020.lrec-1)
Copied to clipboard
| Challenge: | a new version of the Royal Society Corpus covers 300+ years of scientific writing . the corpus is freely available under a Creative Commons license, excluding copy-righted parts . |
| Approach: | They present a new version of the Royal Society Corpus, a diachronic corpus of scientific English covering 300+ years of scientific writing. |
| Outcome: | The extended version of the Royal Society Corpus covers 300+ years of scientific writing . the corpus is freely available under a Creative Commons license, excluding copy-righted parts . |