| Challenge: | specialized newspapers are not well curated in terms of digitization quality, data formatting, completeness, redundancy (de-duplication), supply of metadata, and hence, searchability. |
| Approach: | They propose a workflow that copes with a posteriori digitization problems, inconsistent OCRing and index building for searchability for a major German-language newspaper of the Romantic Age. |
| Outcome: | The proposed workflow copes with a posteriori digitization problems, inconsistent OCRing and index building for searchability. |
Similar Papers
A Lightweight Approach to a Giga-Corpus of Historical Periodicals: The Story of a Slovenian Historical Newspaper Collection (2024.lrec-main)
Copied to clipboard
| Challenge: | a curated corpus of Slovenian historical newspapers is a complex undertaking requiring multiple steps to prepare . a shoestring budget is required to produce a corpus that is billion-words in size . |
| Approach: | They propose a lightweight approach to producing high-quality corpora using OCR . they use noisy OCR-ed data from the National and University Library of Slovenia . |
| Outcome: | The proposed method produces a billion-word giga-corpus of Slovenian historical newspapers from the 18th, 19th and 20th centuries on a shoestring budget. |
taz2024full: Analysing German Newspapers for Gender Bias and Discrimination across Decades (2025.findings-acl)
Copied to clipboard
| Challenge: | a large corpus of German newspaper articles is available for free in other languages, such as English. |
| Approach: | They propose to use taz2024full to analyse gender representation across four decades of reporting. |
| Outcome: | The proposed corpus supports a wide range of applications from diachronic language analysis to critical media studies. |
Leveraging Digitized Newspapers to Collect Summarization Data in Low-Resource Languages (2026.findings-eacl)
Copied to clipboard
| Challenge: | Recent studies suggest that summarization in English may be solved, or even "dead" However, there are no accessible, high-quality summarizing datasets in under-represented languages. |
| Approach: | They propose a method for collecting naturally occurring summaries via front-page teasers, where editors summarize full length articles. |
| Outcome: | The proposed method is suited to varying linguistic resources and is available in seven languages. |
A database of German definitory contexts from selected web sources (L18-1)
Copied to clipboard
| Challenge: | a specialized web corpus and robust pattern-based extraction methods are used to detect definitory contexts. |
| Approach: | They propose to use a web corpus and a database to detect definitory contexts . they describe an experimental setting and front-end for pattern-based definition extraction . |
| Outcome: | The proposed method is based on a web corpus and a robust pattern-based extraction method. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
From FreEM to D’AlemBERT: a Large Corpus and a Language Model for Early Modern French (2022.lrec-1)
Copied to clipboard
Simon Gabay, Pedro Ortiz Suarez, Alexandre Bartz, Alix Chagué, Rachel Bawden, Philippe Gambette, Benoît Sagot
| Challenge: | Anguage models for historical states of language are becoming more complex to process and more scarce in the corpora available. |
| Approach: | They propose to use a contextualised language model to analyse historical states of language in French. |
| Outcome: | The proposed model is based on a corpus of historical texts and is evaluated with an NLP task. |
Fact from Fiction: Finding Serialized Novels in Newspapers (2025.acl-srw)
Copied to clipboard
| Challenge: | Among underrepresented but widely read forms are serialized fiction and feuilleton novels embedded in newspapers rather than published as standalone volumes. |
| Approach: | They propose to annotate 1,394 articles and evaluate classification pipelines using both selected linguistic features and embeddings to identify serialized fiction and feuilleton fiction. |
| Outcome: | The proposed methods achieve F1-scores of 0.91 in an annotated dataset of 1,394 articles and support the construction of alternative literary corpora and contribute to work on modeling the fiction–nonfiction boundary at scale. |
Infini-gram mini: Exact n-gram Search at the Internet Scale with FM-Index (2025.emnlp-main)
Copied to clipboard
| Challenge: | Modern language models are trained on text data downsampled from massive text corpora like Common Crawl. |
| Approach: | They propose an efficient and scalable system that can make petabyte-level text corpora searchable by using the FM-index data structure. |
| Outcome: | The proposed system indexes 83TB of Internet text in 99 days with a single 128-core CPU node (or 19 hours if using 137 such nodes). |
To Boldly Query What No One Has Annotated Before? The Frontiers of Corpus Querying (2020.acl-main)
Copied to clipboard
| Challenge: | a systematic review of corpora and query tools focuses on the query side . annotated corporata are the backbone of many fields in linguistics . |
| Approach: | They propose a chronology of the major interplay between corpus progression and query tool evolution . they focus on the query side and hints at exciting directions for future development . |
| Outcome: | This paper provides a broad overview of the history of corpora and query tools . it focuses on the query side and hints at exciting directions for future development . |
Know thy Corpus! Robust Methods for Digital Curation of Web corpora (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing methods for estimating the lexicon of Web corpora have not been used to train pre-trained models. |
| Approach: | They propose a framework for digital curation of Web corpora to provide robust estimation of their parameters. |
| Outcome: | The proposed framework provides robust estimation of Web corpora's composition and lexicon . the proposed framework is similar to the BNC and ELMO models, but lacks curated categories . |