| Challenge: | a new corpus on the history domain is being created to mine text in the domain . primary motivation for the project is the need to query the material in a non-linear way . |
| Approach: | They propose to use a Brazilian historical-biographical dictionary as a resource for text mining. |
| Outcome: | The proposed corpus is a reference work on the Brazilian history domain . it contains almost 12 millions tokens in about three hundred thousand sentences . the authors argue that the proposed corpu is linguistically motivated . |
Similar Papers
From FreEM to D’AlemBERT: a Large Corpus and a Language Model for Early Modern French (2022.lrec-1)
Copied to clipboard
Simon Gabay, Pedro Ortiz Suarez, Alexandre Bartz, Alix Chagué, Rachel Bawden, Philippe Gambette, Benoît Sagot
| Challenge: | Anguage models for historical states of language are becoming more complex to process and more scarce in the corpora available. |
| Approach: | They propose to use a contextualised language model to analyse historical states of language in French. |
| Outcome: | The proposed model is based on a corpus of historical texts and is evaluated with an NLP task. |
Scalable Construction and Reasoning of Massive Knowledge Bases (N18-6)
Copied to clipboard
| Challenge: | Existing knowledge mining systems assume abundant human annotations for training high quality machine learning models, which is impractical when trying to deploy IE systems to a broad range of domains, settings and languages. |
| Approach: | They introduce how to extract structured facts from text corpora to construct knowledge bases. |
| Outcome: | The proposed methods are weakly-supervised and domain-independent for knowledge base construction across various domains. |
Synthetic Data in the Era of Large Language Models (2025.acl-tutorials)
Copied to clipboard
| Challenge: | 'synthetic data' is a data generated with the assistance of large language models to make dataset construction faster and cheaper. |
| Approach: | This tutorial seeks to build a shared understanding of recent progress in synthetic data generation from NLP and related fields by grouping and describing major methods, applications, and open problems. |
| Outcome: | This tutorial will describe methods, applications, and open problems that have been developed and are being used to improve the quality and efficiency of synthetic data generation. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
To Boldly Query What No One Has Annotated Before? The Frontiers of Corpus Querying (2020.acl-main)
Copied to clipboard
| Challenge: | a systematic review of corpora and query tools focuses on the query side . annotated corporata are the backbone of many fields in linguistics . |
| Approach: | They propose a chronology of the major interplay between corpus progression and query tool evolution . they focus on the query side and hints at exciting directions for future development . |
| Outcome: | This paper provides a broad overview of the history of corpora and query tools . it focuses on the query side and hints at exciting directions for future development . |
Graph Matching and Graph Rewriting: GREW tools for corpus exploration, maintenance and conversion (2021.eacl-demos)
Copied to clipboard
| Challenge: | Graph Rewriting is a mathematical formalism that can be used to describe rule-based transformations on linguistic structures. |
| Approach: | They propose to use graph rewriting to describe rule-based transformations on linguistic structures. |
| Outcome: | The proposed tools can be used to compute rule-based transformations on linguistic structures. |
CLAUSE-ATLAS: A Corpus of Narrative Information to Scale up Computational Literary Analysis (2024.lrec-main)
Copied to clipboard
| Challenge: | XIX and XX century English novels annotated automatically contain 41,715 labeled clauses . a new approach to analyze novels based on clauses captures structural patterns within books, as well as qualitative differences between them. |
| Approach: | They propose to use a corpus of XIX and XX century English novels annotated automatically to study stories as sequences of eventive, subjective and contextual information. |
| Outcome: | The proposed method captures structural patterns within books, as well as qualitative differences between them. |
A Large-Scale Comparison of Historical Text Normalization Systems (N19-1)
Copied to clipboard
| Challenge: | a large study of historical text normalization is done on eight languages . there is no consensus on the state-of-the-art approach to normalization . |
| Approach: | They present a large study of historical text normalization done on eight languages . they evaluate four different systems based on supervised learning on datasets from eight different languages based in the literature . |
| Outcome: | The proposed methods are based on supervised learning and are available online. |
Pretraining Language Models for Diachronic Linguistic Change Discovery (2026.findings-eacl)
Copied to clipboard
| Challenge: | Large language models are increasingly used as knowledge discovery tools . historical linguistics and literary studies often construct arguments on the basis of distinctions between phenomena like time-period or genre. |
| Approach: | They propose to use LLMs to train large language models over modest historical corpora without allowing contamination from anachronistic data. |
| Outcome: | The proposed model better respects historical divisions and is more computationally efficient compared to the standard approach of fine-tuning an existing LLM. |
A Lightweight Approach to a Giga-Corpus of Historical Periodicals: The Story of a Slovenian Historical Newspaper Collection (2024.lrec-main)
Copied to clipboard
| Challenge: | a curated corpus of Slovenian historical newspapers is a complex undertaking requiring multiple steps to prepare . a shoestring budget is required to produce a corpus that is billion-words in size . |
| Approach: | They propose a lightweight approach to producing high-quality corpora using OCR . they use noisy OCR-ed data from the National and University Library of Slovenia . |
| Outcome: | The proposed method produces a billion-word giga-corpus of Slovenian historical newspapers from the 18th, 19th and 20th centuries on a shoestring budget. |