Distant Reading in Digital Humanities: Case Study on the Serbian Part of the ELTeC Collection (2022.lrec-1)
Copied to clipboard
Ranka Stanković, Cvetana Krstev, Branislava Šandrih Todorović, Dusko Vitas, Mihailo Skoric, Milica Ikonić Nešić
| Challenge: | Distant reading is a new scale of description that does not displace previous scales of literary description. |
| Approach: | They present the Serbian part of the ELTeC multilingual corpus . they propose to test various methods and tools for distant reading . |
| Outcome: | The Serbian part of the ELTeC multilingual corpus is being built to test various methods and tools . Several use examples show that this sub-collection is usefull for both close and distant reading approaches. |
Similar Papers
A Lightweight Approach to a Giga-Corpus of Historical Periodicals: The Story of a Slovenian Historical Newspaper Collection (2024.lrec-main)
Copied to clipboard
| Challenge: | a curated corpus of Slovenian historical newspapers is a complex undertaking requiring multiple steps to prepare . a shoestring budget is required to produce a corpus that is billion-words in size . |
| Approach: | They propose a lightweight approach to producing high-quality corpora using OCR . they use noisy OCR-ed data from the National and University Library of Slovenia . |
| Outcome: | The proposed method produces a billion-word giga-corpus of Slovenian historical newspapers from the 18th, 19th and 20th centuries on a shoestring budget. |
The BDCamões Collection of Portuguese Literary Documents: a Research Resource for Digital Humanities and Language Technology (2020.lrec-1)
Copied to clipboard
| Challenge: | a new corpus of literary documents in Portuguese is presented . it includes close to 4 million words from over 200 complete documents . the corpus is suitable for research in language technology and digital humanities . |
| Approach: | They present the BDCames Collection of Portuguese Literary Documents, a new corpus of literary texts written in Portuguese. |
| Outcome: | The BDCames Collection of Portuguese Literary Documents is a new corpus of literary documents written in Portuguese . it includes close to 4 million words from over 200 complete documents from 83 authors in 14 genres . the corpus is suitable for research in language technology and language science and digital humanities . |
Using Bibliodata LODification to Create Metadata-Enriched Literary Corpora in Line with FAIR Principles (2024.lrec-main)
Copied to clipboard
| Challenge: | Literature corpus building is relatively nascent, and standardized procedures for curating literary corpora are not yet developed. |
| Approach: | They propose a workflow for the creation and reuse of literary corpora using a metadata-enriched Polish Novel Corpus from the 19th and 20th centuries. |
| Outcome: | The proposed workflow includes a multi-stage metadata enrichment and verification process and efficient data collection and data sharing according to the FAIR principles and 5- and 7-star data standards. |
SLäNDa: An Annotated Corpus of Narrative and Dialogue in Swedish Literary Fiction (2020.lrec-1)
Copied to clipboard
| Challenge: | a new corpus of literary fiction has been annotated for cited materials with a focus on dialogue. |
| Approach: | They propose to annotate a new corpus of Swedish literary fiction for cited materials with a focus on dialogue. |
| Outcome: | The proposed corpus can be used to train and analyze models for different types of analysis of literary narrative and speech. |
CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation (2024.lrec-main)
Copied to clipboard
| Challenge: | Using a similar crawling setup, the corpora are comparable across the entire South Slavic language space. |
| Approach: | They propose to collect 13 billion tokens of texts from 26 million documents . they are linguistically annotated with a CLASSLA-Stanza pipeline and enriched with document-level genre information via a Transformer-based multilingual classifier. |
| Outcome: | The corpora are linguistically annotated with the state-of-the-art CLASSLA-Stanza linguistic processing pipeline and enriched with document-level genre information via the Transformer-based multilingual X-GENRE classifier. |
Machine Learning and Deep Neural Network-Based Lemmatization and Morphosyntactic Tagging for Serbian (2020.lrec-1)
Copied to clipboard
| Challenge: | The training of new tagger models for Serbian is motivated by the enhancement of the existing tagset with the grammatical category of a gender. |
| Approach: | They propose to use TreeTagger and spaCy taggers to train new Serbian tagger models and to align Serbian morphological dictionaries with the grammatical category of a gender. |
| Outcome: | The proposed models achieve 98% PoS-tagging precision per token, and the annotated dataset will be published. |
Multi-layer Annotation of the Rigveda (L18-1)
Copied to clipboard
| Challenge: | Using a multi-level annotation, we present a corpus of the R. GVEDA . |
| Approach: | They propose a multi-level annotation of the R . GVEDA, a Sanskrit text composed in the 2. millenium BCE, and a basic argument identification algorithm to supplement missing verb-argument links. |
| Outcome: | The proposed model replaces verb-argument links by LSTM based model . the proposed model is based on a LS-based model to supplement missing verb-al arguments. |
A Workflow for HTR-Postprocessing, Labeling and Classifying Diachronic and Regional Variation in Pre-Modern Slavic Texts (2024.lrec-main)
Copied to clipboard
Piroska Lendvai, Maarten van Gompel, Anna Jouravel, Elena Renje, Uwe Reichel, Achim Rabus, Eckhart Arnold
| Challenge: | a workflow for classifying diachronic and regional language variation in medieval texts is currently being developed . the workflow is generic or language-agnostic, but can be applied to other historical languages as well. |
| Approach: | They propose a workflow for classifying diachronic and regional language variation in medieval texts . they use handwritten text recognition and manual transcription to obtain the data . |
| Outcome: | The proposed workflow covers HTR-postprocessing, annotating and classifying medieval texts . it is accessible to humanists with limited experience in research data infrastructures, analysis or NLP . |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
The Reference Corpus of the Contemporary Romanian Language (CoRoLa) (L18-1)
Copied to clipboard
| Challenge: | a four-year project focused on the creation of a big corpus for contemporary Romanian language is underway . the corpus is the largest publicly available corpus of contemporary Romania . |
| Approach: | a four-year project is focusing on the creation of a big corpus for Romanian language . the corpus is the largest publicly available corpus of the language based in the country . authors propose to use the corpora as a tool to query and listen to the results . |
| Outcome: | a four-year project has created the largest publicly available corpus of Romanian language . the corpus is the result of a project focused on the creation of 'corola.racai.ro' the written component contains 1,257,752,812 tokens, distributed in several languages . |