Challenge: Distant reading is a new scale of description that does not displace previous scales of literary description.
Approach: They present the Serbian part of the ELTeC multilingual corpus . they propose to test various methods and tools for distant reading .
Outcome: The Serbian part of the ELTeC multilingual corpus is being built to test various methods and tools . Several use examples show that this sub-collection is usefull for both close and distant reading approaches.

Similar Papers

A Lightweight Approach to a Giga-Corpus of Historical Periodicals: The Story of a Slovenian Historical Newspaper Collection (2024.lrec-main)

Copied to clipboard

Challenge: a curated corpus of Slovenian historical newspapers is a complex undertaking requiring multiple steps to prepare . a shoestring budget is required to produce a corpus that is billion-words in size .
Approach: They propose a lightweight approach to producing high-quality corpora using OCR . they use noisy OCR-ed data from the National and University Library of Slovenia .
Outcome: The proposed method produces a billion-word giga-corpus of Slovenian historical newspapers from the 18th, 19th and 20th centuries on a shoestring budget.
The BDCamões Collection of Portuguese Literary Documents: a Research Resource for Digital Humanities and Language Technology (2020.lrec-1)

Copied to clipboard

Challenge: a new corpus of literary documents in Portuguese is presented . it includes close to 4 million words from over 200 complete documents . the corpus is suitable for research in language technology and digital humanities .
Approach: They present the BDCames Collection of Portuguese Literary Documents, a new corpus of literary texts written in Portuguese.
Outcome: The BDCames Collection of Portuguese Literary Documents is a new corpus of literary documents written in Portuguese . it includes close to 4 million words from over 200 complete documents from 83 authors in 14 genres . the corpus is suitable for research in language technology and language science and digital humanities .
Using Bibliodata LODification to Create Metadata-Enriched Literary Corpora in Line with FAIR Principles (2024.lrec-main)

Copied to clipboard

Challenge: Literature corpus building is relatively nascent, and standardized procedures for curating literary corpora are not yet developed.
Approach: They propose a workflow for the creation and reuse of literary corpora using a metadata-enriched Polish Novel Corpus from the 19th and 20th centuries.
Outcome: The proposed workflow includes a multi-stage metadata enrichment and verification process and efficient data collection and data sharing according to the FAIR principles and 5- and 7-star data standards.
SLäNDa: An Annotated Corpus of Narrative and Dialogue in Swedish Literary Fiction (2020.lrec-1)

Copied to clipboard

Challenge: a new corpus of literary fiction has been annotated for cited materials with a focus on dialogue.
Approach: They propose to annotate a new corpus of Swedish literary fiction for cited materials with a focus on dialogue.
Outcome: The proposed corpus can be used to train and analyze models for different types of analysis of literary narrative and speech.
CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation (2024.lrec-main)

Copied to clipboard

Challenge: Using a similar crawling setup, the corpora are comparable across the entire South Slavic language space.
Approach: They propose to collect 13 billion tokens of texts from 26 million documents . they are linguistically annotated with a CLASSLA-Stanza pipeline and enriched with document-level genre information via a Transformer-based multilingual classifier.
Outcome: The corpora are linguistically annotated with the state-of-the-art CLASSLA-Stanza linguistic processing pipeline and enriched with document-level genre information via the Transformer-based multilingual X-GENRE classifier.
Machine Learning and Deep Neural Network-Based Lemmatization and Morphosyntactic Tagging for Serbian (2020.lrec-1)

Copied to clipboard

Challenge: The training of new tagger models for Serbian is motivated by the enhancement of the existing tagset with the grammatical category of a gender.
Approach: They propose to use TreeTagger and spaCy taggers to train new Serbian tagger models and to align Serbian morphological dictionaries with the grammatical category of a gender.
Outcome: The proposed models achieve 98% PoS-tagging precision per token, and the annotated dataset will be published.
Multi-layer Annotation of the Rigveda (L18-1)

Copied to clipboard

Challenge: Using a multi-level annotation, we present a corpus of the R. GVEDA .
Approach: They propose a multi-level annotation of the R . GVEDA, a Sanskrit text composed in the 2. millenium BCE, and a basic argument identification algorithm to supplement missing verb-argument links.
Outcome: The proposed model replaces verb-argument links by LSTM based model . the proposed model is based on a LS-based model to supplement missing verb-al arguments.
A Workflow for HTR-Postprocessing, Labeling and Classifying Diachronic and Regional Variation in Pre-Modern Slavic Texts (2024.lrec-main)

Copied to clipboard

Challenge: a workflow for classifying diachronic and regional language variation in medieval texts is currently being developed . the workflow is generic or language-agnostic, but can be applied to other historical languages as well.
Approach: They propose a workflow for classifying diachronic and regional language variation in medieval texts . they use handwritten text recognition and manual transcription to obtain the data .
Outcome: The proposed workflow covers HTR-postprocessing, annotating and classifying medieval texts . it is accessible to humanists with limited experience in research data infrastructures, analysis or NLP .
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
The Reference Corpus of the Contemporary Romanian Language (CoRoLa) (L18-1)

Copied to clipboard

Challenge: a four-year project focused on the creation of a big corpus for contemporary Romanian language is underway . the corpus is the largest publicly available corpus of contemporary Romania .
Approach: a four-year project is focusing on the creation of a big corpus for Romanian language . the corpus is the largest publicly available corpus of the language based in the country . authors propose to use the corpora as a tool to query and listen to the results .
Outcome: a four-year project has created the largest publicly available corpus of Romanian language . the corpus is the result of a project focused on the creation of 'corola.racai.ro' the written component contains 1,257,752,812 tokens, distributed in several languages .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations