Latvian National Corpora Collection – Korpuss.lv (2022.lrec-1)

Copied to clipboard

Challenge: Latvian National Corpora Collection (LNCC) is a multi-institutional and multi-project effort supporting the Latvian language research and language modelling.
Approach: They propose to use Latvian corpora for linguistic research and language modelling.
Outcome: LNCC is a multi-institutional and multi-project effort supported by the Digital Humanities and Language Technology communities in Latvia.

Similar Papers

BalsuTalka.lv - Boosting the Common Voice Corpus for Low-Resource Languages (2024.lrec-main)

Copied to clipboard

Challenge: Latvian is a low-resource language for many NLP tasks, but most speech corpora are closed data . a crowdsourcing campaign to create a relatively large, diverse and open speech corpus for Latvian has been launched .
Approach: a crowdsourcing campaign is helping to create an open speech corpus for Latvian . the goal is to enlarge the datasets and make them more diverse . authors use the opensource Mozilla Common Voice platform to validate speech samples .
Outcome: a crowdsourcing initiative has increased the size and speaker diversity of the Latvian Common Voice 17.0 dataset by more than tenfold in less than a year.
Towards Latvian WordNet (2022.lrec-1)

Copied to clipboard

Challenge: Currently the dataset consists of 6432 words linked in 5528 synsets . the goal is to provide a structured lexical-semantic resource for Latvian word sense disambiguation .
Approach: They propose to use Princeton's word sense definition and sense linking principles to create a Latvian wordnet . they use corpus evidence and an online dictionary to build a lexical-semantic resource .
Outcome: The proposed resource is based on the Princeton WordNet and is available in Latvian . the initial portion of the data is available for download .
A Computational Model of Latvian Morphology (2024.lrec-main)

Copied to clipboard

Challenge: a computational model of Latvian morphology provides a formal structure for Latvian word form inflection . the model explicitly enumerates and handles the many exceptions to the general Latvian inflation principles .
Approach: They propose a computational model of Latvian morphology that provides a formal structure for Latvian word form inflection.
Outcome: The proposed model provides a good coverage for modern Latvian literary language and potential to extend to Latgalian language.
LaVA – Latvian Language Learner corpus (2022.lrec-1)

Copied to clipboard

Challenge: a corpus of 1015 essays from foreigners learning Latvian as a foreign language is available at http://www.korpuss.lv/id/LaVA.
Approach: They propose to create a Latvian Language Learner Corpus (LaVA) which contains 1015 essays from Latvian students with different language backgrounds.
Outcome: The LaVA corpus contains 1015 essays from foreigners studying at Latvian higher education institutions and reaching the A1 (possibly A2) Latvian language proficiency level.
A Progress Report on Activities at the Linguistic Data Consortium Benefitting the LREC Community (2020.lrec-1)

Copied to clipboard

Challenge: Linguistic Data Consortium (LDC) activities include the collection, annotation, processing, distribution, archiving and curation of language resources.
Approach: a new report sketches the activities of a data center devoted to supporting the work of LREC attendees . 96 new corpora released in 2018-2020 to date, a technology evaluation campaign and innovations to advance methodology for language data collection and annotation.
Outcome: 96 new corpora released in 2018-2020 to date, new technology evaluation campaign and innovations to advance methodology of language data collection and annotation.
The LREC Workshops Map (L18-1)

Copied to clipboard

Challenge: a corpus of workshops titles and related presentations has been retrieved from the conference's website . data is used to analyze the research presented at the conferences over the years 1998-2016 .
Approach: a paper aims to present an overview of the research presented at the LREC workshops over the years 1998-2016.
Outcome: The aim of the present study is to shed light on the community represented by workshop participants over the years 1998-2016.
CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation (2024.lrec-main)

Copied to clipboard

Challenge: Using a similar crawling setup, the corpora are comparable across the entire South Slavic language space.
Approach: They propose to collect 13 billion tokens of texts from 26 million documents . they are linguistically annotated with a CLASSLA-Stanza pipeline and enriched with document-level genre information via a Transformer-based multilingual classifier.
Outcome: The corpora are linguistically annotated with the state-of-the-art CLASSLA-Stanza linguistic processing pipeline and enriched with document-level genre information via the Transformer-based multilingual X-GENRE classifier.
Code-Mixed Text Augmentation for Latvian ASR (2024.lrec-main)

Copied to clipboard

Challenge: a new study attempts to tackle code-mixed speech recognition by improving the language model of a hybrid system.
Approach: They propose an inflected transliteration and phonetic transcription model for code-mixed Latvian sentences . they leverage a large human-translated English-Latvian parallel text corpus to generate synthetic Latvian phrases .
Outcome: The proposed system improves on a human-translated English-Latvian parallel text corpus . the results show that the proposed system can generate code-mixed Latvian sentences .
Introducing the CURLICAT Corpora: Seven-language Domain Specific Annotated Corpora from Curated Sources (2022.lrec-1)

Copied to clipboard

Challenge: The CURLICAT CEF Telecom project aims to collect and deeply annotate a set of large corpora from selected domains.
Approach: They present the results of the CURLICAT CEF Telecom project . they propose to collect and deeply annotate a set of large corpora from selected domains .
Outcome: The CURLICAT CEF Telecom project provides a set of large corpora from selected domains . the corporatized corporates are tokenized, lemmatized and morphologically analysed .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations