Baiba Saulite, Roberts Darģis, Normunds Gruzitis, Ilze Auzina, Kristīne Levāne-Petrova, Lauma Pretkalniņa, Laura Rituma, Peteris Paikens, Arturs Znotins, Laine Strankale, Kristīne Pokratniece, Ilmārs Poikāns, Guntis Barzdins, Inguna Skadiņa, Anda Baklāne, Valdis Saulespurēns, Jānis Ziediņš
| Challenge: | Latvian National Corpora Collection (LNCC) is a multi-institutional and multi-project effort supporting the Latvian language research and language modelling. |
| Approach: | They propose to use Latvian corpora for linguistic research and language modelling. |
| Outcome: | LNCC is a multi-institutional and multi-project effort supported by the Digital Humanities and Language Technology communities in Latvia. |
Similar Papers
BalsuTalka.lv - Boosting the Common Voice Corpus for Low-Resource Languages (2024.lrec-main)
Copied to clipboard
Roberts Dargis, Arturs Znotins, Ilze Auzina, Baiba Saulite, Sanita Reinsone, Raivis Dejus, Antra Klavinska, Normunds Gruzitis
| Challenge: | Latvian is a low-resource language for many NLP tasks, but most speech corpora are closed data . a crowdsourcing campaign to create a relatively large, diverse and open speech corpus for Latvian has been launched . |
| Approach: | a crowdsourcing campaign is helping to create an open speech corpus for Latvian . the goal is to enlarge the datasets and make them more diverse . authors use the opensource Mozilla Common Voice platform to validate speech samples . |
| Outcome: | a crowdsourcing initiative has increased the size and speaker diversity of the Latvian Common Voice 17.0 dataset by more than tenfold in less than a year. |
Towards Latvian WordNet (2022.lrec-1)
Copied to clipboard
Peteris Paikens, Mikus Grasmanis, Agute Klints, Ilze Lokmane, Lauma Pretkalniņa, Laura Rituma, Madara Stāde, Laine Strankale
| Challenge: | Currently the dataset consists of 6432 words linked in 5528 synsets . the goal is to provide a structured lexical-semantic resource for Latvian word sense disambiguation . |
| Approach: | They propose to use Princeton's word sense definition and sense linking principles to create a Latvian wordnet . they use corpus evidence and an online dictionary to build a lexical-semantic resource . |
| Outcome: | The proposed resource is based on the Princeton WordNet and is available in Latvian . the initial portion of the data is available for download . |
A Computational Model of Latvian Morphology (2024.lrec-main)
Copied to clipboard
| Challenge: | a computational model of Latvian morphology provides a formal structure for Latvian word form inflection . the model explicitly enumerates and handles the many exceptions to the general Latvian inflation principles . |
| Approach: | They propose a computational model of Latvian morphology that provides a formal structure for Latvian word form inflection. |
| Outcome: | The proposed model provides a good coverage for modern Latvian literary language and potential to extend to Latgalian language. |
LaVA – Latvian Language Learner corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | a corpus of 1015 essays from foreigners learning Latvian as a foreign language is available at http://www.korpuss.lv/id/LaVA. |
| Approach: | They propose to create a Latvian Language Learner Corpus (LaVA) which contains 1015 essays from Latvian students with different language backgrounds. |
| Outcome: | The LaVA corpus contains 1015 essays from foreigners studying at Latvian higher education institutions and reaching the A1 (possibly A2) Latvian language proficiency level. |
Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018) (L18-1)
Copied to clipboard
| Challenge: | null |
| Approach: | null |
| Outcome: | null |
A Progress Report on Activities at the Linguistic Data Consortium Benefitting the LREC Community (2020.lrec-1)
Copied to clipboard
Christopher Cieri, James Fiumara, Stephanie Strassel, Jonathan Wright, Denise DiPersio, Mark Liberman
| Challenge: | Linguistic Data Consortium (LDC) activities include the collection, annotation, processing, distribution, archiving and curation of language resources. |
| Approach: | a new report sketches the activities of a data center devoted to supporting the work of LREC attendees . 96 new corpora released in 2018-2020 to date, a technology evaluation campaign and innovations to advance methodology for language data collection and annotation. |
| Outcome: | 96 new corpora released in 2018-2020 to date, new technology evaluation campaign and innovations to advance methodology of language data collection and annotation. |
The LREC Workshops Map (L18-1)
Copied to clipboard
| Challenge: | a corpus of workshops titles and related presentations has been retrieved from the conference's website . data is used to analyze the research presented at the conferences over the years 1998-2016 . |
| Approach: | a paper aims to present an overview of the research presented at the LREC workshops over the years 1998-2016. |
| Outcome: | The aim of the present study is to shed light on the community represented by workshop participants over the years 1998-2016. |
CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation (2024.lrec-main)
Copied to clipboard
| Challenge: | Using a similar crawling setup, the corpora are comparable across the entire South Slavic language space. |
| Approach: | They propose to collect 13 billion tokens of texts from 26 million documents . they are linguistically annotated with a CLASSLA-Stanza pipeline and enriched with document-level genre information via a Transformer-based multilingual classifier. |
| Outcome: | The corpora are linguistically annotated with the state-of-the-art CLASSLA-Stanza linguistic processing pipeline and enriched with document-level genre information via the Transformer-based multilingual X-GENRE classifier. |
Code-Mixed Text Augmentation for Latvian ASR (2024.lrec-main)
Copied to clipboard
| Challenge: | a new study attempts to tackle code-mixed speech recognition by improving the language model of a hybrid system. |
| Approach: | They propose an inflected transliteration and phonetic transcription model for code-mixed Latvian sentences . they leverage a large human-translated English-Latvian parallel text corpus to generate synthetic Latvian phrases . |
| Outcome: | The proposed system improves on a human-translated English-Latvian parallel text corpus . the results show that the proposed system can generate code-mixed Latvian sentences . |
Introducing the CURLICAT Corpora: Seven-language Domain Specific Annotated Corpora from Curated Sources (2022.lrec-1)
Copied to clipboard
Tamás Váradi, Bence Nyéki, Svetla Koeva, Marko Tadić, Vanja Štefanec, Maciej Ogrodniczuk, Bartłomiej Nitoń, Piotr Pęzik, Verginica Barbu Mititelu, Elena Irimia, Maria Mitrofan, Dan Tufiș, Radovan Garabík, Simon Krek, Andraž Repar
| Challenge: | The CURLICAT CEF Telecom project aims to collect and deeply annotate a set of large corpora from selected domains. |
| Approach: | They present the results of the CURLICAT CEF Telecom project . they propose to collect and deeply annotate a set of large corpora from selected domains . |
| Outcome: | The CURLICAT CEF Telecom project provides a set of large corpora from selected domains . the corporatized corporates are tokenized, lemmatized and morphologically analysed . |