Papers by Roberts Darģis
The Use of Text Alignment in Semi-Automatic Error Analysis: Use Case in the Development of the Corpus of the Latvian Language Learners (L18-1)
Copied to clipboard
| Challenge: | Using error annotation methods, the corpus of the Latvian language learners can be adapted for other languages with relatively free word order. |
| Approach: | They propose a method for creating error annotated corpora using text correction, automated morphological analysis, automated text alignment and error annotation. |
| Outcome: | The proposed method has been approbated in the development of the corpus of the Latvian language learners. |
LaVA – Latvian Language Learner corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | a corpus of 1015 essays from foreigners learning Latvian as a foreign language is available at http://www.korpuss.lv/id/LaVA. |
| Approach: | They propose to create a Latvian Language Learner Corpus (LaVA) which contains 1015 essays from Latvian students with different language backgrounds. |
| Outcome: | The LaVA corpus contains 1015 essays from foreigners studying at Latvian higher education institutions and reaching the A1 (possibly A2) Latvian language proficiency level. |
Quality Focused Approach to a Learner Corpus Development (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing methods for annotating learner corpus with errors are not well defined and can be repetitive. |
| Approach: | The paper proposes a quality focused approach to a learner corpus development . the approach includes comparison of digitized texts, text correction, automated morphological analysis and manual review of annotations. |
| Outcome: | The proposed method is used to create a learner corpus in Latvian . it reduces the amount of mistakes that could be introduced due to inconsistent correction or carelessness. |
Development and Evaluation of Speech Synthesis Corpora for Latvian (2020.lrec-1)
Copied to clipboard
| Challenge: | Recent advances in neural speech synthesis have enabled the development of text to speech systems for all languages. |
| Approach: | They propose to obtain a suitable corpus from unannotated Latvian audio recordings using automated speech recognition and speaker segmentation and identification. |
| Outcome: | The proposed method and software tools are applied and evaluated on a Latvian public radio archive data. |
Latvian National Corpora Collection – Korpuss.lv (2022.lrec-1)
Copied to clipboard
Baiba Saulite, Roberts Darģis, Normunds Gruzitis, Ilze Auzina, Kristīne Levāne-Petrova, Lauma Pretkalniņa, Laura Rituma, Peteris Paikens, Arturs Znotins, Laine Strankale, Kristīne Pokratniece, Ilmārs Poikāns, Guntis Barzdins, Inguna Skadiņa, Anda Baklāne, Valdis Saulespurēns, Jānis Ziediņš
| Challenge: | Latvian National Corpora Collection (LNCC) is a multi-institutional and multi-project effort supporting the Latvian language research and language modelling. |
| Approach: | They propose to use Latvian corpora for linguistic research and language modelling. |
| Outcome: | LNCC is a multi-institutional and multi-project effort supported by the Digital Humanities and Language Technology communities in Latvia. |