Papers by Inguna Skadiņa

5 papers
MultiLeg: Dataset for Text Sanitisation in Less-resourced Languages (2024.lrec-main)

Copied to clipboard

Challenge: Text sanitization is the task of detecting and removing personal information from the text.
Approach: They propose a dataset for multilingual named entities that can be used for text sanitization.
Outcome: The proposed dataset is available in 8 languages and contains 3082 parallel text segments for each language.
The Competitiveness Analysis of the European Language Technology Market (2020.lrec-1)

Copied to clipboard

Challenge: The study focuses on three LT areas of the greatest interest for the ECmachine translation (MT), speech technology, and cross-lingual search.
Approach: This paper presents the key results of a competitiveness analysis of the European language technology market for three areas – Machine Translation, speech technology, and cross-lingual search.
Outcome: The study focuses on three LT areas of the greatest interest for the ECmachine translation (MT), speech technology, and cross-lingual search.
Assessing Multilinguality of Publicly Accessible Websites (2022.lrec-1)

Copied to clipboard

Challenge: multilingualism on the Web is a problem not only at the world level, but also at the European and regional level.
Approach: They propose a tool that automatically analyses the language diversity of the Web and propose indicators and methodologies to measure multilingualism of European websites.
Outcome: The proposed tool can be independently run at set intervals and concludes that multilingualism on the Web is still a problem not only at the world level, but also at the European and regional level.
Latvian National Corpora Collection – Korpuss.lv (2022.lrec-1)

Copied to clipboard

Challenge: Latvian National Corpora Collection (LNCC) is a multi-institutional and multi-project effort supporting the Latvian language research and language modelling.
Approach: They propose to use Latvian corpora for linguistic research and language modelling.
Outcome: LNCC is a multi-institutional and multi-project effort supported by the Digital Humanities and Language Technology communities in Latvia.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations