Papers by Rinalds Vīksna
Annotations for Exploring Food Tweets from Multiple Aspects (2024.lrec-main)
Copied to clipboard
| Challenge: | The Latvian Twitter Eater Corpus (LTEC) is a collection of tweets gathered by following the appearance of 363 keywords related to food, drinks, eating and drinking in various valid word forms in the Latvian language. |
| Approach: | They build upon the Latvian Twitter Eater Corpus which is focused on the narrow domain of tweets related to food, drinks, eating and drinking. |
| Outcome: | The Latvian Twitter Eater Corpus (LTEC) is a collection of tweets gathered by following the appearance of 363 keywords related to food and eating inflected in various valid word forms in the Latvian language. |
MultiLeg: Dataset for Text Sanitisation in Less-resourced Languages (2024.lrec-main)
Copied to clipboard
| Challenge: | Text sanitization is the task of detecting and removing personal information from the text. |
| Approach: | They propose a dataset for multilingual named entities that can be used for text sanitization. |
| Outcome: | The proposed dataset is available in 8 languages and contains 3082 parallel text segments for each language. |
Assessing Multilinguality of Publicly Accessible Websites (2022.lrec-1)
Copied to clipboard
| Challenge: | multilingualism on the Web is a problem not only at the world level, but also at the European and regional level. |
| Approach: | They propose a tool that automatically analyses the language diversity of the Web and propose indicators and methodologies to measure multilingualism of European websites. |
| Outcome: | The proposed tool can be independently run at set intervals and concludes that multilingualism on the Web is still a problem not only at the world level, but also at the European and regional level. |