Papers by Rinalds Vīksna

3 papers
Annotations for Exploring Food Tweets from Multiple Aspects (2024.lrec-main)

Copied to clipboard

Challenge: The Latvian Twitter Eater Corpus (LTEC) is a collection of tweets gathered by following the appearance of 363 keywords related to food, drinks, eating and drinking in various valid word forms in the Latvian language.
Approach: They build upon the Latvian Twitter Eater Corpus which is focused on the narrow domain of tweets related to food, drinks, eating and drinking.
Outcome: The Latvian Twitter Eater Corpus (LTEC) is a collection of tweets gathered by following the appearance of 363 keywords related to food and eating inflected in various valid word forms in the Latvian language.
MultiLeg: Dataset for Text Sanitisation in Less-resourced Languages (2024.lrec-main)

Copied to clipboard

Challenge: Text sanitization is the task of detecting and removing personal information from the text.
Approach: They propose a dataset for multilingual named entities that can be used for text sanitization.
Outcome: The proposed dataset is available in 8 languages and contains 3082 parallel text segments for each language.
Assessing Multilinguality of Publicly Accessible Websites (2022.lrec-1)

Copied to clipboard

Challenge: multilingualism on the Web is a problem not only at the world level, but also at the European and regional level.
Approach: They propose a tool that automatically analyses the language diversity of the Web and propose indicators and methodologies to measure multilingualism of European websites.
Outcome: The proposed tool can be independently run at set intervals and concludes that multilingualism on the Web is still a problem not only at the world level, but also at the European and regional level.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations