Papers by Ranka Stanković

5 papers
Using English Baits to Catch Serbian Multi-Word Terminology (L18-1)

Copied to clipboard

Challenge: a new method for bilingual terminology extraction is proposed for a source language and a target language.
Approach: They propose to use a bilingual terminology extraction approach for a source language and a target language to extract the terminology for sri lanka.
Outcome: The proposed method extracts terminology for a source language and a target language from it.
Bridging Computational Lexicography and Corpus Linguistics: A Query Extension for OntoLex-FrAC (2024.lrec-main)

Copied to clipboard

Challenge: OntoLex is the dominant community standard for machine-readable lexical resources . it is currently extended with a designated module for Frequency, Attestations and Corpus-based Information .
Approach: They propose a module for Frequency, Attestations and Corpus-based Information for OntoLex . the module enables RDF-based web services to exchange corpus queries dynamically .
Outcome: The proposed module addresses the incorporation of corpus queries for linking dictionaries with corpus engines and enabling RDF-based web services to exchange corpus query data dynamically.
Distant Reading in Digital Humanities: Case Study on the Serbian Part of the ELTeC Collection (2022.lrec-1)

Copied to clipboard

Challenge: Distant reading is a new scale of description that does not displace previous scales of literary description.
Approach: They present the Serbian part of the ELTeC multilingual corpus . they propose to test various methods and tools for distant reading .
Outcome: The Serbian part of the ELTeC multilingual corpus is being built to test various methods and tools . Several use examples show that this sub-collection is usefull for both close and distant reading approaches.
A Multilingual Evaluation Dataset for Monolingual Word Sense Alignment (2020.lrec-1)

Copied to clipboard

Challenge: a new dataset aims to align monolingual dictionaries with a single sense level for 15 languages . this dataset covers a wide range of languages and resources .
Approach: They propose to manually align monolingual dictionaries with possible semantic relationships . they use 15 languages to create a new baseline for the task of monolingual word sense alignment .
Outcome: The proposed dataset covers 15 languages and covers the more challenging task of linking general-purpose language.
MultiLexBATS: Multilingual Dataset of Lexical Semantic Relations (2024.lrec-main)

Copied to clipboard

Challenge: Prior work has focused on analysing lexical semantic relations in word embeddings or probing pretrained language models (PLMs) with some exceptions.
Approach: They propose to use a multilingual parallel dataset of lexical semantic relations adapted from BATS in 15 languages including low-resource languages such as Bambara, Lithuanian, and Albanian as an experiment on cross-lingual transfer of relational knowledge.
Outcome: The proposed dataset is adapted from a BATS-based dataset in 15 languages including low-resource languages such as Bambara, Lithuanian, and Albanian.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations