Papers by Marko Tadić

6 papers
M2SA: Multimodal and Multilingual Model for Sentiment Analysis of Tweets (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies on sentiment analysis of tweets focus on the English language . however, there is still a challenge of processing lower-resourced languages .
Approach: They transform tweet sentiment dataset into a multimodal format through a straightforward curation process.
Outcome: The proposed approach performs exceptionally well in unimodal and multimodal configurations.
Introducing the CURLICAT Corpora: Seven-language Domain Specific Annotated Corpora from Curated Sources (2022.lrec-1)

Copied to clipboard

Challenge: The CURLICAT CEF Telecom project aims to collect and deeply annotate a set of large corpora from selected domains.
Approach: They present the results of the CURLICAT CEF Telecom project . they propose to collect and deeply annotate a set of large corpora from selected domains .
Outcome: The CURLICAT CEF Telecom project provides a set of large corpora from selected domains . the corporatized corporates are tokenized, lemmatized and morphologically analysed .
Evaluating Language Tools for Fifteen EU-official Under-resourced Languages (2020.lrec-1)

Copied to clipboard

Challenge: Evaluation of language tools available for 15 EU-official under-resourced languages . evaluation of NERC systems was problematic because of lack of universally or cross-lingually applicable named entities classification scheme.
Approach: They evaluate language tools available for 15 EU-official under-resourced languages . they focus on existing NLP platforms that provide models for under-represented languages - stanton core, nl cube, uDPipe .
Outcome: The evaluation of language tools for 15 under-resourced languages is reproducible . the results are below what was reported in the literature and in some cases even better than the ones reported previously.
The MARCELL Legislative Corpus (2020.lrec-1)

Copied to clipboard

Challenge: MARCELL corpus provides a rich and valuable source for further studies and developments in machine learning, cross-lingual terminological data extraction and classification.
Approach: They present the results of the project MARCELL CEF Telecom . they aim to collect and deeply annotate a large comparable corpus of legal documents .
Outcome: The MARCELL corpus includes 7 monolingual sub-corpora containing the body of respective national legislative documents.
Building the Spanish-Croatian Parallel Corpus (2020.lrec-1)

Copied to clipboard

Challenge: The first unidirectional parallel corpus spanish-croatia was built at the Faculty of Humanities and Social Sciences of the University of Zagreb.
Approach: They describe the building of the first Spanish-Croatian unidirectional parallel corpus at the Faculty of Humanities and Social Sciences of the University of Zagreb.
Outcome: The proposed corpus is a bilingual unidirectional (SpanishCroatian) parallel corpus . it contains 11 Spanish novels and their translations to Croatian done by six translators .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations