Papers by Marko Tadić
M2SA: Multimodal and Multilingual Model for Sentiment Analysis of Tweets (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing studies on sentiment analysis of tweets focus on the English language . however, there is still a challenge of processing lower-resourced languages . |
| Approach: | They transform tweet sentiment dataset into a multimodal format through a straightforward curation process. |
| Outcome: | The proposed approach performs exceptionally well in unimodal and multimodal configurations. |
The European Language Technology Landscape in 2020: Language-Centric and Human-Centric AI for Cross-Cultural Communication in Multilingual Europe (2020.lrec-1)
Copied to clipboard
Georg Rehm, Katrin Marheinecke, Stefanie Hegele, Stelios Piperidis, Kalina Bontcheva, Jan Hajič, Khalid Choukri, Andrejs Vasiļjevs, Gerhard Backfried, Christoph Prinz, José Manuel Gómez-Pérez, Luc Meertens, Paul Lukowicz, Josef van Genabith, Andrea Lösch, Philipp Slusallek, Morten Irgens, Patrick Gatellier, Joachim Köhler, Laure Le Bars, Dimitra Anastasiou, Albina Auksoriūtė, Núria Bel, António Branco, Gerhard Budin, Walter Daelemans, Koenraad De Smedt, Radovan Garabík, Maria Gavriilidou, Dagmar Gromann, Svetla Koeva, Simon Krek, Cvetana Krstev, Krister Lindén, Bernardo Magnini, Jan Odijk, Maciej Ogrodniczuk, Eiríkur Rögnvaldsson, Mike Rosner, Bolette Pedersen, Inguna Skadiņa, Marko Tadić, Dan Tufiș, Tamás Váradi, Kadri Vider, Andy Way, François Yvon
| Challenge: | Language Technologies (LTs) are a powerful means to break down language barriers impacting business, cross-lingual and cross-cultural communication in Europe. |
| Approach: | They present an overview of the European LT landscape and the current state of play in industry and the LT market. |
| Outcome: | The present study outlines funding programmes, activities, actions and challenges in the different countries with regard to LT, including the current state of play in industry and the LT market. |
Introducing the CURLICAT Corpora: Seven-language Domain Specific Annotated Corpora from Curated Sources (2022.lrec-1)
Copied to clipboard
Tamás Váradi, Bence Nyéki, Svetla Koeva, Marko Tadić, Vanja Štefanec, Maciej Ogrodniczuk, Bartłomiej Nitoń, Piotr Pęzik, Verginica Barbu Mititelu, Elena Irimia, Maria Mitrofan, Dan Tufiș, Radovan Garabík, Simon Krek, Andraž Repar
| Challenge: | The CURLICAT CEF Telecom project aims to collect and deeply annotate a set of large corpora from selected domains. |
| Approach: | They present the results of the CURLICAT CEF Telecom project . they propose to collect and deeply annotate a set of large corpora from selected domains . |
| Outcome: | The CURLICAT CEF Telecom project provides a set of large corpora from selected domains . the corporatized corporates are tokenized, lemmatized and morphologically analysed . |
Evaluating Language Tools for Fifteen EU-official Under-resourced Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Evaluation of language tools available for 15 EU-official under-resourced languages . evaluation of NERC systems was problematic because of lack of universally or cross-lingually applicable named entities classification scheme. |
| Approach: | They evaluate language tools available for 15 EU-official under-resourced languages . they focus on existing NLP platforms that provide models for under-represented languages - stanton core, nl cube, uDPipe . |
| Outcome: | The evaluation of language tools for 15 under-resourced languages is reproducible . the results are below what was reported in the literature and in some cases even better than the ones reported previously. |
The MARCELL Legislative Corpus (2020.lrec-1)
Copied to clipboard
Tamás Váradi, Svetla Koeva, Martin Yamalov, Marko Tadić, Bálint Sass, Bartłomiej Nitoń, Maciej Ogrodniczuk, Piotr Pęzik, Verginica Barbu Mititelu, Radu Ion, Elena Irimia, Maria Mitrofan, Vasile Păiș, Dan Tufiș, Radovan Garabík, Simon Krek, Andraz Repar, Matjaž Rihtar, Janez Brank
| Challenge: | MARCELL corpus provides a rich and valuable source for further studies and developments in machine learning, cross-lingual terminological data extraction and classification. |
| Approach: | They present the results of the project MARCELL CEF Telecom . they aim to collect and deeply annotate a large comparable corpus of legal documents . |
| Outcome: | The MARCELL corpus includes 7 monolingual sub-corpora containing the body of respective national legislative documents. |
Building the Spanish-Croatian Parallel Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | The first unidirectional parallel corpus spanish-croatia was built at the Faculty of Humanities and Social Sciences of the University of Zagreb. |
| Approach: | They describe the building of the first Spanish-Croatian unidirectional parallel corpus at the Faculty of Humanities and Social Sciences of the University of Zagreb. |
| Outcome: | The proposed corpus is a bilingual unidirectional (SpanishCroatian) parallel corpus . it contains 11 Spanish novels and their translations to Croatian done by six translators . |