Papers by Simon Krek
The European Language Technology Landscape in 2020: Language-Centric and Human-Centric AI for Cross-Cultural Communication in Multilingual Europe (2020.lrec-1)
Copied to clipboard
Georg Rehm, Katrin Marheinecke, Stefanie Hegele, Stelios Piperidis, Kalina Bontcheva, Jan Hajič, Khalid Choukri, Andrejs Vasiļjevs, Gerhard Backfried, Christoph Prinz, José Manuel Gómez-Pérez, Luc Meertens, Paul Lukowicz, Josef van Genabith, Andrea Lösch, Philipp Slusallek, Morten Irgens, Patrick Gatellier, Joachim Köhler, Laure Le Bars, Dimitra Anastasiou, Albina Auksoriūtė, Núria Bel, António Branco, Gerhard Budin, Walter Daelemans, Koenraad De Smedt, Radovan Garabík, Maria Gavriilidou, Dagmar Gromann, Svetla Koeva, Simon Krek, Cvetana Krstev, Krister Lindén, Bernardo Magnini, Jan Odijk, Maciej Ogrodniczuk, Eiríkur Rögnvaldsson, Mike Rosner, Bolette Pedersen, Inguna Skadiņa, Marko Tadić, Dan Tufiș, Tamás Váradi, Kadri Vider, Andy Way, François Yvon
| Challenge: | Language Technologies (LTs) are a powerful means to break down language barriers impacting business, cross-lingual and cross-cultural communication in Europe. |
| Approach: | They present an overview of the European LT landscape and the current state of play in industry and the LT market. |
| Outcome: | The present study outlines funding programmes, activities, actions and challenges in the different countries with regard to LT, including the current state of play in industry and the LT market. |
Gigafida 2.0: The Reference Corpus of Written Standard Slovene (2020.lrec-1)
Copied to clipboard
Simon Krek, Špela Arhar Holdt, Tomaž Erjavec, Jaka Čibej, Andraz Repar, Polona Gantar, Nikola Ljubešić, Iztok Kosem, Kaja Dobrovoljc
| Challenge: | Gigafida reference corpus of Slovene is updated with new material and tools . focus of upgrade was on transformation from general reference corp to standard reference corp . |
| Approach: | We present a new version of the Gigafida reference corpus of Slovene . the upgrade includes new material and better tools for annotating it . |
| Outcome: | The new version of the Gigafida reference corpus of Slovene is described . the whole Gigido corpus was deduplicated for the first time . |
Introducing the CURLICAT Corpora: Seven-language Domain Specific Annotated Corpora from Curated Sources (2022.lrec-1)
Copied to clipboard
Tamás Váradi, Bence Nyéki, Svetla Koeva, Marko Tadić, Vanja Štefanec, Maciej Ogrodniczuk, Bartłomiej Nitoń, Piotr Pęzik, Verginica Barbu Mititelu, Elena Irimia, Maria Mitrofan, Dan Tufiș, Radovan Garabík, Simon Krek, Andraž Repar
| Challenge: | The CURLICAT CEF Telecom project aims to collect and deeply annotate a set of large corpora from selected domains. |
| Approach: | They present the results of the CURLICAT CEF Telecom project . they propose to collect and deeply annotate a set of large corpora from selected domains . |
| Outcome: | The CURLICAT CEF Telecom project provides a set of large corpora from selected domains . the corporatized corporates are tokenized, lemmatized and morphologically analysed . |
What’s the Meaning of Superhuman Performance in Today’s NLU? (2023.acl-long)
Copied to clipboard
Simone Tedeschi, Johan Bos, Thierry Declerck, Jan Hajič, Daniel Hershcovich, Eduard Hovy, Alexander Koller, Simon Krek, Steven Schockaert, Rico Sennrich, Ekaterina Shutova, Roberto Navigli
| Challenge: | Recent research has focused on developing larger pretrained language models and introducing benchmarks such as SuperGLUE and SQuAD to measure their abilities. |
| Approach: | They propose to use benchmarks such as SuperGLUE and SQUAD to evaluate PLMs' abilities in language understanding, reasoning, and reading comprehension to assess their performance. |
| Outcome: | The proposed benchmarks have serious limitations affecting comparison between humans and PLMs and provide recommendations for fairer and more transparent benchmarks. |
A Multilingual Evaluation Dataset for Monolingual Word Sense Alignment (2020.lrec-1)
Copied to clipboard
Sina Ahmadi, John Philip McCrae, Sanni Nimb, Fahad Khan, Monica Monachini, Bolette Pedersen, Thierry Declerck, Tanja Wissik, Andrea Bellandi, Irene Pisani, Thomas Troelsgård, Sussi Olsen, Simon Krek, Veronika Lipp, Tamás Váradi, László Simon, András Gyorffy, Carole Tiberius, Tanneke Schoonheim, Yifat Ben Moshe, Maya Rudich, Raya Abu Ahmad, Dorielle Lonke, Kira Kovalenko, Margit Langemets, Jelena Kallas, Oksana Dereza, Theodorus Fransen, David Cillessen, David Lindemann, Mikel Alonso, Ana Salgado, José Luis Sancho, Rafael-J. Ureña-Ruiz, Jordi Porta Zamorano, Kiril Simov, Petya Osenova, Zara Kancheva, Ivaylo Radev, Ranka Stanković, Andrej Perdih, Dejan Gabrovsek
| Challenge: | a new dataset aims to align monolingual dictionaries with a single sense level for 15 languages . this dataset covers a wide range of languages and resources . |
| Approach: | They propose to manually align monolingual dictionaries with possible semantic relationships . they use 15 languages to create a new baseline for the task of monolingual word sense alignment . |
| Outcome: | The proposed dataset covers 15 languages and covers the more challenging task of linking general-purpose language. |
SUK 1.0: A New Training Corpus for Linguistic Annotation of Modern Standard Slovene (2024.lrec-main)
Copied to clipboard
Špela Arhar Holdt, Jaka Čibej, Kaja Dobrovoljc, Tomaž Erjavec, Polona Gantar, Simon Krek, Tina Munda, Nejc Robida, Luka Terčon, Slavko Zitnik
| Challenge: | a training corpus for linguistic annotation of modern standard Slovene has been in continuous development for 15 years. |
| Approach: | They introduce an upgrade of a training corpus for linguistic annotation of modern standard Slovene. |
| Outcome: | The revised corpus, built on its predecessor, doubles in size and depth of annotation layers. |
The MARCELL Legislative Corpus (2020.lrec-1)
Copied to clipboard
Tamás Váradi, Svetla Koeva, Martin Yamalov, Marko Tadić, Bálint Sass, Bartłomiej Nitoń, Maciej Ogrodniczuk, Piotr Pęzik, Verginica Barbu Mititelu, Radu Ion, Elena Irimia, Maria Mitrofan, Vasile Păiș, Dan Tufiș, Radovan Garabík, Simon Krek, Andraz Repar, Matjaž Rihtar, Janez Brank
| Challenge: | MARCELL corpus provides a rich and valuable source for further studies and developments in machine learning, cross-lingual terminological data extraction and classification. |
| Approach: | They present the results of the project MARCELL CEF Telecom . they aim to collect and deeply annotate a large comparable corpus of legal documents . |
| Outcome: | The MARCELL corpus includes 7 monolingual sub-corpora containing the body of respective national legislative documents. |