Papers by Jan Hajič
Semantic-pragmatic Annotations in the Prague Dependency Treebank (2026.findings-acl)
Copied to clipboard
Marie Mikulová, Eva Hajicova, Jiří Mírovský, Anna Nedoluzhko, Michal Novák, Pavlína Synková, Jan Štěpánek, Barbora Štěpánková, Jan Hajič
| Challenge: | morphology and syntax work on sentence level, but semantic-pragmatic phenomena are often related to two or more neighbouring sentences and possibly to an extra-linguistic context. |
| Approach: | They present semantic-pragmatic specification and annotations in the Prague Dependency Treebank - Consolidated 2.0 release2 by annotating the entire corpus. |
| Outcome: | The proposed annotations are based on the Prague Dependency Treebank -Consolidated 2.0 (PDT-C 2.0) the dataset contains more than 3 million tokens (of Czech) manually annotated from morphology to surface and deep syntax including several types of semantic-pragmatic annotations. |
The European Language Technology Landscape in 2020: Language-Centric and Human-Centric AI for Cross-Cultural Communication in Multilingual Europe (2020.lrec-1)
Copied to clipboard
Georg Rehm, Katrin Marheinecke, Stefanie Hegele, Stelios Piperidis, Kalina Bontcheva, Jan Hajič, Khalid Choukri, Andrejs Vasiļjevs, Gerhard Backfried, Christoph Prinz, José Manuel Gómez-Pérez, Luc Meertens, Paul Lukowicz, Josef van Genabith, Andrea Lösch, Philipp Slusallek, Morten Irgens, Patrick Gatellier, Joachim Köhler, Laure Le Bars, Dimitra Anastasiou, Albina Auksoriūtė, Núria Bel, António Branco, Gerhard Budin, Walter Daelemans, Koenraad De Smedt, Radovan Garabík, Maria Gavriilidou, Dagmar Gromann, Svetla Koeva, Simon Krek, Cvetana Krstev, Krister Lindén, Bernardo Magnini, Jan Odijk, Maciej Ogrodniczuk, Eiríkur Rögnvaldsson, Mike Rosner, Bolette Pedersen, Inguna Skadiņa, Marko Tadić, Dan Tufiș, Tamás Váradi, Kadri Vider, Andy Way, François Yvon
| Challenge: | Language Technologies (LTs) are a powerful means to break down language barriers impacting business, cross-lingual and cross-cultural communication in Europe. |
| Approach: | They present an overview of the European LT landscape and the current state of play in industry and the LT market. |
| Outcome: | The present study outlines funding programmes, activities, actions and challenges in the different countries with regard to LT, including the current state of play in industry and the LT market. |
Bridging the LAPPS Grid and CLARIN (L18-1)
Copied to clipboard
Erhard Hinrichs, Nancy Ide, James Pustejovsky, Jan Hajič, Marie Hinrichs, Mohammad Fazleh Elahi, Keith Suderman, Marc Verhagen, Kyeongmin Rim, Pavel Straňák, Jozef Mišutka
| Challenge: | The LAPPS-CLARIN project is creating a "trust network" between the Language Applications Grid and WebLicht workflow engine . the goal is to allow users on one side of the bridge to gain appropriately authenticated access to the other . |
| Approach: | The LAPPS-CLARIN project is creating a "trust network" between the Language Applications Grid and WebLicht workflow engine hosted by the CLARIN-D Center in Tübingen. |
| Outcome: | The LAPPS-CLARIN project is creating a "trust network" between the Language Applications (LAPPS) Grid and the WebLicht workflow engine hosted by the CLARIN-D Center in Tübingen. |
Building a Broad Infrastructure for Uniform Meaning Representations (2024.lrec-main)
Copied to clipboard
Julia Bonn, Matthew J. Buchholz, Jayeol Chun, Andrew Cowell, William Croft, Lukas Denk, Sijia Ge, Jan Hajič, Kenneth Lai, James H. Martin, Skatje Myers, Alexis Palmer, Martha Palmer, Claire Benet Post, James Pustejovsky, Kristine Stenzel, Haibo Sun, Zdeňka Urešová, Rosa Vallejos, Jens E. L. Van Gysel, Meagan Vigus, Nianwen Xue, Jin Zhao
| Challenge: | This paper reports the first release of the UMR data set for six languages . it includes annotations for six different languages that vary greatly in terms of their linguistic properties and resource availability. |
| Approach: | They report the first release of the UMR data set for six languages . they describe on-going efforts to enlarge the data set and extend it to other languages - including Navajo, Navájo, and Sanapaná . |
| Outcome: | The first release of the UMR data set includes annotations for six languages . the language dataset is available for free and can be extended to other languages if needed . |
An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT) (2025.acl-long)
Copied to clipboard
Laurie Burchell, Ona De Gibert Bonet, Nikolay Arefyev, Mikko Aulamo, Marta Bañón, Pinzhen Chen, Mariia Fedorova, Liane Guillou, Barry Haddow, Jan Hajič, Jindřich Helcl, Erik Henriksson, Mateusz Klimaszewski, Ville Komulainen, Andrey Kutuzov, Joona Kytöniemi, Veronika Laippala, Petter Mæhlum, Bhavitvya Malik, Farrokh Mehryary, Vladislav Mikhailov, Nikita Moghe, Amanda Myntti, Dayyán O’Brien, Stephan Oepen, Proyag Pal, Jousia Piha, Sampo Pyysalo, Gema Ramírez-Sánchez, David Samuel, Pavel Stepachev, Jörg Tiedemann, Dušan Variš, Tereza Vojtěchová, Jaume Zaragoza-Bernabeu
| Challenge: | a large number of textual data is needed to train state-of-the-art large language models. |
| Approach: | They propose a collection of monolingual and parallel corpora from the Internet Archive . they document the entire data pipeline and release the code to reproduce it . |
| Outcome: | The proposed collection of monolingual and parallel corpora is based on the HPLT v2 dataset . it includes 8T tokens covering 193 languages and 380M sentence pairs covering 51 languages . |
Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection (2020.lrec-1)
Copied to clipboard
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajič, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, Daniel Zeman
| Challenge: | Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages. |
| Approach: | They describe version 2 of the universal guidelines and discuss major changes from UD v1 to UD 2 . they propose a morphological layer, a syntactic layer and a word segmentation layer . |
| Outcome: | The proposed treebanks are available for 90 languages and have been updated to meet the needs of multilingual parsers and researchers. |
What’s the Meaning of Superhuman Performance in Today’s NLU? (2023.acl-long)
Copied to clipboard
Simone Tedeschi, Johan Bos, Thierry Declerck, Jan Hajič, Daniel Hershcovich, Eduard Hovy, Alexander Koller, Simon Krek, Steven Schockaert, Rico Sennrich, Ekaterina Shutova, Roberto Navigli
| Challenge: | Recent research has focused on developing larger pretrained language models and introducing benchmarks such as SuperGLUE and SQuAD to measure their abilities. |
| Approach: | They propose to use benchmarks such as SuperGLUE and SQUAD to evaluate PLMs' abilities in language understanding, reasoning, and reading comprehension to assess their performance. |
| Outcome: | The proposed benchmarks have serious limitations affecting comparison between humans and PLMs and provide recommendations for fairer and more transparent benchmarks. |
Creating a Verb Synonym Lexicon Based on a Parallel Corpus (L18-1)
Copied to clipboard
| Challenge: | a new lexical resource called CzEngClass is being built to help define synonyms in a bilingual context. |
| Approach: | They propose to group verb senses into bilingual verbal synonym groups and use a parallel dependency corpus to explore semantic 'equivalence' they argue that existence of core argument mappings and adjunct mappings to a common set of semantic roles is a suitable criterion for a reasonable verb synonymy definition . |
| Outcome: | The proposed resource will be available by mid-2018 . |
Tools for Building an Interlinked Synonym Lexicon Network (L18-1)
Copied to clipboard
| Challenge: | a new lexicon is being developed for cross-lingual (Czech and English) synonyms based on their syntactic and semantic behavior in (bilingual) context. |
| Approach: | They propose to build a new interlinked verbal synonym lexicon called CzEngClass using a tool that helps to keep cross-lingual synonym classes consistent. |
| Outcome: | The proposed lexicon captures cross-lingual (Czech and English) synonyms . the tool, called Synonym Class Editor -SynEd, is customized to build and edit entries . |
European Language Grid: An Overview (2020.lrec-1)
Copied to clipboard
Georg Rehm, Maria Berger, Ela Elsholz, Stefanie Hegele, Florian Kintzel, Katrin Marheinecke, Stelios Piperidis, Miltos Deligiannis, Dimitris Galanis, Katerina Gkirtzou, Penny Labropoulou, Kalina Bontcheva, David Jones, Ian Roberts, Jan Hajič, Jana Hamrlová, Lukáš Kačena, Khalid Choukri, Victoria Arranz, Andrejs Vasiļjevs, Orians Anvari, Andis Lagzdiņš, Jūlija Meļņika, Gerhard Backfried, Erinç Dikici, Miroslav Janosik, Katja Prinz, Christoph Prinz, Severin Stampler, Dorothea Thomas-Aniola, José Manuel Gómez-Pérez, Andres Garcia Silva, Christian Berrío, Ulrich Germann, Steve Renals, Ondrej Klejch
| Challenge: | European LT business is dominated by hundreds of SMEs and a few large players, with technologies that outperform the global players. |
| Approach: | European Language Grid (ELG) project addresses this by establishing the ELG as the primary platform for LT in Europe. |
| Outcome: | European Language Grid (ELG) will be primary platform for LT in Europe . it will provide access to hundreds of commercial and non-commercial LTs for all European languages, including running tools and services as well as data sets and resources. |
SumeCzech: Large Czech News-Based Summarization Dataset (L18-1)
Copied to clipboard
| Challenge: | Summarization of documents is a well-studied NLP task, but only a few datasets are available for Czech. |
| Approach: | They propose to use a Czech news-based summarization dataset to evaluate document summarizing . they propose a language-agnostic variant of the ROUGE metric to enable automatic evaluation . |
| Outcome: | The proposed dataset contains more than a million Czech news articles . the proposed approach is strong abstractive and language-agnostic . |
Synonymy in Bilingual Context: The CzEngClass Lexicon (C18-1)
Copied to clipboard
| Challenge: | Existing lexical resources for semantic annotation of synonyms are lacking in computational language processing. |
| Approach: | They describe a bilingual lexical resource being built to investigate verbal synonymy in bilingual context and relate semantic roles common to one synonym class to verb arguments. |
| Outcome: | The proposed resource is based on English and Czech WordNet, FrameNet, PropBank, VerbNet (SemLink), and valency lexicons for Czech and English (PDT-Vallex, Vallex, and EngValleX). |
Prague Dependency Treebank - Consolidated 1.0 (2020.lrec-1)
Copied to clipboard
Jan Hajič, Eduard Bejček, Jaroslava Hlavacova, Marie Mikulová, Milan Straka, Jan Štěpánek, Barbora Štěpánková
| Challenge: | Using the standard PDT scheme, the Prague Dependency Treebank-Consolidated 1.0 contains 4 different datasets of Czech, uniformly annotated using the standard scheme. |
| Approach: | They present a richly annotated and genre-diversified language resource, the Prague Dependency Treebank-Consolidated 1.0, which contains 4 different datasets of Czech, uniformly annnotated using the standard PDT scheme. |
| Outcome: | The Prague Dependency Treebank-Consolidated 1.0 contains 4 datasets of Czech, uniformly annotated using the standard PDT scheme. |
Meaning Representations for Natural Languages: Design, Models and Applications (2024.lrec-tutorials)
Copied to clipboard
| Challenge: | a tutorial reviews the design of common meaning representations and SoTA models for predicting meaning representation. |
| Approach: | This tutorial reviews the design of common meaning representations and SoTA models for predicting meaning representation. authors propose a cutting-edge, full-day tutorial for all stakeholders in the AI community. |
| Outcome: | This tutorial reviews the design of common meaning representations and SoTA models for predicting meaning representation models . it also reviews the applications of meaning representation in downstream NLP tasks and real-world applications . |
LemmaTag: Jointly Tagging and Lemmatizing for Morphologically Rich Languages with BRNNs (D18-1)
Copied to clipboard
| Challenge: | We compare morphologically rich languages with analytical languages like English due to the large vocabulary size and data sparsity. |
| Approach: | They propose a featureless neural network architecture that generates part-of-speech tags and lemmas for sentences by using bidirectional RNNs with character-level and word-level embeddings. |
| Outcome: | The proposed model outperforms state-of-the-art models in Czech, German, and Arabic. |
Textual Coverage of Eventive Entries in Lexical Semantic Resources (2024.lrec-main)
Copied to clipboard
| Challenge: | Several English, German, Spanish and Czech lexical semantic resources (which, for the most part, focus on verbs and predicates) have been selected for this experiment. |
| Approach: | They propose to quantify coverage gaps in lexical semantic resources when applied to running texts taken from the internet. |
| Outcome: | The proposed resources cover eventive entries (verbs, predicates, etc.) of well-known lexical semantic resources when applied to running texts taken from the internet. |
Diacritics Restoration Using Neural Networks (L18-1)
Copied to clipboard
| Challenge: | a novel combination of character-level recurrent neural network and language model is proposed . people often replace characters with diacritics with their ASCII counterparts . |
| Approach: | They propose a character-level recurrent neural network-based model and a language model for diacritics restoration. |
| Outcome: | The proposed model reduces error of current best systems by 20% to 64% on four languages . it is also able to restore diacritical marks on a number of languages using the same model . |