Papers by Jan Hajič

17 papers
Semantic-pragmatic Annotations in the Prague Dependency Treebank (2026.findings-acl)

Copied to clipboard

Challenge: morphology and syntax work on sentence level, but semantic-pragmatic phenomena are often related to two or more neighbouring sentences and possibly to an extra-linguistic context.
Approach: They present semantic-pragmatic specification and annotations in the Prague Dependency Treebank - Consolidated 2.0 release2 by annotating the entire corpus.
Outcome: The proposed annotations are based on the Prague Dependency Treebank -Consolidated 2.0 (PDT-C 2.0) the dataset contains more than 3 million tokens (of Czech) manually annotated from morphology to surface and deep syntax including several types of semantic-pragmatic annotations.
Bridging the LAPPS Grid and CLARIN (L18-1)

Copied to clipboard

Challenge: The LAPPS-CLARIN project is creating a "trust network" between the Language Applications Grid and WebLicht workflow engine . the goal is to allow users on one side of the bridge to gain appropriately authenticated access to the other .
Approach: The LAPPS-CLARIN project is creating a "trust network" between the Language Applications Grid and WebLicht workflow engine hosted by the CLARIN-D Center in Tübingen.
Outcome: The LAPPS-CLARIN project is creating a "trust network" between the Language Applications (LAPPS) Grid and the WebLicht workflow engine hosted by the CLARIN-D Center in Tübingen.
Building a Broad Infrastructure for Uniform Meaning Representations (2024.lrec-main)

Copied to clipboard

Challenge: This paper reports the first release of the UMR data set for six languages . it includes annotations for six different languages that vary greatly in terms of their linguistic properties and resource availability.
Approach: They report the first release of the UMR data set for six languages . they describe on-going efforts to enlarge the data set and extend it to other languages - including Navajo, Navájo, and Sanapaná .
Outcome: The first release of the UMR data set includes annotations for six languages . the language dataset is available for free and can be extended to other languages if needed .
An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT) (2025.acl-long)

Copied to clipboard

Challenge: a large number of textual data is needed to train state-of-the-art large language models.
Approach: They propose a collection of monolingual and parallel corpora from the Internet Archive . they document the entire data pipeline and release the code to reproduce it .
Outcome: The proposed collection of monolingual and parallel corpora is based on the HPLT v2 dataset . it includes 8T tokens covering 193 languages and 380M sentence pairs covering 51 languages .
Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection (2020.lrec-1)

Copied to clipboard

Challenge: Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages.
Approach: They describe version 2 of the universal guidelines and discuss major changes from UD v1 to UD 2 . they propose a morphological layer, a syntactic layer and a word segmentation layer .
Outcome: The proposed treebanks are available for 90 languages and have been updated to meet the needs of multilingual parsers and researchers.
What’s the Meaning of Superhuman Performance in Today’s NLU? (2023.acl-long)

Copied to clipboard

Challenge: Recent research has focused on developing larger pretrained language models and introducing benchmarks such as SuperGLUE and SQuAD to measure their abilities.
Approach: They propose to use benchmarks such as SuperGLUE and SQUAD to evaluate PLMs' abilities in language understanding, reasoning, and reading comprehension to assess their performance.
Outcome: The proposed benchmarks have serious limitations affecting comparison between humans and PLMs and provide recommendations for fairer and more transparent benchmarks.
Creating a Verb Synonym Lexicon Based on a Parallel Corpus (L18-1)

Copied to clipboard

Challenge: a new lexical resource called CzEngClass is being built to help define synonyms in a bilingual context.
Approach: They propose to group verb senses into bilingual verbal synonym groups and use a parallel dependency corpus to explore semantic 'equivalence' they argue that existence of core argument mappings and adjunct mappings to a common set of semantic roles is a suitable criterion for a reasonable verb synonymy definition .
Outcome: The proposed resource will be available by mid-2018 .
Tools for Building an Interlinked Synonym Lexicon Network (L18-1)

Copied to clipboard

Challenge: a new lexicon is being developed for cross-lingual (Czech and English) synonyms based on their syntactic and semantic behavior in (bilingual) context.
Approach: They propose to build a new interlinked verbal synonym lexicon called CzEngClass using a tool that helps to keep cross-lingual synonym classes consistent.
Outcome: The proposed lexicon captures cross-lingual (Czech and English) synonyms . the tool, called Synonym Class Editor -SynEd, is customized to build and edit entries .
European Language Grid: An Overview (2020.lrec-1)

Copied to clipboard

Challenge: European LT business is dominated by hundreds of SMEs and a few large players, with technologies that outperform the global players.
Approach: European Language Grid (ELG) project addresses this by establishing the ELG as the primary platform for LT in Europe.
Outcome: European Language Grid (ELG) will be primary platform for LT in Europe . it will provide access to hundreds of commercial and non-commercial LTs for all European languages, including running tools and services as well as data sets and resources.
SumeCzech: Large Czech News-Based Summarization Dataset (L18-1)

Copied to clipboard

Challenge: Summarization of documents is a well-studied NLP task, but only a few datasets are available for Czech.
Approach: They propose to use a Czech news-based summarization dataset to evaluate document summarizing . they propose a language-agnostic variant of the ROUGE metric to enable automatic evaluation .
Outcome: The proposed dataset contains more than a million Czech news articles . the proposed approach is strong abstractive and language-agnostic .
Synonymy in Bilingual Context: The CzEngClass Lexicon (C18-1)

Copied to clipboard

Challenge: Existing lexical resources for semantic annotation of synonyms are lacking in computational language processing.
Approach: They describe a bilingual lexical resource being built to investigate verbal synonymy in bilingual context and relate semantic roles common to one synonym class to verb arguments.
Outcome: The proposed resource is based on English and Czech WordNet, FrameNet, PropBank, VerbNet (SemLink), and valency lexicons for Czech and English (PDT-Vallex, Vallex, and EngValleX).
Prague Dependency Treebank - Consolidated 1.0 (2020.lrec-1)

Copied to clipboard

Challenge: Using the standard PDT scheme, the Prague Dependency Treebank-Consolidated 1.0 contains 4 different datasets of Czech, uniformly annotated using the standard scheme.
Approach: They present a richly annotated and genre-diversified language resource, the Prague Dependency Treebank-Consolidated 1.0, which contains 4 different datasets of Czech, uniformly annnotated using the standard PDT scheme.
Outcome: The Prague Dependency Treebank-Consolidated 1.0 contains 4 datasets of Czech, uniformly annotated using the standard PDT scheme.
Meaning Representations for Natural Languages: Design, Models and Applications (2024.lrec-tutorials)

Copied to clipboard

Challenge: a tutorial reviews the design of common meaning representations and SoTA models for predicting meaning representation.
Approach: This tutorial reviews the design of common meaning representations and SoTA models for predicting meaning representation. authors propose a cutting-edge, full-day tutorial for all stakeholders in the AI community.
Outcome: This tutorial reviews the design of common meaning representations and SoTA models for predicting meaning representation models . it also reviews the applications of meaning representation in downstream NLP tasks and real-world applications .
LemmaTag: Jointly Tagging and Lemmatizing for Morphologically Rich Languages with BRNNs (D18-1)

Copied to clipboard

Challenge: We compare morphologically rich languages with analytical languages like English due to the large vocabulary size and data sparsity.
Approach: They propose a featureless neural network architecture that generates part-of-speech tags and lemmas for sentences by using bidirectional RNNs with character-level and word-level embeddings.
Outcome: The proposed model outperforms state-of-the-art models in Czech, German, and Arabic.
Textual Coverage of Eventive Entries in Lexical Semantic Resources (2024.lrec-main)

Copied to clipboard

Challenge: Several English, German, Spanish and Czech lexical semantic resources (which, for the most part, focus on verbs and predicates) have been selected for this experiment.
Approach: They propose to quantify coverage gaps in lexical semantic resources when applied to running texts taken from the internet.
Outcome: The proposed resources cover eventive entries (verbs, predicates, etc.) of well-known lexical semantic resources when applied to running texts taken from the internet.
Diacritics Restoration Using Neural Networks (L18-1)

Copied to clipboard

Challenge: a novel combination of character-level recurrent neural network and language model is proposed . people often replace characters with diacritics with their ASCII counterparts .
Approach: They propose a character-level recurrent neural network-based model and a language model for diacritics restoration.
Outcome: The proposed model reduces error of current best systems by 20% to 64% on four languages . it is also able to restore diacritical marks on a number of languages using the same model .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations