Papers by Attila Novák

5 papers
Much Ado About Nothing – Identification of Zero Copulas in Hungarian Using an NMT Model (2020.lrec-1)

Copied to clipboard

Challenge: Zero copulas are the phenomenon that nominal predicates lack an explicit verbal copule in default present tense 3rd person indicative cases.
Approach: They propose a tool that can identify and mark the location of zero copulas in Hungarian clauses that contain nominal predicates at the right position.
Outcome: The proposed tool can identify and mark the location of zero copulas, i.e. where an overt copulan would appear in the non-default cases.
E-magyar – A Digital Language Processing System (L18-1)

Copied to clipboard

Challenge: e-magyar is a free, open, modular text processing pipeline for Hungarian . existing tools were overhauled to operate in the pipeline with a uniform encoding and run in the same Java platform.
Approach: e-magyar is a free, open, modular text processing pipeline for Hungarian . it was created by a collaborative effort by the language technology community . the system is aimed at a broad range of users, from language developers to researchers .
Outcome: The proposed tool is open source and available for download on the HFST framework.
NerKor+Cars-OntoNotes++ (2022.lrec-1)

Copied to clipboard

Challenge: In this paper, we present an upgraded version of the Hungarian NYTK-NerKor named entity corpus . it contains twice as many annotated spans and 7 times as many distinct entity types as the original version.
Approach: They present an upgraded version of the Hungarian NYTK-NerKor named entity corpus with an extended OntoNotes 5 annotation scheme.
Outcome: The enhanced version of the corpus contains twice as many annotated spans and 7 times more distinct entity types than the original version.
CBOW-tag: a Modified CBOW Algorithm for Generating Embedding Models from Annotated Corpora (2020.lrec-1)

Copied to clipboard

Challenge: Using word2vec, we train distributional semantic models that predict a word from the context or vice versa.
Approach: They propose a modified version of the CBOW algorithm implemented in the fastText framework that includes the representation of original word forms and their annotation at the same time.
Outcome: The proposed model can answer questions such as What do we eat?, What can we do with a skeleton?, etc.
Cross-Lingual Generation and Evaluation of a Wide-Coverage Lexical Semantic Resource (L18-1)

Copied to clipboard

Challenge: Neural word embedding models are not interpretable for humans by themselves . we present a method that assigns explicit symbolic semantic features to words .
Approach: They propose a method that assigns explicit symbolic semantic features to words in an embedding model . they use a finite list of terms to make the model interpretable for humans .
Outcome: The proposed method is shown to be very efficient for word embedding models . it can be applied across languages and can be used as a searchable semantic annotation .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations