Tamás Váradi, Eszter Simon, Bálint Sass, Iván Mittelholcz, Attila Novák, Balázs Indig, Richárd Farkas, Veronika Vincze
| Challenge: | e-magyar is a free, open, modular text processing pipeline for Hungarian . existing tools were overhauled to operate in the pipeline with a uniform encoding and run in the same Java platform. |
| Approach: | e-magyar is a free, open, modular text processing pipeline for Hungarian . it was created by a collaborative effort by the language technology community . the system is aimed at a broad range of users, from language developers to researchers . |
| Outcome: | The proposed tool is open source and available for download on the HFST framework. |
Similar Papers
Parsivar: A Language Processing Toolkit for Persian (L18-1)
Copied to clipboard
| Challenge: | a preprocessing step is required to convert text into a standard format for NLP tasks. |
| Approach: | They propose a Persian preprocessing toolkit that performs various kinds of activities . they use a plagiarism detection system to exploit the proposed toolkit . |
| Outcome: | The proposed tool outperforms available Persian preprocessing tools by about 8 percent in terms of F1 . the proposed toolkit performs normalization, space correction, tokenization, stemming, parts of speech tagging and shallow parsing tasks. |
lingvis.io - A Linguistic Visual Analytics Framework (P19-3)
Copied to clipboard
Mennatallah El-Assady, Wolfgang Jentner, Fabian Sperrle, Rita Sevastjanova, Annette Hautli-Janisz, Miriam Butt, Daniel Keim
| Challenge: | Using a modular framework, linguistic visual analytics applications can be rapidly prototypized using a web-based framework. |
| Approach: | They propose a modular framework for rapid prototyping of linguistic, web-based, visual analytics applications. |
| Outcome: | The proposed framework supports rapid prototyping of linguistic, web-based, visual analytics applications. |
Identification and Analysis of Personification in Hungarian: The PerSECorp project (2022.lrec-1)
Copied to clipboard
| Challenge: | despite recent findings on the conceptual and linguistic organization of personification, we have relatively little knowledge about its lexical patterns and grammatical templates. |
| Approach: | They propose a corpus-driven approach to personification analysis in cognitive linguistics . they use a semi-automatically processed corpus to annotate personifying linguistic structures . |
| Outcome: | The proposed method consists of annotating a semi-automatic corpus of car reviews in Hungarian . the corpus is structured and annotated manually, and gives an overview of possible data types . |
Texar: A Modularized, Versatile, and Extensible Toolkit for Text Generation (P19-3)
Copied to clipboard
Zhiting Hu, Haoran Shi, Bowen Tan, Wentao Wang, Zichao Yang, Tiancheng Zhao, Junxian He, Lianhui Qin, Di Wang, Xuezhe Ma, Zhengzhong Liu, Xiaodan Liang, Wanrong Zhu, Devendra Sachan, Eric Xing
| Challenge: | Texar is an open-source text generation toolkit that supports a broad set of text generation tasks. |
| Approach: | They introduce Texar, an open-source text generation toolkit that supports text generation tasks. |
| Outcome: | Texar supports machine translation, summarization, dialog, content manipulation, and more. |
ET: A Workstation for Querying, Editing and Evaluating Annotated Corpora (2021.emnlp-demo)
Copied to clipboard
| Challenge: | Using annotated corpora is costly for humans alone and requires a large amount of time and effort to manipulate. |
| Approach: | They propose to use annotated corpora as a tool for linguistic research and natural language processing. |
| Outcome: | The proposed work is based on two integrated environments – Interrogatório and Julgamento . the open-source environment is used in several linguistic and NLP-related studies . |
A Free/Open-Source Morphological Analyser and Generator for Sakha (2022.lrec-1)
Copied to clipboard
| Challenge: | a morphological transducer for Sakha is being developed for use in downstream tasks . the marginalised language is subject to increasing economic and cultural peril due to climate change . |
| Approach: | They describe the development of a morphological analyser and generator for Sakha . the transducer has coverage of solidly above 90%, and high precision . it is already being used in downstream tasks such as linguistic maintenance . |
| Outcome: | The proposed morphological analyser has coverage of 90% and high precision . it is already being used in computer assisted language learning applications . |
OpenHuEval: Evaluating Large Language Model on Hungarian Specifics (2025.findings-acl)
Copied to clipboard
Haote Yang, Xingjian Wei, Jiang Wu, Noémi Ligeti-Nagy, Jiaxing Sun, Yinfan Wang, Győző Zijian Yang, Junyuan Gao, Jingchao Wang, Bowen Jiang, Shasha Wang, Nanjun Yu, Zihao Zhang, Shixin Hong, Hongwei Liu, Wei Li, Songyang Zhang, Dahua Lin, Lijun Wu, Gábor Prószéky, Conghui He
| Challenge: | Recent advances in Large Language Models (LLMs) represent significant strides toward artificial general intelligence (AGI). |
| Approach: | They introduce OpenHuEval, the first benchmark for LLMs focusing on the Hungarian language and specifics. |
| Outcome: | The framework reveals intrinsic patterns and mechanisms of LLMs in non-English languages, with Hungarian serving as an example. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
Making Metadata Fit for Next Generation Language Technology Platforms: The Metadata Schema of the European Language Grid (2020.lrec-1)
Copied to clipboard
Penny Labropoulou, Katerina Gkirtzou, Maria Gavriilidou, Miltos Deligiannis, Dimitris Galanis, Stelios Piperidis, Georg Rehm, Maria Berger, Valérie Mapelli, Michael Rigault, Victoria Arranz, Khalid Choukri, Gerhard Backfried, José Manuel Gómez-Pérez, Andres Garcia-Silva
| Challenge: | Metadata are a key factor in the management, sharing and usage of digital assets . the European Language Grid project aims to be the primary hub and marketplace for industry-relevant Language Technology in Europe. |
| Approach: | They propose a rich metadata schema catering for the description of Language Resources and Technologies. |
| Outcome: | The proposed schema powers the European Language Grid platform that aims to be the primary hub and marketplace for industry-relevant Language Technology in Europe. |
Stanza: A Python Natural Language Processing Toolkit for Many Human Languages (2020.acl-demos)
Copied to clipboard
| Challenge: | Existing tools that support only a few major languages are under-optimized for accuracy due to a focus on efficiency or use of less powerful models. |
| Approach: | They introduce a Python natural language processing toolkit that supports 66 languages . they train Stanza on 112 datasets and show it generalizes well on all languages compared to other tools . |
| Outcome: | The proposed toolkit performs well on 112 datasets and is compatible with the popular Java CoreNLP software. |