Papers by Daniel Baumartz

4 papers
LTV: Labeled Topic Vector (C18-2)

Copied to clipboard

Challenge: Using nnDDC, we generate labeled topic classifications based on the Dewey Decimal Classification (DDC) Unlike related approaches, we use classifiers to define the dimensions of CISS, which are directly labeles by the underlying target class.
Approach: They propose a website and API that generates labeled topic classifications based on the Dewey Decimal Classification (DDC) they propose nnDDC, a largely language-independent natural network-based classifier for DDC, which is language-dependent .
Outcome: The proposed model is language-independent and performs well in 40 languages.
FastSense: An Efficient Word Sense Disambiguation Classifier (L18-1)

Copied to clipboard

Challenge: Word Sense Disambiguation (WSD) is a task that is often overlooked by NLP pipelines because of its complexity and complexity.
Approach: They propose a neural network-based tool for word sense disambiguation called fastSense.
Outcome: The proposed tool can process huge amounts of data quickly and surpasses state-of-the-art tools in terms of F-measure.
Unlocking the Heterogeneous Landscape of Big Data NLP with DUUI (2023.findings-emnlp)

Copied to clipboard

Challenge: Automated analysis of large corpora is a complex task, especially in terms of time efficiency.
Approach: They propose a framework for automatic distributed analysis of text corpora that leverages Big Data experience and virtualization with Docker.
Outcome: The proposed framework is scalable, flexible, lightweight, and feature-rich for automatic distributed analysis of text corpora.
Dependencies over Times and Tools (DoTT) (2024.lrec-main)

Copied to clipboard

Challenge: Using the examples of English and German, we examine how parsers trained on modern variants of these languages can be transferred to older language levels without loss.
Approach: They develop a treebank of diachronic corpora enriched with dependency annotations using 3 parsers, 6 pre-trained language models, 5 newly trained models for German, and two tag sets.
Outcome: The proposed treebank covers the time period from 1800 until today and is based on the DependencyAnnotator annotation tool.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations