Challenge: »textklang« aims to explore the relationship between written text and its potential and actual sonic realisation in lyric poetry . the platform will combine three modalities: the poetic text, the audio signal of a recorded recitation and, at a later stage, music scores of . musical setting of lyrical poetry.
Approach: They propose to combine a multi-modal corpus of German lyric poetry from the Romantic era with a platform for systematic exploration.
Outcome: The platform will combine the poetic text, the audio signal of a recorded recitation and, at a later stage, music scores of . a musical setting of lyric poetry.

Similar Papers

A Corpus Linguistic Perspective on Contemporary German Pop Lyrics with the Multi-Layer Annotated “Songkorpus” (2020.lrec-1)

Copied to clipboard

Challenge: TEI-compliant song lyrics are used as primary data, linguistically and literary motivated annotations, and extralinguistic metadata.
Approach: They propose to annotate a multiply annotated corpus of German lyrics as a publicly available basis for multidisciplinary research.
Outcome: The proposed corpus of german lyrics is available for evaluation and analysis using TEI-compliant, linguistically and literary motivated annotations and extralinguistic metadata.
AET: Web-based Adjective Exploration Tool for German (L18-1)

Copied to clipboard

Challenge: AET enables research on the modificational behavior of German adjectives and adverbs . currently available online corpus query tools for German do not lend themselves specifically to research on adjectives - e.g., syntactic relationships or morphological properties.
Approach: They propose a web-based corpus query tool that can be used to query German corpus . they extracted modifiers and modifiees from a print media corpus and stored them in a database .
Outcome: The proposed tool can be transferred to other languages and modification phenomena.
Metrical Tagging in the Wild: Building and Annotating Poetry Corpora with Rhythmic Features (2021.eacl-main)

Copied to clipboard

Challenge: a prerequisite for the computational study of literature is the availability of properly digitized texts with reliable meta-data and ground-truth annotation.
Approach: They propose to annotate prosodic features in large poetry corpora for English and German and train corpus driven neural models that enable large scale analysis.
Outcome: The proposed models outperform baseline and BERT-based approaches in English and german and show that they learn foot boundaries better when jointly predicting syllable stress, aesthetic emotions and verse measures benefit from each other.
The D-WISE Tool Suite: Multi-Modal Machine-Learning-Powered Tools Supporting and Enhancing Digital Discourse Analysis (2023.acl-demo)

Copied to clipboard

Challenge: The D-WISE Tool Suite addresses limitations of current DH tools due to the ever-increasing amount of heterogeneous, unstructured, and multi-modal data in which discourses of contemporary societies are encoded.
Approach: They propose to use D-WISE Tool Suite to analyze heterogeneous, unstructured, and multi-modal data in the Digital Humanities (DH)
Outcome: The proposed tool leverages state-of-the-art machine learning technologies from Natural Language Processing and Com-puter Vision to ensure its usability for modernDH research.
FEIDEGGER: A Multi-modal Corpus of Fashion Images and Descriptions in German (L18-1)

Copied to clipboard

Challenge: Recent years have seen a renewed interest in text-image multi-modality . paired text-picture datasets are often limited to English language text .
Approach: They propose a multi-modal corpus that pairs images and textual descriptions of their content in German to enable study of these challenges.
Outcome: The proposed dataset focuses on the domain of fashion items and their visual descriptions in German.
Placing multi-modal, and multi-lingual Data in the Humanities Domain on the Map: the Mythotopia Geo-tagged Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Using mythology as a starting point, visitors of Northern Greece will have a multi-faceted experience using a corpus of textual data supplemented with images, and video.
Approach: They propose to integrate a multi-lingual corpus with a dedicated database with advanced indexing, linking and search functionalities into a platform for scholarly research in the digital humanities.
Outcome: The proposed infrastructure will be integrated into a platform aimed at providing a multi-faceted experience to visitors of Northern Greece using mythology as a starting point.
A Lightweight Modeling Middleware for Corpus Processing (L18-1)

Copied to clipboard

Challenge: Present-day empirical research in computational or theoretical linguistics has richly annotated and diverse corpus resources.
Approach: They propose a framework for modeling arbitrary multi-modal corpus resources in a unified form for processing tools.
Outcome: The proposed framework allows researchers to explore and query more diverse corpus resources and artifacts through a single interactive interface.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
Multimodality for NLP-Centered Applications: Resources, Advances and Frontiers (2022.lrec-1)

Copied to clipboard

Challenge: resurgence of multimodal datasets has attracted significant research interest, but there is no comprehensive survey for this task.
Approach: They present a survey of a multimodal dataset with different modalities according to the applications.
Outcome: The proposed datasets are available online and discuss the new frontier and motivate future researches.
Wikinflection Corpus: A (Better) Multilingual, Morpheme-Annotated Inflectional Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Inflectional corpora with annotated morpheme boundaries are scarce in the NLP community . a generated, multilingual inflectional lexicon with morphological features is not as good as UniMorph's .
Approach: They evaluate a multilingual inflectional corpus with morpheme boundaries from the English Wiktionary and the UniMorph project's inflection corpus.
Outcome: The generated Wikinflection corpus is not as good as UniMorph's, but extracts significant amount of words from the intersection of the two corpora.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations