Papers by Giuseppe Abrami

8 papers
I still have Time(s): Extending HeidelTime for German Texts (2022.lrec-1)

Copied to clipboard

Challenge: HeidelTime is a widely used tool for detecting temporal expressions in texts.
Approach: They propose to extend HeidelTime's pattern matching system by observing false negatives within real world texts and various time banks.
Outcome: The proposed extension can detect expressions in texts and time banks in a convenient way.
German Parliamentary Corpus (GerParCor) (2022.lrec-1)

Copied to clipboard

Challenge: German parliaments have a large and partly unexploited treasure trove of publicly accessible texts.
Approach: a new corpus of German-language parliamentary protocols is made available in XMI format . the corpus is genre-specific and contains conversions of scanned protocols . a researcher at the university of berlin and a professor at the berlin university created the corpuus .
Outcome: the German Parliamentary Corpus is a genre-specific corpus of German-language parliamentary protocols from three centuries and four countries.
German Parliamentary Corpus (GerParCor) Reloaded (2024.lrec-main)

Copied to clipboard

Challenge: In 2022, the largest German-speaking corpus of parliamentary protocols from three different centuries has been published - GerParCor.
Approach: They propose to update the largest German-speaking corpus of parliamentary protocols from three different centuries, on a national and federal level, from Germany, Austria, Switzerland and Liechtenstein, and to make them available in XMI format.
Outcome: The updated corpus includes all new parliamentary protocols and adds and preprocesses further parliamentary protocol not covered in the previous version.
A UIMA Database Interface for Managing NLP-related Text Annotations (L18-1)

Copied to clipboard

Challenge: despite the use of UIMA as a document-based schema, it does not provide native database support.
Approach: They develop a database interface to allow generic use of UIMA documents in database systems.
Outcome: The framework is evaluated in relation to file system-based storage and provides data protection.
TextAnnotator: A UIMA Based Tool for the Simultaneous and Collaborative Annotation of Texts (2020.lrec-1)

Copied to clipboard

Challenge: Existing annotation tools are not efficient for the annotation of corpora and are not error-free.
Approach: They propose to extend existing annotation tools by evaluating their flexibility and efficiency.
Outcome: The proposed system performs platform-independent multimodal annotations and annotates complex textual structures.
Unlocking the Heterogeneous Landscape of Big Data NLP with DUUI (2023.findings-emnlp)

Copied to clipboard

Challenge: Automated analysis of large corpora is a complex task, especially in terms of time efficiency.
Approach: They propose a framework for automatic distributed analysis of text corpora that leverages Big Data experience and virtualization with Docker.
Outcome: The proposed framework is scalable, flexible, lightweight, and feature-rich for automatic distributed analysis of text corpora.
Dependencies over Times and Tools (DoTT) (2024.lrec-main)

Copied to clipboard

Challenge: Using the examples of English and German, we examine how parsers trained on modern variants of these languages can be transferred to older language levels without loss.
Approach: They develop a treebank of diachronic corpora enriched with dependency annotations using 3 parsers, 6 pre-trained language models, 5 newly trained models for German, and two tag sets.
Outcome: The proposed treebank covers the time period from 1800 until today and is based on the DependencyAnnotator annotation tool.
TreeAnnotator: Versatile Visual Annotation of Hierarchical Text Relations (L18-1)

Copied to clipboard

Challenge: TREEANNOTATOR is a browser-based tool for annotating tree-like structures . it provides a wider range of formats and provides graphical annotations .
Approach: They evaluate TREEANNOTATOR, a browser-based tool for annotating tree-like structures, in particular structures that jointly map dependency relations and inclusion hierarchies, as used by Rhetorical Structure Theory.
Outcome: The GUI interface is user-friendly and provides two visualization modes.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations