Papers by Dominique Stutzmann

2 papers
Books of Hours. the First Liturgical Data Set for Text Segmentation. (2020.lrec-1)

Copied to clipboard

Challenge: Until now, the book of hours has been scarcely studied because of its manuscript nature, its length and its complex content.
Approach: They propose to use Handwritten Text Recognition to generate a corpus of Latin transcriptions of 300 books of hours generated by OCR for handwritten and not printed texts.
Outcome: The proposed structure and state-of-the-art methods are compared with existing methods and are based on the results of a systematic evaluation of two books of hours.
Hierarchical Text Segmentation for Medieval Manuscripts (2020.coling-main)

Copied to clipboard

Challenge: Until now, text segmentation methods have only addressed data sets lying within the scope of narrative and expository texts or user dialogues texts.
Approach: They propose a bottom-up greedy approach that enhances the results . they argue that books of hours exhibit a complex hierarchical entangled structure .
Outcome: The proposed bottom-up greedy approach significantly enhances the results.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations