Papers by Dominique Stutzmann
Books of Hours. the First Liturgical Data Set for Text Segmentation. (2020.lrec-1)
Copied to clipboard
Amir Hazem, Beatrice Daille, Christopher Kermorvant, Dominique Stutzmann, Marie-Laurence Bonhomme, Martin Maarand, Mélodie Boillet
| Challenge: | Until now, the book of hours has been scarcely studied because of its manuscript nature, its length and its complex content. |
| Approach: | They propose to use Handwritten Text Recognition to generate a corpus of Latin transcriptions of 300 books of hours generated by OCR for handwritten and not printed texts. |
| Outcome: | The proposed structure and state-of-the-art methods are compared with existing methods and are based on the results of a systematic evaluation of two books of hours. |
Hierarchical Text Segmentation for Medieval Manuscripts (2020.coling-main)
Copied to clipboard
| Challenge: | Until now, text segmentation methods have only addressed data sets lying within the scope of narrative and expository texts or user dialogues texts. |
| Approach: | They propose a bottom-up greedy approach that enhances the results . they argue that books of hours exhibit a complex hierarchical entangled structure . |
| Outcome: | The proposed bottom-up greedy approach significantly enhances the results. |