Papers by Beatrice Daille
Books of Hours. the First Liturgical Data Set for Text Segmentation. (2020.lrec-1)
Copied to clipboard
Amir Hazem, Beatrice Daille, Christopher Kermorvant, Dominique Stutzmann, Marie-Laurence Bonhomme, Martin Maarand, Mélodie Boillet
| Challenge: | Until now, the book of hours has been scarcely studied because of its manuscript nature, its length and its complex content. |
| Approach: | They propose to use Handwritten Text Recognition to generate a corpus of Latin transcriptions of 300 books of hours generated by OCR for handwritten and not printed texts. |
| Outcome: | The proposed structure and state-of-the-art methods are compared with existing methods and are based on the results of a systematic evaluation of two books of hours. |
Cross-lingual and Cross-domain Transfer Learning for Automatic Term Extraction from Low Resource Data (2022.lrec-1)
Copied to clipboard
| Challenge: | Automatic Term Extraction (ATE) is a key component for domain knowledge understanding and can be used for further NLP applications. |
| Approach: | They propose to fine-tune pre-trained BERT models for automatic Term Extraction (ATE) using cross-lingual and cross-domain transfer learning to extract single and multi-word terms. |
| Outcome: | The proposed models can capture cross-domain and cross-lingual terminologically-marked contexts shared by terms, opening a new design-pattern for ATE. |
Towards Reliable Paper Contributions Annotation in the ACL Rolling Review (2026.findings-acl)
Copied to clipboard
| Challenge: | Identifying the types of contributions an article makes can help readers grasp its significance. |
| Approach: | They propose to use a typology to categorize articles by their contributions to improve review quality and fairness. |
| Outcome: | The ACL Rolling Review (ARR) introduced a typology requiring authors to specify their contributions to improve review quality and fairness. |
Hierarchical Text Segmentation for Medieval Manuscripts (2020.coling-main)
Copied to clipboard
| Challenge: | Until now, text segmentation methods have only addressed data sets lying within the scope of narrative and expository texts or user dialogues texts. |
| Approach: | They propose a bottom-up greedy approach that enhances the results . they argue that books of hours exhibit a complex hierarchical entangled structure . |
| Outcome: | The proposed bottom-up greedy approach significantly enhances the results. |