Papers by Beatrice Daille

4 papers
Books of Hours. the First Liturgical Data Set for Text Segmentation. (2020.lrec-1)

Copied to clipboard

Challenge: Until now, the book of hours has been scarcely studied because of its manuscript nature, its length and its complex content.
Approach: They propose to use Handwritten Text Recognition to generate a corpus of Latin transcriptions of 300 books of hours generated by OCR for handwritten and not printed texts.
Outcome: The proposed structure and state-of-the-art methods are compared with existing methods and are based on the results of a systematic evaluation of two books of hours.
Cross-lingual and Cross-domain Transfer Learning for Automatic Term Extraction from Low Resource Data (2022.lrec-1)

Copied to clipboard

Challenge: Automatic Term Extraction (ATE) is a key component for domain knowledge understanding and can be used for further NLP applications.
Approach: They propose to fine-tune pre-trained BERT models for automatic Term Extraction (ATE) using cross-lingual and cross-domain transfer learning to extract single and multi-word terms.
Outcome: The proposed models can capture cross-domain and cross-lingual terminologically-marked contexts shared by terms, opening a new design-pattern for ATE.
Towards Reliable Paper Contributions Annotation in the ACL Rolling Review (2026.findings-acl)

Copied to clipboard

Challenge: Identifying the types of contributions an article makes can help readers grasp its significance.
Approach: They propose to use a typology to categorize articles by their contributions to improve review quality and fairness.
Outcome: The ACL Rolling Review (ARR) introduced a typology requiring authors to specify their contributions to improve review quality and fairness.
Hierarchical Text Segmentation for Medieval Manuscripts (2020.coling-main)

Copied to clipboard

Challenge: Until now, text segmentation methods have only addressed data sets lying within the scope of narrative and expository texts or user dialogues texts.
Approach: They propose a bottom-up greedy approach that enhances the results . they argue that books of hours exhibit a complex hierarchical entangled structure .
Outcome: The proposed bottom-up greedy approach significantly enhances the results.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations