Papers by Ladislav Lenc
Czech Text Document Corpus v 2.0 (L18-1)
Copied to clipboard
| Challenge: | a corpus of text documents for automatic document classification in Czech is presented . paper aims to facilitate a straightforward comparison of document classification approaches on Czech data . |
| Approach: | This paper introduces a collection of text documents for automatic document classification in Czech language. |
| Outcome: | The proposed corpus is based on the Czech news agency's real newspaper articles . it is used for evaluation of multi-label document classification approaches . |
COMICORDA: Dialogue Act Recognition in Comic Books (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing work on dialogue act recognition from images is limited to speech balloon segmentation and optical character recognition. |
| Approach: | They propose a novel DA recognition approach for comic books using speech balloon segmentation, optical character recognition and DA classification. |
| Outcome: | The proposed method achieves 98% average precision for speech balloon segmentation and exceeds 70% accuracy for the DA recognition task. |