Papers by Pavel Kral

4 papers
Czech Dataset for Complex Aspect-Based Sentiment Analysis Tasks (2024.lrec-main)

Copied to clipboard

Challenge: 3.1K reviews are manually annotated for aspect-based sentiment analysis (ABSA) ABSA is a fine-grained task that aims to identify the sentiment associated with each aspect or characteristic of a text.
Approach: They propose a new Czech dataset for aspect-based sentiment analysis . the new dataset is built upon the older Czech dataset . authors provide 24M reviews without annotations suitable for unsupervised learning .
Outcome: The proposed dataset is built upon the older dataset, but is specifically designed for more complex tasks.
COMICORDA: Dialogue Act Recognition in Comic Books (2024.lrec-main)

Copied to clipboard

Challenge: Existing work on dialogue act recognition from images is limited to speech balloon segmentation and optical character recognition.
Approach: They propose a novel DA recognition approach for comic books using speech balloon segmentation, optical character recognition and DA classification.
Outcome: The proposed method achieves 98% average precision for speech balloon segmentation and exceeds 70% accuracy for the DA recognition task.
LACA: Improving Cross-lingual Aspect-Based Sentiment Analysis with LLM Data Augmentation (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to cross-lingual aspect-based sentiment analysis depend on translation tools.
Approach: They propose a cross-lingual aspect-based sentiment analysis framework that leverages a large language model to generate pseudo-labelled data in target language.
Outcome: The proposed approach outperforms translation-based approaches in six languages and five backbone models.
Czech Historical Named Entity Corpus v 1.0 (2020.lrec-1)

Copied to clipboard

Challenge: a lack of annotated historical data for named entity recognition is an obstacle to research in this area.
Approach: They propose to create an annotated corpus for named entity recognition in historical documents . they define domain-specific named entity types and create an annotation manual .
Outcome: The proposed corpus is available for research and is available to download . it is the first annotated historical corpus for named entity recognition (NER)

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations