Papers by Nikita Mediankin
SumeCzech: Large Czech News-Based Summarization Dataset (L18-1)
Copied to clipboard
| Challenge: | Summarization of documents is a well-studied NLP task, but only a few datasets are available for Czech. |
| Approach: | They propose to use a Czech news-based summarization dataset to evaluate document summarizing . they propose a language-agnostic variant of the ROUGE metric to enable automatic evaluation . |
| Outcome: | The proposed dataset contains more than a million Czech news articles . the proposed approach is strong abstractive and language-agnostic . |