Papers by Dwaipayan Roy
A Topic-Aligned Multilingual Corpus of Wikipedia Articles for Studying Information Asymmetry in Low Resource Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | a systematic comparison of information coverage in English Wikipedia and Wikipedias in eight other widely spoken languages is needed to bridge the information gap. |
| Approach: | They compare information coverage in English Wikipedia and Wikipedias in eight other widely spoken languages. |
| Outcome: | The analysis quantifies and provides useful insights about the information gap that exists between different language editions of Wikipedia and offers a roadmap for the IR community to bridge this gap. |
MAKED: Multi-lingual Automatic Keyword Extraction Dataset (2022.lrec-1)
Copied to clipboard
| Challenge: | a large dataset of news articles spanning 20 languages is lacking for keyword extraction. |
| Approach: | They propose a large-scale multi-lingual keyword extraction dataset for 11 of 20 languages . authors believe it will help advance the field of automatic keyword extraction . |
| Outcome: | The proposed dataset is the first for 11 of 20 languages and is based on 540K+ news articles from the BBC News network. |