Papers by Shantipriya Parida
POLAR: A Benchmark for Multilingual, Multicultural, and Multi-Event Online Polarization (2026.findings-acl)
Copied to clipboard
Usman Naseem, Robert Geislinger, Juan Ren, Sarah Kohail, Rudy Alexandro Garrido Veliz, P Sam Sahil, Yiran Zhang, Idris Abdulmumin, Marco Antonio Stranisci, Özge Alacam, Cengiz Acarturk, Aisha Jabr, Saba Anwar, Abinew Ali Ayele, Simona Frenda, Alessandra Teresa Cignarella, Elena Tutubalina, Oleg Rogov, Aung Kyaw Htet, Xintong Wang, Surendrabikram Thapa, Kritesh Rauniyar, Tanmoy Chakraborty, MD Arfeen Zeeshan, Dheeraj Kodati, Satya Keerthi, Sahar Moradizeyveh, Firoj Alam, Md Arid Hasan, Syed Ishtiaque Ahmed, Ye Kyaw Thu, Shantipriya Parida, Ihsan Ayyub Qazi, Lilian Diana Awuor Wanzare, Nelson Odhiambo Onyango, Clemencia Siro, Jane Wanjiru Kimani, Ibrahim Said Ahmad, Adem Chanie Ali, Martin Semmann, Chris Biemann, Shamsuddeen Hassan Muhammad, Seid Muhie Yimam
| Challenge: | polarization is a pervasive threat to democratic institutions, civil discourse, and social cohesion worldwide . most existing datasets focus on English or high-resource languages, reflecting a widespread trend across NLP tasks . |
| Approach: | They propose a multilingual, multicultural, and multi-event dataset with over 110K instances in 22 languages drawn from diverse online platforms and real-world events. |
| Outcome: | The proposed dataset analyzes polarization detection, type, and manifestation using a variety of annotation platforms adapted to each cultural context. |
HaVQA: A Dataset for Visual Question Answering and Multimodal Research in Hausa Language (2023.findings-acl)
Copied to clipboard
Shantipriya Parida, Idris Abdulmumin, Shamsuddeen Hassan Muhammad, Aneesh Bose, Guneet Singh Kohli, Ibrahim Said Ahmad, Ketan Kotwal, Sayan Deb Sarkar, Ondřej Bojar, Habeebah Kakudi
| Challenge: | Existing models for visual question answering are limited to the English language. |
| Approach: | They present a multimodal dataset for visual question answering tasks in the Hausa language. |
| Outcome: | The proposed dataset provides 12,044 gold standard English-Hausa parallel sentences that are semantically identical to the corresponding visual information. |
Hausa Visual Genome: A Dataset for Multi-Modal English to Hausa Machine Translation (2022.lrec-1)
Copied to clipboard
Idris Abdulmumin, Satya Ranjan Dash, Musa Abdullahi Dawud, Shantipriya Parida, Shamsuddeen Muhammad, Ibrahim Sa’id Ahmad, Subhadarshi Panda, Ondřej Bojar, Bashir Shehu Galadanci, Bello Shehu Bello
| Challenge: | Hausa is considered a low resource language in natural language processing due to lack of resources. |
| Approach: | They propose a dataset that contains the description of an image in Hausa and its equivalent in English. |
| Outcome: | The Hausa Visual Genome is the first dataset of its kind . it can be used for Hausa-English machine translation, multi-modal research, image description . |
Overview of the 6th Workshop on Asian Translation (D19-52)
Copied to clipboard
Toshiaki Nakazawa, Nobushige Doi, Shohei Higashiyama, Chenchen Ding, Raj Dabre, Hideya Mino, Isao Goto, Win Pa Pa, Anoop Kunchukuttan, Yusuke Oda, Shantipriya Parida, Ondřej Bojar, Sadao Kurohashi
| Challenge: | The 6th workshop on Asian translation (WAT2019) was held in hong kong, hongkong, and hong kong. |
| Approach: | They present the results of the shared tasks from the 6th workshop on Asian translation (WAT2019) 25 teams participated in the shared task and 10 research paper submissions were accepted . |
| Outcome: | The results of the 6th workshop on Asian translation (WAT2019) include JaEn, JaZh scientific paper translation subtasks, Ja'En, ja'Ko, Ja’En patent translation sub tasks, Hi'En and My'En patent subtask and Ru'Ja news commentary translation task. |
Abstract Text Summarization: A Low Resource Challenge (D19-1)
Copied to clipboard
| Challenge: | Existing datasets for multilingual text summarization are difficult to construct and lack of human knowledge and language processing abilities in computers makes text summaries a challenging task. |
| Approach: | They propose an iterative data augmentation approach which uses synthetic data along with the real summarization data for the German language. |
| Outcome: | The proposed system improves on the development and test sets on the German language text using the state-of-the-art “Transformer” model. |
Idiap NMT System for WAT 2019 Multimodal Translation Task (D19-52)
Copied to clipboard
| Challenge: | In the past few decades, multi-modality has received critical attention in translation studies, although the benefit of visual modality in machine translation is still in debate. |
| Approach: | They propose to use the Transformer model and IITB English-Hindi parallel corpus as additional data sources for the evaluation and challenge test sets. |
| Outcome: | The proposed system outperforms systems that consider visual information in the English-Hindi Multi-Modal Translation task. |