Papers by Shantipriya Parida

6 papers
POLAR: A Benchmark for Multilingual, Multicultural, and Multi-Event Online Polarization (2026.findings-acl)

Copied to clipboard

Challenge: polarization is a pervasive threat to democratic institutions, civil discourse, and social cohesion worldwide . most existing datasets focus on English or high-resource languages, reflecting a widespread trend across NLP tasks .
Approach: They propose a multilingual, multicultural, and multi-event dataset with over 110K instances in 22 languages drawn from diverse online platforms and real-world events.
Outcome: The proposed dataset analyzes polarization detection, type, and manifestation using a variety of annotation platforms adapted to each cultural context.
HaVQA: A Dataset for Visual Question Answering and Multimodal Research in Hausa Language (2023.findings-acl)

Copied to clipboard

Challenge: Existing models for visual question answering are limited to the English language.
Approach: They present a multimodal dataset for visual question answering tasks in the Hausa language.
Outcome: The proposed dataset provides 12,044 gold standard English-Hausa parallel sentences that are semantically identical to the corresponding visual information.
Hausa Visual Genome: A Dataset for Multi-Modal English to Hausa Machine Translation (2022.lrec-1)

Copied to clipboard

Challenge: Hausa is considered a low resource language in natural language processing due to lack of resources.
Approach: They propose a dataset that contains the description of an image in Hausa and its equivalent in English.
Outcome: The Hausa Visual Genome is the first dataset of its kind . it can be used for Hausa-English machine translation, multi-modal research, image description .
Overview of the 6th Workshop on Asian Translation (D19-52)

Copied to clipboard

Challenge: The 6th workshop on Asian translation (WAT2019) was held in hong kong, hongkong, and hong kong.
Approach: They present the results of the shared tasks from the 6th workshop on Asian translation (WAT2019) 25 teams participated in the shared task and 10 research paper submissions were accepted .
Outcome: The results of the 6th workshop on Asian translation (WAT2019) include JaEn, JaZh scientific paper translation subtasks, Ja'En, ja'Ko, Ja’En patent translation sub tasks, Hi'En and My'En patent subtask and Ru'Ja news commentary translation task.
Abstract Text Summarization: A Low Resource Challenge (D19-1)

Copied to clipboard

Challenge: Existing datasets for multilingual text summarization are difficult to construct and lack of human knowledge and language processing abilities in computers makes text summaries a challenging task.
Approach: They propose an iterative data augmentation approach which uses synthetic data along with the real summarization data for the German language.
Outcome: The proposed system improves on the development and test sets on the German language text using the state-of-the-art “Transformer” model.
Idiap NMT System for WAT 2019 Multimodal Translation Task (D19-52)

Copied to clipboard

Challenge: In the past few decades, multi-modality has received critical attention in translation studies, although the benefit of visual modality in machine translation is still in debate.
Approach: They propose to use the Transformer model and IITB English-Hindi parallel corpus as additional data sources for the evaluation and challenge test sets.
Outcome: The proposed system outperforms systems that consider visual information in the English-Hindi Multi-Modal Translation task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations