Papers by Souvika Sarkar
Zero-Shot Multi-Label Topic Inference with Sentence Encoders and LLMs (2023.emnlp-main)
Copied to clipboard
| Challenge: | In this paper, we focus on Zero-shot approaches for inferring topics from documents where both the document and topics were never seen by a model previously. |
| Approach: | They propose to use Sentence Encoders and Large Language Models to perform a "definition-wild zero-shot topic inference" where users define or provide topics of interest in real-time. |
| Outcome: | The proposed methods outperform ChatGPT-3.5 and PaLM and Sentence-BERT on the definition-wild zero-shot topic inference task on seven datasets. |
Exploring Universal Sentence Encoders for Zero-shot Text Classification (2022.aacl-short)
Copied to clipboard
| Challenge: | Universal Sentence Encoder (USE) has gained popularity as a general-purpose sentence encoding technique. |
| Approach: | They propose to use Universal Sentence Encoder (USE) to learn a general-purpose sentence encoding technique. |
| Outcome: | The proposed technique outperforms topic-based inference in zero-shot text classification tasks. |
Benchmarking LLMs on Semantic Overlap Summarization (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are the most capable text generation models in a variety of tasks and fields. |
| Approach: | They benchmark Large Language Models (LLMs) on SOS and introduce PrivacyPolicyPairs (3P) a dataset of 135 high-quality privacy policy documents is used to evaluate the model. |
| Outcome: | The proposed dataset complements existing resources and broadens domain coverage. |
LLMs as Meta-Reviewers’ Assistants: A Case Study (2025.naacl-long)
Copied to clipboard
Eftekhar Hossain, Sanjeev Kumar Sinha, Naman Bansal, R. Alexander Knipper, Souvika Sarkar, John Salvador, Yash Mahajan, Sri Ram Pavan Kumar Guttikonda, Mousumi Akter, Md. Mahadi Hassan, Matthew Freestone, Matthew C. Williams Jr., Dongji Feng, Santu Karmaker
| Challenge: | Meta-reviews are a critical step in the overall scientific peer-reviewed process, which focuses on understanding the consensus of expert opinions on a scholarly work and making informed judgments on its scientific merit. |
| Approach: | They propose to use large language models to generate a controlled multi-perspective-summary (MPS) of their opinions to help meta-reviewers better comprehend multiple experts' perspectives. |
| Outcome: | The proposed model can help meta-reviewers better comprehend multiple experts’ perspectives by generating a controlled multi-perspective-summary (MPS) of their opinions. |
On Evaluation of Bangla Word Analogies (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing word embeddings in Bangla struggle to perform well on low-resource data sets. |
| Approach: | They propose to use a benchmark dataset of Bangla word analogies to evaluate the quality of existing Bangla embeddings. |
| Outcome: | The proposed evaluation set includes 16,678 unique word analogies in Bangla and a translated and curated version of the original Mikolov dataset (10,594 samples) . |