Papers by Atharva Kulkarni

9 papers
SynthDST: Synthetic Data is All You Need for Few-Shot Dialog State Tracking (2024.eacl-long)

Copied to clipboard

Challenge: In-context learning with Large Language Models (LLMs) is a promising avenue of research in Dialog State Tracking (DST).
Approach: They propose a data generation framework tailored for Dialog State Tracking that uses large language models to synthesize natural, coherent, and free-flowing dialogues with DST annotations.
Outcome: The proposed framework improves joint goal accuracy by 4-5% over the zero-shot baseline on MultiWOZ 2.1 and 2.4.
Characterizing the Entities in Harmful Memes: Who is the Hero, the Villain, the Victim? (2023.eacl-main)

Copied to clipboard

Challenge: A common problem associated with meme comprehension lies in detecting the entities referenced and characterizing the role of each of these entities.
Approach: They propose to use a memes dataset on US Politics and Covid-19 memes to characterize the role of harmful entities in memes.
Outcome: The proposed model improves 4% over baseline and 1% over competing models.
When did you become so smart, oh wise one?! Sarcasm Explanation in Multi-modal Multi-party Dialogues (2022.acl-long)

Copied to clipboard

Challenge: Indirect speech achieves a constellation of discourse goals in human communication, but it is challenging for AI agents to comprehend such idiosyncrasies.
Approach: They propose a task to generate natural language explanations of satirical conversations using a multimodal and code-mixed dataset to capture multimodality.
Outcome: The proposed task generates natural language explanations of satirical conversations in a multimodal and code-mixed setting and surpasses baselines on almost all metrics.
Evaluating Evaluation Metrics – The Mirage of Hallucination Detection (2025.findings-emnlp)

Copied to clipboard

Challenge: a large-scale empirical evaluation of hallucination detection metrics is conducted . hallucinosity is a significant obstacle to the reliability and widespread adoption of language models .
Approach: They conduct large-scale empirical evaluation of hallucination detection metrics . they compare hallucinian language models, language models and decoding methods .
Outcome: The results show that the evaluations of hallucination detection metrics fail to align with human judgments, they say . they also show that evaluations with LLM-based evaluation yield the best overall results .
The student becomes the master: Outperforming GPT3 on Scientific Factual Error Correction (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for Factual Claim Correction rely on a verification model to guide the correction process.
Approach: They propose a claim correction system that does not require a verifier but outperforms existing methods by a considerable margin.
Outcome: The proposed system outperforms existing methods by a considerable margin on the SciFact dataset, 77% on SciFACT-Open and 72.75% on the CovidFact data set.
Counting the Bugs in ChatGPT’s Wugs: A Multilingual Investigation into the Morphological Capabilities of a Large Language Model (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies on large language models (LLMs) ignore the remarkable ability of humans to generalize and focus only on English.
Approach: They conduct the first rigorous analysis of the morphological capabilities of ChatGPT in four typologically varied languages.
Outcome: The proposed model massively underperforms purpose-built systems, particularly in English.
Empowering the Fact-checkers! Automatic Identification of Claim Spans on Twitter (2022.emnlp-main)

Copied to clipboard

Challenge: Current vogue is to employ manual fact-checkers to efficiently classify and verify such data to combat this avalanche of misinformation and fake news.
Approach: They propose a large-scale Twitter corpus with token-level claim spans on more than 7.5k tweets and a model that automatically detects and extracts the snippets of misinformation.
Outcome: The proposed model outperforms baseline systems on several evaluation metrics, improving by 1.5 points.
Still Not Quite There! Evaluating Large Language Models for Comorbid Mental Health Diagnosis (2024.emnlp-main)

Copied to clipboard

Challenge: ANGST is a benchmark for depression-anxiety comorbidity classification from social media posts.
Approach: They propose a social media-based benchmark for depression-anxiety comorbidity classification . ANGST enables multi-label classification, allowing each post to be simultaneously identified as indicating depression and/or anxiety.
Outcome: The proposed dataset enables multi-label classification of depression and anxiety . it outperforms existing models but none achieves an F1 score exceeding 72% .
Music for All: Representational Bias and Cross-Cultural Adaptability of Music Generation Models (2025.findings-naacl)

Copied to clipboard

Challenge: Existing music generation models are limited in their coverage of the musical genres and cultures of the world.
Approach: They propose to use parametric fine tuning techniques to mitigat the bias in existing music datasets.
Outcome: The proposed models are able to perform well across genres and cultures.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations