Papers by Sougata Saha
Meta-Cultural Competence: Climbing the Right Hill of Cultural Awareness (2025.naacl-long)
Copied to clipboard
| Challenge: | Recent studies show that Large Language Models are biased towards a Western and Anglo-centric worldview. |
| Approach: | They propose to extend the Octopus test to measure "cultural awareness" they argue that cultural awareness is needed for AI systems to be useful across cultures . |
| Outcome: | The proposed method argues that cultural awareness is not cultural knowledge, but meta-cultural competence . the proposed method is based on the octopus test, which shows it is impossible to learn meaning from real-world concepts without knowing intent and meaning . |
Reading between the Lines: Can LLMs Identify Cross-Cultural Communication Gaps? (2025.naacl-long)
Copied to clipboard
| Challenge: | a new study examines the extent and patterns of gaps in understandability of book reviews . 83% of the reviews had at least one culturally-specific difficult-to-understand element . |
| Approach: | They examine extent and patterns of gaps in understandability of book reviews . 83% of reviews had at least one culturally-specific difficult-to-understand element . authors say they have a significant scope for improvement . |
| Outcome: | The proposed approach improves the understanding of book reviews from different cultures . 83% of the reviews had at least one culturally-specific difficult-to-understand element . |
Using Multi-Encoder Fusion Strategies to Improve Personalized Response Selection (2022.coling-1)
Copied to clipboard
| Challenge: | Existing systems that focus on persona do not explore well the correlation between persona and empathy. |
| Approach: | They propose a suite of fusion strategies that capture interaction between persona, emotion, and entailment information of the utterances. |
| Outcome: | The proposed model outperforms the previous methods by 2.3% on original personas and 1.9% on revised persona models in terms of hits@1 accuracy. |
Dialo-AP: A Dependency Parsing Based Argument Parser for Dialogues (2022.coling-1)
Copied to clipboard
| Challenge: | a recent work on argument mining has focused on parsing monologues, while neglecting dialogues. |
| Approach: | They propose an end-to-end argument parser that constructs argument graphs from dialogues . they use extensive pre-training and curriculum learning to train AM . |
| Outcome: | The proposed system performs all sub-tasks of AM and achieves significant improvements . it is compared to existing systems and validated through human evaluation . |
User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs (2025.findings-acl)
Copied to clipboard
| Challenge: | We argue that knowledge-retrieval and reasoning tasks are not ideal for measuring generalization, as LLMs are not trained for specific tasks. |
| Approach: | They propose a statistically motivated framework using personalization to assess generalization in Large Language Models. |
| Outcome: | The proposed framework outperforms existing models on movie and music recommendation datasets, but all models have room for improvement, especially Llama. |
CULTURALLY YOURS: A Reading Assistant for Cross-Cultural Content (2025.coling-demos)
Copied to clipboard
| Challenge: | Culturally Yours (CY) is a cultural reading assistant that helps users from diverse cultural backgrounds understand content from online sources that are written by people from a different culture. |
| Approach: | They propose to use culturally sensitive language to personalize a cultural reading assistant tool that can identify cultural-specific items for users from varying cultural contexts. |
| Outcome: | The tool personalizes to the user’s preferences based on the interaction of the user with the tool. |
Women, Infamous, and Exotic Beings: A Comparative Study of Honorific Usages in Wikipedia and LLMs for Bengali and Hindi (2025.emnlp-main)
Copied to clipboard
| Challenge: | Honorifics encode nuanced socio-pragmatic cues such as power, age, gender, fame, and cultural distance. |
| Approach: | They propose to study third-person honorific usage across 10,000 Hindi and Bengali Wikipedia articles . honorifics are more prevalent in Bengali than in Hindi, while non-honorifics dominate . |
| Outcome: | The authors show that large language models internalize similar socio-pragmatic norms . their analysis shows that honorifics are more prevalent in Bengali than in Hindi . |
ArgU: A Controllable Factual Argument Generator (2023.acl-long)
Copied to clipboard
| Challenge: | Effective argumentation is essential towards a purposeful conversation with a satisfactory outcome. |
| Approach: | They propose a controllable neural argument generator capable of producing factual arguments from input facts and real-world concepts that can be explicitly controlled for stance and argument structure. |
| Outcome: | The proposed model produces factual arguments from input facts and real-world concepts that can be explicitly controlled for stance and argument structure using Walton’s argument scheme-based control codes. |
Diving Deep into Modes of Fact Hallucinations in Dialogue Systems (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Knowledge Graph(KG) grounded conversations often use large pre-trained models and suffer from fact hallucination. |
| Approach: | They propose to use a human feedback analysis to identify various modes of hallucination in KG chatbots. |
| Outcome: | The proposed system provides fine-grained signals that control fallacious content while generating responses. |
Integrating Argumentation and Hate-Speech-based Techniques for Countering Misinformation (2024.emnlp-main)
Copied to clipboard
| Challenge: | scalable strategies to combat online misinformation are short-term and insufficient, authors say . current reactive approaches, like content flagging and banning, do little to change perception of misinformants . human evaluations show that our framework generates expert-like responses . |
| Approach: | They propose a framework that generates persuasive responses from hate-speech counter-responses . human evaluations show that the framework generates expert-like responses . |
| Outcome: | The proposed framework generates expert-like responses and is 14% more engaging, 21% more natural, and 18% more factual than the best available alternatives. |