Papers by Kokil Jaidka

12 papers
Diachronic degradation of language models: Insights from social media (P18-2)

Copied to clipboard

Challenge: Existing studies have explored whether and how language models degrade over time, i.e. why they fail to work on contemporary language.
Approach: They investigate the accuracy of pre-trained language models for downstream tasks in machine learning and user profiling.
Outcome: The results show that it is possible to measure diachronic drifts within social media and within the span of a few years.
Offer a Different Perspective: Modeling the Belief Alignment of Arguments in Multi-party Debates (2022.emnlp-main)

Copied to clipboard

Challenge: Existing work on persuasion in online forums focuses on identifying debate winners and winning negotiation games.
Approach: They adopt a hierarchical generative Variational Autoencoder model to model winning arguments . they propose competing hypotheses about the nature of argumentation .
Outcome: The proposed model predicts winning arguments in reddit debates . it uses a hierarchical generative Variational Autoencoder to model argumentation .
Beyond Text: Leveraging Multi-Task Learning and Cognitive Appraisal Theory for Post-Purchase Intention Analysis (2024.findings-acl)

Copied to clipboard

Challenge: Recent studies have shown that user-level features can carry more task-related information than the text itself.
Approach: They evaluate multi-task learning frameworks grounded in Cognitive Appraisal Theory to predict user behavior as a function of users’ self-expression and psychological attributes.
Outcome: The proposed models improve on the language and traits of users, while lacking rich annotations of other attributes.
“Thinking” Fair and Slow: On the Efficacy of Structured Prompts for Debiasing Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Existing debiasing techniques are typically training-based or require access to the model’s internals and output distributions, so they are inaccessible to end-users looking to adapt LLM outputs for their particular needs.
Approach: They propose a system-based iterative framework that uses System 2 thinking processes to induce logical, reflective, and critical text generation with single, multi-step, instruction, and role-based variants.
Outcome: The proposed framework significantly improves over other frameworks demonstrating lower mean bias in the outputs with competitive performance on the downstream tasks.
I am PsyAM: Modeling Happiness with Cognitive Appraisal Dimensions (2023.findings-acl)

Copied to clipboard

Challenge: Emotions are an indicator of psychological states, such as happiness, which can be modelled through cognitive appraisal theory (CAT)
Approach: They propose to use adaptor modules in a sequential multi-task learning setup to generate high-dimensional feature representations of hedonic well-being (momentary happiness) they propose to apply feature fusion methods to model emotion in text .
Outcome: The proposed framework has cross-task validity and generalizability and is robust against traditional methods and BERT baselines.
Learning Through Dialogue: Engagement and Efficacy Matter More Than Explanations (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly used as conversational partners for learning, yet the interactional dynamics supporting users’ learning and engagement are understudied.
Approach: They analyze linguistic and interactional features from LLM and participant chats to identify the mechanisms and conditions under which LLM explanations shape changes in political knowledge and confidence.
Outcome: The results show that LLM explanations shape political knowledge and confidence . they also show that their effects are highly conditional and vary by political efficacy .
WikiTalkEdit: A Dataset for modeling Editors’ behaviors on Wikipedia (2021.naacl-main)

Copied to clipboard

Challenge: Using the WikiTalkEdit dataset, we show how positive emotion and the use of first-person pronouns predict a positive emotional change in a Wikipedia contributor.
Approach: They introduce and analyze WikiTalkEdit, a dataset of conversations and edit histories from Wikipedia, for research in online cooperation and conversation modeling.
Outcome: The proposed dataset supports the classic understanding of style matching, where positive emotion and the use of first-person pronouns predict a positive emotional change in a Wikipedia contributor.
Beyond Context to Cognitive Appraisal: Emotion Reasoning as a Theory of Mind Benchmark for Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Recent studies have shown that large language models (LLMs) reason about others' emotional states using contextual information, within a Theory-of-Mind framework.
Approach: They propose to use large language models to reason about others’ emotional states using contextual information within a Theory-of-Mind framework.
Outcome: The proposed models can reason about situations and appraisals, but are poor at associating situational outcomes and appraisal with specific emotions.
The PEACE-Reviews dataset: Modeling Cognitive Appraisals in Emotion Text Analysis (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent studies have delved into its significance, yet the interplay between various forms of cognitive appraisal and specific emotions, such as joy and anger, remains an area of exploration in consumption contexts.
Approach: They propose to construct a dataset to model the evaluations people make about their situations based on annotated autobiographical accounts of their emotional and appraisal experiences .
Outcome: The proposed model incorporates emotion, cognition, individual traits, and demographic data.
Developing A Multilabel Corpus for the Quality Assessment of Online Political Talk (2022.lrec-1)

Copied to clipboard

Challenge: a corpus of political tweets labeled for its deliberative characteristics is presented . the dataset offers a first step in building dictionaries to aid in the measurement of the Discourse Quality Index .
Approach: They present a Twitter Deliberative Politics dataset that measures the quality of political tweets . they propose to use machine learning to analyze tweets and to use it to build dictionaries .
Outcome: The proposed dataset is useful to linguists, political scientists, and social scientists . it offers a first step in building dictionaries for the quality assessment of political talk in english .
Identifying Locus of Control in Social Media Language (D18-1)

Copied to clipboard

Challenge: lexical features outperform syntactic features in expressing control in social media . authors communicate internal locus of control when they ascribe control to themselves .
Approach: They examine the role of syntax and semantics in expressing users’ sense of control in annotated Facebook posts.
Outcome: The proposed language outperforms syntactic features in identifying whether or not a user is in control of their circumstances.
Disentangling Codemixing in Chats: The NUS ABC Codemixed Corpus (2026.findings-acl)

Copied to clipboard

Challenge: Existing studies on code-mixing have not been able to model human interactions in context.
Approach: They propose to use a general-purpose code-mixing corpus to model human interactions and relationships in context while maintaining ethical standards.
Outcome: The proposed corpus includes over 355,641 messages spanning various code-mixing patterns, with a primary focus on English, Mandarin, and other languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations