Papers by Niyati Chhaya

17 papers
Counterfactuals to Control Latent Disentangled Text Representations for Style Transfer (2021.acl-short)

Copied to clipboard

Challenge: Existing methods for unsupervised text style transfer focus on transferring a specific attribute, but this technique has never been explored in natural language generation tasks.
Approach: They propose a counterfactual-based method to modify latent representations by posing a ‘what-if’ scenario.
Outcome: The proposed method is tested on multiple attribute transfer tasks like Sentiment, Formality and Excitement to support the hypothesis.
AUTOSUMM: Automatic Model Creation for Text Summarization (2021.emnlp-main)

Copied to clipboard

Challenge: Recent efforts to develop deep learning models for text generation tasks are challenging for non-experts.
Approach: They propose methods to automatically create deep learning models for extractive and abstractive summarization tasks using large language models.
Outcome: The proposed methods achieve near state-of-the-art performance on a range of datasets.
Diachronic degradation of language models: Insights from social media (P18-2)

Copied to clipboard

Challenge: Existing studies have explored whether and how language models degrade over time, i.e. why they fail to work on contemporary language.
Approach: They investigate the accuracy of pre-trained language models for downstream tasks in machine learning and user profiling.
Outcome: The results show that it is possible to measure diachronic drifts within social media and within the span of a few years.
IndicIRSuite: Multilingual Dataset and Neural Information Models for Indian Languages (2024.acl-short)

Copied to clipboard

Challenge: IndicIRSuite is the first attempt at building large-scale Neural Information Retrieval resources for a large number of Indian languages.
Approach: They introduce Neural Information Retrieval resources for 11 widely spoken Indian Languages from two major Indian language families.
Outcome: Experiments show that Indic-ColBERT improves on INDIC-MARCO datasets for 11 languages, and that it can be used to improve IR for Indian languages.
Offer a Different Perspective: Modeling the Belief Alignment of Arguments in Multi-party Debates (2022.emnlp-main)

Copied to clipboard

Challenge: Existing work on persuasion in online forums focuses on identifying debate winners and winning negotiation games.
Approach: They adopt a hierarchical generative Variational Autoencoder model to model winning arguments . they propose competing hypotheses about the nature of argumentation .
Outcome: The proposed model predicts winning arguments in reddit debates . it uses a hierarchical generative Variational Autoencoder to model argumentation .
Leveraging Mental Health Forums for User-level Depression Detection on Social Media (2022.lrec-1)

Copied to clipboard

Challenge: Existing methods to detect depression on social media platforms are limited due to the vastness of social media content and the lack of linguistic features.
Approach: They propose to optimize the performance of user-level depression classification to lessen the burden on computational resources.
Outcome: The proposed system outperforms baselines across standard metrics for the task of depression detection in text.
“Let’s not Quote out of Context”: Unified Vision-Language Pretraining for Context Assisted Image Captioning (2023.acl-industry)

Copied to clipboard

Challenge: Large enterprises have several teams to create their content for the purpose of marketing, campaigning, or even maintaining a brand presence.
Approach: They propose a new unified Vision-Language (VL) model with a focus on context-assisted image captioning where the caption is generated based on both the image and its context.
Outcome: The proposed model achieves state-of-the-art with an improvement of up to 8.34 CIDEr score on the benchmark news image captioning datasets.
Multi-label Categorization of Accounts of Sexism using a Neural Framework (D19-1)

Copied to clipboard

Challenge: Sexism manifests in blatant as well as subtle ways, authors say . existing work on sexism classification has limitations in terms of categories used . authors: categorization of accounts of sexist behavior can aid in countering sextism .
Approach: They propose a neural solution that can combine sentence representations with distributional and linguistic word embeddings.
Outcome: a new method outperforms deep learning and traditional methods by an appreciable margin . the proposed method outpersforms several deep learning as well as traditional baselines by an approval margin compared to baselines .
WikiTalkEdit: A Dataset for modeling Editors’ behaviors on Wikipedia (2021.naacl-main)

Copied to clipboard

Challenge: Using the WikiTalkEdit dataset, we show how positive emotion and the use of first-person pronouns predict a positive emotional change in a Wikipedia contributor.
Approach: They introduce and analyze WikiTalkEdit, a dataset of conversations and edit histories from Wikipedia, for research in online cooperation and conversation modeling.
Outcome: The proposed dataset supports the classic understanding of style matching, where positive emotion and the use of first-person pronouns predict a positive emotional change in a Wikipedia contributor.
Open-World Factually Consistent Question Generation (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods for question generation suffer from factual inconsistencies and incorrect entities and are not answerable from the input paragraph.
Approach: They propose a data processing technique based on de-lexicalization for consistent question generation across domains and a model that is generic across question-generation models.
Outcome: The proposed method produces entity-level factually consistent questions without significant impact on traditional metrics.
EmpathBERT: A BERT-based Framework for Demographic-aware Empathy Prediction (2021.eacl-main)

Copied to clipboard

Challenge: EmpathBERT is a demographic-aware framework for empathy prediction based on BERT.
Approach: They propose a demographic-aware framework for empathy prediction based on BERT and utilize user demographics to analyze user responses to stimulative news articles.
Outcome: The proposed framework surpasses machine learning and deep learning models and highlights the importance of demographic information in the responses.
Aff2Vec: Affect–Enriched Distributional Word Representations (C18-1)

Copied to clipboard

Challenge: Affective word distributions are not well understood in literature.
Approach: They propose a model that embeds affective word interpretations into enriched word embeddings.
Outcome: The proposed model outperforms the state-of-the-art in word-similarity tasks and in emotion analysis, personality detection, and frustration prediction tasks.
He is very intelligent, she is very beautiful? On Mitigating Social Biases in Language Modelling and Generation (2021.findings-acl)

Copied to clipboard

Challenge: Existing studies have focused on mitigating social biases in context-free representations, with recent shift to contextual ones.
Approach: They propose an approach to mitigate social biases in a large pre-trained contextual language model . they propose lexical co-occurrence-based bias penalization in the decoder units .
Outcome: The proposed approach reduces biases in fill-in-the-blank sentences and summarizes . it also reduces the biased representations in the frameworks, the authors show .
A Neural CRF-based Hierarchical Approach for Linear Text Segmentation (2023.findings-eacl)

Copied to clipboard

Challenge: Existing methods to segment unformatted text and transcripts explicitly train to predict segment boundaries, but they fail to provide a large annotated dataset.
Approach: They propose a method to generate hierarchical segmentation structures based on Wikipedia annotations by using a neural conditional random field.
Outcome: The proposed method outperforms or achieves competitive performance when compared to previous state-of-the-art algorithms.
DialogueGCN: A Graph Convolutional Neural Network for Emotion Recognition in Conversation (D19-1)

Copied to clipboard

Challenge: Emotion recognition in conversation (ERC) has received much attention lately due to its potential widespread applications in diverse areas, such as health-care, education, and human resources.
Approach: They propose a graph neural network-based approach to emotion recognition in conversation that leverages self and inter-speaker dependency of the interlocutors to model conversational context.
Outcome: The proposed method outperforms the current state-of-the-art on a number of benchmark emotion classification datasets while minimizing context propagation issues.
Semi-supervised Multi-task Learning for Multi-label Fine-grained Sexism Classification (2020.coling-main)

Copied to clipboard

Challenge: Sexism is a form of oppression based on one's sex and is reported online in numerous ways.
Approach: They propose a multi-task approach for fine-grained multi-label sexism classification that leverages several supporting tasks without incurring manual labeling cost.
Outcome: The proposed method outperforms the state-of-the-art for multi-label sexism classification on a recently released dataset across five standard metrics.
CaM-Gen: Causally Aware Metric-Guided Text Generation (2022.findings-acl)

Copied to clipboard

Challenge: Content is created for a well-defined purpose, often described by a metric or signal . external metrics and content tend to have inherent relationships and not all of them may be of consequence.
Approach: They propose a mechanism to guide generative models by user-defined target metrics . authors propose generative networks guided by causally significant aspects of text .
Outcome: The proposed models beat baselines in terms of the target metric control while maintaining fluency and language quality of the generated text.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations