Papers by Krishnapriya Vishnubhotla
Improving Automatic Quotation Attribution in Literary Novels (2023.acl-short)
Copied to clipboard
| Challenge: | Existing methods for quotation attribution in literary novels require varying levels of available information. |
| Approach: | They propose to train and evaluate models for character identification, coreference resolution, quotation identification and speaker attribution tasks using an annotated dataset. |
| Outcome: | The proposed model scores on speaker attribution task on the same scale as state-of-the-art models. |
An Evaluation of Disentangled Representation Learning for Texts (2021.findings-acl)
Copied to clipboard
| Challenge: | Disentangled representations of texts encode information pertaining to different aspects of the text in separate vector embeddings. |
| Approach: | They propose to use a highly-structured natural language dataset to evaluate disentangled representations for texts. |
| Outcome: | The proposed models are well-suited for learning disentangled representations of texts on a synthetic natural language dataset. |
SemRel2024: A Collection of Semantic Textual Relatedness Datasets for 13 Languages (2024.findings-acl)
Copied to clipboard
Nedjma Ousidhoum, Shamsuddeen Muhammad, Mohamed Abdalla, Idris Abdulmumin, Ibrahim Ahmad, Sanchit Ahuja, Alham Aji, Vladimir Araujo, Abinew Ayele, Pavan Baswani, Meriem Beloucif, Chris Biemann, Sofia Bourhim, Christine Kock, Genet Dekebo, Oumaima Hourrane, Gopichand Kanumolu, Lokesh Madasu, Samuel Rutunda, Manish Shrivastava, Thamar Solorio, Nirmal Surange, Hailegnaw Tilaye, Krishnapriya Vishnubhotla, Genta Winata, Seid Yimam, Saif Mohammad
| Challenge: | SemRel datasets are annotated by native speakers across 13 languages . they are used to characterise the relationship between two units of text . |
| Approach: | They propose to use a semantic relatedness dataset to measure the degree of semantic textual relatedness between sentences in Afrikaans, Algerian Arabic, Amharic, English, Hausa, Hindi, Indonesian, Kinyarwanda, Marathi, Moroccan Arabic, Modern Standard Arabic, Spanish, and Telugu. |
| Outcome: | The proposed datasets are annotated by native speakers across 13 languages and represent the semantic relatedness of 13 languages. |
Emotion Granularity from Text: An Aggregate-Level Indicator of Mental Health (2024.emnlp-main)
Copied to clipboard
| Challenge: | Emotions play a central role in how we construct meaning and communicate with others. |
| Approach: | They propose to use temporally-ordered speaker utterances to measure emotion granularity in social media to determine whether they are effective as mental health markers. |
| Outcome: | The proposed measures of emotion granularity function as markers of mental health conditions. |
The Project Dialogism Novel Corpus: A Dataset for Quotation Attribution in Literary Texts (2022.lrec-1)
Copied to clipboard
| Challenge: | Annotated quotations are used to model attribution and coreference of literary texts . authors present project Dialogism Novel Corpus, or PDNC, for 22 novels . |
| Approach: | They present an annotated dataset of quotations for English literary texts . they use natural language processing to model aspects of narrative, events, and characters . |
| Outcome: | The project Dialogism Novel Corpus contains annotations for 35,978 quotations across 22 novels . authors show that NLP can be used to model aspects of narrative, events, and characters . |
The Emotion Dynamics of Literary Novels (2024.findings-acl)
Copied to clipboard
| Challenge: | a new study examines the emotional journeys of characters in novels . previous studies have considered a novel as representing a single story arc . |
| Approach: | They analyze the emotion arcs of English literary novels using Utterance Emotion Dynamics . they find that narration and dialogue largely express disparate emotions through the course of a novel . |
| Outcome: | The analysis of English literary novels shows that narration and dialogue express disparate emotions . the commonalities or differences in the emotional arcs are more accurately captured by individual characters . |
What Makes Sentences Semantically Related? A Textual Relatedness Dataset and Empirical Study (2023.eacl-main)
Copied to clipboard
| Challenge: | Existing work on semantic relatedness has focused on semantic similarity because of a lack of relatedness datasets. |
| Approach: | They propose a dataset for semantic relatedness that has 5,500 English sentence pairs manually annotated using a comparative annotation framework. |
| Outcome: | The proposed dataset has 5,500 English sentence pairs manually annotated using a comparative annotation framework. |
Tweet Emotion Dynamics: Emotion Word Usage in Tweets from US and Canada (2022.lrec-1)
Copied to clipboard
| Challenge: | a dataset of 45 million geo-located tweets from the US and Canada is used to analyze emotions . early work identified tweets as a crucial indicator of public sentiment . |
| Approach: | They propose a dataset of more than 45 million geo-located tweets from US and Canada . they also introduce Tweet Emotion Dynamics (TED) metrics to capture patterns of emotions associated with tweets . |
| Outcome: | The proposed dataset includes more than 45 million geo-located tweets from US and Canada . it shows that Canadian tweets tend to have higher valence, lower arousal, and higher dominance than the US tweets . |