Papers by Pushpak Bhattacharyya
Copied to clipboard
| Challenge: | Existing studies on suicide notes have not explored the topic of emotion detection. |
| Approach: | They develop a fine-grained emotion annotated corpus of suicide notes in English and use it to perform emotion detection on a curated dataset. |
| Outcome: | The proposed model performs emotion detection on a curated dataset of 205 suicide notes in English. |
Copied to clipboard
| Challenge: | Existing research indicates that disfluencies can constitute up to 5.9% of words in spontaneous speech, with repetitions accounting for over half of these disfluency. |
| Approach: | They propose to use a dataset to analyze reduplication and repetition in speech using computational linguistics to evaluate transformer-based models. |
| Outcome: | The proposed models achieve macro F1 scores of up to 85.62% in Hindi, 83.95% in Telugu, and 84.82% in Marathi for reduplication-repetition classification. |
Copied to clipboard
| Challenge: | Existing word embedding models mix semantic similarity with other types of relatedness. |
| Approach: | They propose a model that leverages relational knowledge available in a knowledge resource to improve word embeddings. |
| Outcome: | The proposed model improves word embeddings on synonymy, antonymy and hypernymy relations in WordNet and significantly improves lexical entailment detection task. |
Copied to clipboard
| Challenge: | Empirical results show the efficacy of our proposed multi-task framework over existing state-of-the-art systems. |
| Approach: | They propose a multi-task, multi-modal deep learning framework to solve multiple tasks simultaneously. |
| Outcome: | The proposed framework performs better than existing state-of-the-art systems on a complicated form of information, i.e., memes. |
Copied to clipboard
| Challenge: | A study of multilingual fine-tuning yields better performance on downstream NLP applications . low resource languages such as Oriya and Punjabi are found to be the largest beneficiaries of multi-lingual fine tuning. |
| Approach: | They propose to leverage the relatedness of languages that belong to the same family in NLP models by multilingual fine-tuning. |
| Outcome: | The proposed approach improves performance on downstream NLP tasks by 15% compared to monolingual fine-tuning. |
Copied to clipboard
| Challenge: | Existing Document-level machine translation systems struggle to handle discourse-level phenomena such as pronoun resolution, lexical cohesion, and ellipsis. |
| Approach: | They propose a graph-based document-level machine translation framework that leverages Large Language Models to model translation flow and discourse structure. |
| Outcome: | The proposed framework outperforms commercial and closed systems in eight languages and six domains. |
Copied to clipboard
| Challenge: | Large Vision Language Models lack domain-specific data for reasoning on complex problems. |
| Approach: | They propose to use explicit knowledge-infused questions, answers, and reasons to answer and reason upon the questions. |
| Outcome: | The proposed model improves by 25% over the baseline model. |
Copied to clipboard
| Challenge: | Prior work on ML based lemmatization focused on high resource languages, where data sets (word forms) are readily available. |
| Approach: | They propose to use neural methods to relate inflected forms of words to their dictionary form to reduce the sparse data problem. |
| Outcome: | The proposed methods can give competitive accuracy even in low resource setting. |
Copied to clipboard
| Challenge: | Automated Grammatical Error Correction (GEC) is a scarcely explored low-resource language . a recent study focused on English, but it focused on Hindi, which presents unique challenges due to its complex syntax and intricate morphology. |
| Approach: | They propose to use a human-edited dataset to generate Hindi GEC data . they also investigate round trip translation using diverse languages for the technique . |
| Outcome: | The proposed method outperforms other methods in Hindi, showing that it is highly efficient. |
Copied to clipboard
| Challenge: | Existing text-to-image frameworks for figurative illustration rely on proprietary models or human supervision to achieve adequate alignment. |
| Approach: | They propose a critique-driven framework that uses VLM feedback to refine visual elaborations for figurative image generation. |
| Outcome: | The proposed framework outperforms existing figurative image-to-text pipelines on human-supervised visual elaborations. |
Copied to clipboard
| Challenge: | Visual metaphors are a complex vision–language phenomenon that requires both perceptual and conceptual reasoning to understand. |
| Approach: | They introduce a visual metaphor dataset featuring 2177 synthetic and 350 human-annotated images and benchmark several SOTA VLMs on two tasks: Visual Metaphor Captioning (VMC) and Visual Metamorphosis VQA (VM-VQA). |
| Outcome: | The proposed model outperforms standard few-shot baselines on visual metaphors and VM-VQA tasks. |
Copied to clipboard
| Challenge: | Existing peer review system is not straightforward and requires domain knowledge, expertise, and intelligence of human reviewers, which is somewhat elusive with the current state of AI. |
| Approach: | They propose to use peer review texts to predict acceptance or rejection of a manuscript based on reviewer sentiment. |
| Outcome: | The proposed deep neural architecture achieves significant performance improvement over baselines (29% error reduction) in a recently released dataset of peer reviews. |
Copied to clipboard
| Challenge: | Recent laws like “right to explanations” have spurred research in developing interpretable models . a recent study has shown that multimodal explanations improve performance in generating textual justifications . |
| Approach: | They propose to use visual and textual modalities to explain why a given meme is cyberbullying . they use a Contrastive Language-Image Pretraining approach to generate textual justifications . |
| Outcome: | The proposed model improves performance in visual and textual explanations and identifies the visual evidence supporting a decision. |
Copied to clipboard
| Challenge: | Humor is an essential aspect of daily conversation, and people try to provoke humor in their talks. |
| Approach: | They propose a multitask framework that annotates Hindi utterances with sentiment and emotion classes. |
| Outcome: | The proposed framework improves on the recently released Hindi Humor dataset . it takes sentiment and emotion into account to understand humor . |
Copied to clipboard
| Challenge: | Using a dataset of 931 videos with 4021 code-mixed Hindi-English utterances, we find that video content with multiple modalities is more accurate and more accurate than textual content. |
| Approach: | They propose to use a dataset to analyze toxic content in video content in non-English languages by leveraging language models. |
| Outcome: | The proposed framework achieves an Accuracy and Weighted F1 score of 94.29% and 94.35% for the first time in its class. |
Copied to clipboard
| Challenge: | Existing methods for hallucination detection depend on knowledge sources that are explicit such as Wikipedia or knowledge graphs. |
| Approach: | They propose a cognitive approach that leverages gaze signals from humans to detect hallucinations in natural language processing (NLP) they collect and introduce an eye tracking corpus consisting of 500 instances, annotated by five annotators for hallucinism detection. |
| Outcome: | The proposed approach achieves a balanced accuracy of 87.1% on a FactCC dataset. |
Copied to clipboard
| Challenge: | Existing methods rely on annotated labels but overlook the reasoning process humans naturally use to interpret implicit meaning. |
| Approach: | They propose a dataset that includes explicit reasoning for both correct and incorrect interpretations and propose supervised fine-tuning to improve their performance. |
| Outcome: | The proposed dataset improves LLMs' pragmatic understanding by 11.12% across model families and 16.10% over label trained models. |
Copied to clipboard
| Challenge: | Existing methods for predicting semantic labels for noun compound interpretation are difficult. |
| Approach: | They propose to predict semantic labels in a continuous embedding space using FrameNet data. |
| Outcome: | The proposed method performs well on unseen labels, with 5% and 2% improvement over baselines for frame and FE prediction. |
Copied to clipboard
| Challenge: | Existing research on stereotypical biases ignores literature on them and results in resource wastage. |
| Approach: | They argue that stereotypes are social constructs shaping human perception and behavior that can produce harmful outcomes under specific conditions. |
| Outcome: | The proposed models can inherit and amplify stereotypes under certain conditions. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are widely used in industry but still produce hallucinations, limiting their reliability in critical applications. |
| Approach: | They propose to reduce hallucinations in consumer grievance chatbots by reducing their token accuracy by 0.4159 per turn. |
| Outcome: | The proposed system achieves an F1 score of 68.92% outperforming baseline detectors by 22.47% while maintaining the highest token accuracy. |
Copied to clipboard
| Challenge: | a novel retrofitting method to induce emotion aspects into pre-trained language models is proposed . the models are computationally less expensive and open, but do not capture affective aspects of human communication well. |
| Approach: | They propose a retrofitting method to induce emotion aspects into pre-trained language models . they retrofit text fragments exhibiting similar emotions into pretrained networks . |
| Outcome: | The proposed method produces emotion-aware text representations for sentiment analysis and sarcasm detection tasks. |
Copied to clipboard
| Challenge: | Disfluencies in conversational speech can affect performance of downstream NLP tasks. |
| Approach: | They propose a disfluency correction model that converts disfluent to fluent text . they use unsupervised encoder-decoder models to generate semi-supervised models . |
| Outcome: | The proposed model achieves a BLEU score of 79.39 on the Switchboard corpus test set and 85.28 with semi-supervision. |
Copied to clipboard
| Challenge: | IndicIRSuite is the first attempt at building large-scale Neural Information Retrieval resources for a large number of Indian languages. |
| Approach: | They introduce Neural Information Retrieval resources for 11 widely spoken Indian Languages from two major Indian language families. |
| Outcome: | Experiments show that Indic-ColBERT improves on INDIC-MARCO datasets for 11 languages, and that it can be used to improve IR for Indian languages. |
Copied to clipboard
| Challenge: | In recent past, social media has emerged as an active platform in the context of healthcare and medicine. |
| Approach: | They propose to use a novel adversarial learning approach to capture medical sentiments expressed in a medical blog to analyze the user's opinions on health-related issues. |
| Outcome: | The proposed framework can capture the user's opinions on health-related issues at a medical blog level. |
Copied to clipboard
| Challenge: | Existing approaches to generate general and aspect-specific opinion summarization are limited due to their reliance on human-specified aspects and seed words. |
| Approach: | They propose synthetic dataset creation approaches for general and aspect-specific opinion summarization . general opinion summaries struggle to generate faithful to the input reviews, they say . aspect- specific opinion summarisation models are limited due to reliance on human-specified aspects . |
| Outcome: | The proposed approach outperforms existing models on three e-commerce test sets on general and aspect-specific opinion summarization. |
Copied to clipboard
| Challenge: | Existing approaches to train multiple languages with a shared encoder and multiple decoders are based on denoising autoencoding of each language and back-translating between English and multiple non-English languages. |
| Approach: | They propose a multilingual unsupervised NMT scheme which trains multiple languages with a shared encoder and multiple decoders. |
| Outcome: | The proposed model performs better than the separately trained bilingual models on monolingual corpora and improves by 1.48 BLEU points on WMT test sets. |
Copied to clipboard
| Challenge: | Wordnets are rich lexico-semantic resources. Linked wordnets link similar concepts in wordnet of different languages. |
| Approach: | They propose to map 18 Indian wordnets linked with Princeton WordNet . they use expansion approach with Hindi Wordnet as pivot . |
| Outcome: | The proposed mappings of 18 Indian wordnets are based on Princeton WordNet . they show that availability of such resources will have a direct impact on NLP progress . |
Copied to clipboard
| Challenge: | Existing evaluation methods for opinion summarizations lack adequate opinion summary evaluation datasets. |
| Approach: | They propose a dataset that combines 7 dimensions crucial to opinion summaries . they propose OP-I-PROMPT, a dimension-independent prompt, and OP PROMPTS, . |
| Outcome: | The proposed model achieves a Spearman correlation of 0.70 with human judgments, surpassing prior methods. |
Copied to clipboard
| Challenge: | Existing task-oriented conversational agents assume that end-users will always have a pre-determined and servable task goal, which results in dialogue failure in hostile scenarios, such as goal unavailability. |
| Approach: | They propose to build an end-to-end multi-modal persuasive dialogue system incorporating a personalized persuasive module aided goal controller and goal persuader. |
| Outcome: | The proposed system achieves user tasks even in goal unavailability scenarios by persuading them towards a similar and servable goal. |
Copied to clipboard
| Challenge: | Existing reports are labor-intensive and expert-intensive, resulting in inconsistencies and a lack of patient-centered insight. |
| Approach: | They propose a multimodal prompt-driven report generation framework that integrates diverse data modalities to produce comprehensive and context-aware radiology reports. |
| Outcome: | The proposed framework improves report quality, improves understandability and could foster better patient-doctor communication. |
Copied to clipboard
| Challenge: | Existing surveys on RRG emphasize deep learning while overlooking the critical role of causality. |
| Approach: | They propose to analyze biases across the RRG pipeline and formalize it as a causal modeling problem and review representative causal techniques from the literature. |
| Outcome: | The proposed model can mitigate biases and yield fair, reliable systems with clinically meaningful outputs. |
Copied to clipboard
| Challenge: | Existing datasets for emotion recognition in dialogues are in English . existing datasets are limited to a few languages like Hindi . |
| Approach: | They propose a large conversational dataset in Hindi for multi-label emotion and intensity recognition in conversations . they use a Wizard-of-Oz manner to annotate dialogues with 16 emotion labels . |
| Outcome: | The proposed dataset contains 1,814 dialogues with 44,247 utterances in Hindi . it is based on a Wizard-of-Oz manner and can detect emotions in conversation . |
Copied to clipboard
| Challenge: | Query-focused Summarization (QfS) is a system that generates summaries from document(s) based on a query. |
| Approach: | They propose a Query-focused Summarization approach that uses a generalization of Reinforcement Learning (RL) for Natural Language Generation and a better semantic similarity reward. |
| Outcome: | The proposed approach improves on the ROUGE-L metric and in a benchmark dataset. |
Copied to clipboard
| Challenge: | Existing studies on multilingual automatic post-editing systems for low-resource Indo-Aryan languages have focused on different models for different language pairs. |
| Approach: | They propose to use a multilingual automatic post-editing system to improve machine translations for low-resource Indo-Aryan languages. |
| Outcome: | The proposed model outperforms English-Hindi and English-Marathi models by 2.5 and 2.39 TER points. |
Copied to clipboard
| Challenge: | Word embeddings that encode lexical-semantic relations do not capture emotion aspects of words. |
| Approach: | They propose a retrofitting method to update the vectors of emotion bearing words . they find that the retrofitted embeddings achieve better distances between clusters . |
| Outcome: | The proposed method achieves better distances between clusters and clusters for words having the same emotions. |
Copied to clipboard
| Challenge: | Disfluencies can be introduced in conversational speech due to the conversational nature of speech and/or speech impairments such as stuttering. |
| Approach: | They propose an adversarial sequence-tagging model for Disfluency Correction . they evaluate it in Bengali, Hindi, and Marathi languages and use it to correct stuttering disfluencies . |
| Outcome: | The proposed technique improves in Bengali, Hindi, and Marathi languages . it also removes stuttering disfluencies in ASR transcripts introduced by speech impairments . |
Copied to clipboard
| Challenge: | Multilingual models are widely used for machine translation, but their effectiveness for extremely low-resource languages (ELRLs) is dependent on how related languages are incorporated during fine-tuning. |
| Approach: | They propose a source-side mixing strategy that combines related ELRLs during fine-tuning while constraining the decoder to a single target language. |
| Outcome: | The proposed approach improves performance in high-resource to ELRL translations and in mid-resourced to MT translations. |
Copied to clipboard
| Challenge: | despite growing interest in GEC, most research has focused on English due to the lack of benchmark datasets for low-resource lan-guages. |
| Approach: | They propose a new approach to generate high-quality synthetic data for GEC using monolingual corpora. |
| Outcome: | The proposed framework outperforms other monolingual methods in English, Hindi, Bengali, Marathi, and Tamil. |
Copied to clipboard
| Challenge: | Speech Act Classification determining the communicative intent of an utterance has been investigated widely over the years as a standalone task. |
| Approach: | They propose a multi-modal, emotion-TA dataset called EmoTA from open-source Twitter dataset and a Dyadic Attention Mechanism framework that integrates intra-modal and inter-modal attention to fuse multiple modalities. |
| Outcome: | The proposed framework boosts the performance of the primary task, i.e., TA classification (TAC), by benefitting from the two secondary tasks, namely, Sentiment and Emotion Analysis compared to its uni-modal and single task TAC variants. |
Copied to clipboard
| Challenge: | In this paper, we show that the combination of Phrase Pair Injection and Corpus Filtering boosts the performance of Neural Machine Translation systems. |
| Approach: | They propose to combine Phrase Pair Injection and Corpus Filtering to boost performance of Neural Machine Translation systems. |
| Outcome: | The proposed method improves machine translation models on low-resource language pairs . BLEU score improves over models trained with whole pseudo-parallel corpus augmented with parallel corpus. |
Copied to clipboard
| Challenge: | Cognates are words that have a common etymological origin and can facilitate the Second Language Acquisition (SLA) however, they also pose a challenge to various NLP applications such as Machine Translation and Cross-lingual Sense Disambiguation. |
| Approach: | They create two cognate datasets for twelve Indian languages and use them to generate cognate sets. |
| Outcome: | The proposed datasets are curated using previously available baseline cognate detection approaches and evaluated with the help of lexicographers. |
Copied to clipboard
| Challenge: | Existing studies do not focus on linguistically grounded attacks, but pre-trained models are susceptible to these perturbations. |
| Approach: | They propose to examine whether pre-trained language models are agnostic to linguistically grounded attacks . they find that PLMs are less susceptible to linguistic perturbations than non-linguistic ones . |
| Outcome: | The proposed model is agnostic to linguistically grounded attacks, but is less susceptible to linguist attacks than non-linguistic models. |
Copied to clipboard
| Challenge: | Stereotypes are known to have harmful effects, making their detection critical . current research focuses on detecting and evaluating stereotypical biases . |
| Approach: | They propose a five-tuple definition and provide precise terminologies disentangling stereotypes, antistereotypes, stereotypical bias, and general bias. |
| Outcome: | The proposed framework disentangles stereotypes, antistereotypes, stereotypical bias, and general bias. |
Copied to clipboard
| Challenge: | Existing approaches to detect metaphor and hyperbole independently have not explored their relationship computationally. |
| Approach: | They propose a multi-task deep learning framework to detect hyperbole and metaphor simultaneously by annotating two hyperbolic datasets with metaphor labels. |
| Outcome: | The proposed framework improves state-of-the-art hyperbole detection by 12% over existing methods. |
Copied to clipboard
| Challenge: | Existing studies show that transfer learning works best when the languages are related. |
| Approach: | They propose to pre-order assisting language sentences to match the word order of the source language and train the parent model. |
| Outcome: | The proposed model can improve translation quality in low-resource scenarios by pre-ordering the assisting language sentences to match the word order of the source language and training the parent model. |
Copied to clipboard
| Challenge: | More than half of the world's population is presumed to be bilingual . spoken translation of code-switched speech has been under-explored . |
| Approach: | They propose an end-to-end model architecture CoSTA that scaffolds on pretrained ASR and MT modules. |
| Outcome: | The proposed model outperforms existing models by 3.5 BLEU points in spoken translation of code-switched speech. |
Copied to clipboard
| Challenge: | Contemporary deep learning models handle languages with diverse morphology . morphological complexity of languages is closely linked with positional encodings . |
| Approach: | They propose to use positional encodings to integrate morphological complexity into deep learning models. |
| Outcome: | The proposed model improves on 22 languages and 5 downstream tasks. |
Copied to clipboard
| Challenge: | Automated essay grading (AEG) is one of the most challenging activities in natural language processing (NLP). |
| Approach: | They propose to annotate the ASAP AEG dataset and use it to score different attributes of the essays. |
| Outcome: | The proposed resource is based on the ASAP++ dataset, which contains scores for different attributes of the essays, such as content, word choice, organization, sentence fluency, etc. |
Copied to clipboard
| Challenge: | Existing systems that control concept transitions in a conversation lack a persona-aware topic transition dataset. |
| Approach: | They propose a persona-aware topic-guiding conversational system that leads the conversation to drift to a set of target concepts depending on the persona of the speaker and the context of the conversation. |
| Outcome: | The proposed system produces fluent responses with no useful information and is based on a conversational dataset with a human-in-loop only quality checks. |
Copied to clipboard
| Challenge: | Intent classification is crucial for conversational agents, and deep learning models perform well in this area due to the lack of suitable benchmark data. |
| Approach: | They propose a technique to augment text samples from intent classification datasets with word-level explanations by marking main predicates and their arguments as explanation signals. |
| Outcome: | The proposed method augments text samples from intent classification datasets with word-level explanations. |
Copied to clipboard
| Challenge: | Large enterprises have several teams to create their content for the purpose of marketing, campaigning, or even maintaining a brand presence. |
| Approach: | They propose a new unified Vision-Language (VL) model with a focus on context-assisted image captioning where the caption is generated based on both the image and its context. |
| Outcome: | The proposed model achieves state-of-the-art with an improvement of up to 8.34 CIDEr score on the benchmark news image captioning datasets. |
Copied to clipboard
| Challenge: | Existing methods to train neural network-based models for code-mixing are limited due to language specificity of code-mixed text. |
| Approach: | They propose a deep learning approach to generate code-mixed text from English to multiple languages without any parallel data. |
| Outcome: | The proposed approach generates a code-mixed text from English to multiple languages without any parallel data. |
Copied to clipboard
| Challenge: | In automatic essay grading, essay traits are important for scoring the essay holistically . a single-task learning system gives the best results for scoring essays holistically and scoring essay traits. |
| Approach: | They propose a way to score essays using a multi-task learning approach . they compare the MTL-based BiLSTM system to a single-task Learning approach based on LSTMs and BiLStms . |
| Outcome: | The proposed system gives better results for scoring essay holistically and scoring essay traits. |
Copied to clipboard
| Challenge: | Humblebragging is a phenomenon in which individuals present self-promotional statements under the guise of modesty or complaints. |
| Approach: | They propose a task of automatically detecting humblebragging in text and propose '4-tuple definition' they also propose machine learning, deep learning, and large language models to perform the task . |
| Outcome: | The proposed model achieves an F1-score of 0.88 and is non-trivial even for humans. |
Copied to clipboard
| Challenge: | Considerable work on Dialogue Act Classification (DAC) has been done on textual inputs. |
| Approach: | They propose to use a multimodal Emotion aware Dialogue Act dataset to explore the role of multi-modality and emotion recognition in DAC. |
| Outcome: | The proposed dataset shows that multi-modality and emotion recognition improves DAC performance compared to uni-modal and single task DAC variants. |
Copied to clipboard
| Challenge: | Existing methods for multi-modal sentiment analysis are limited due to the use of text, visual and acoustic inputs. |
| Approach: | They propose a recurrent neural network based multi-modal attention framework that leverages contextual information for utterance-level sentiment prediction. |
| Outcome: | The proposed framework performs better on two multi-modal sentiment analysis benchmark datasets with accuracies of 82.31% and 79.80% for the MOSI and MOSEI datasets. |
Copied to clipboard
| Challenge: | Current methods for mental disorder prediction split data into chunks and use limited context length . mental health professionals lack the skills to diagnose and treat mental disorders . |
| Approach: | They propose a framework which compresses chronologically ordered social media posts into a series of numbers and uses this time variant representation for mental disorder classification. |
| Outcome: | The proposed framework outperforms existing models in depression, self-harm and anorexia . it also shows that the proposed framework can be used across domains . |
Copied to clipboard
| Challenge: | Recent studies have shown that Vision-Language models cannot understand visual metaphors in memes and adverts. |
| Approach: | They propose a task to describe visual metaphors in videos using a manually created dataset and a new metric called Average Concept Distance to automatically evaluate creativity. |
| Outcome: | The proposed system performs comparable to existing video language models on the proposed task and can be used for future research. |
Copied to clipboard
| Challenge: | Existing datasets for humour classification are limited due to the subjectivity of the content and the multiple interpretations of the data. |
| Approach: | They propose to annotate a multi-modal humour-annotated dataset using stand-up comedy clips and compute a humor quotient using the audience's laughter. |
| Outcome: | The proposed scoring mechanism is validated by comparing with manual scoring methods and achieves an accuracy of 0.813 in terms of QWK. |
Copied to clipboard
| Challenge: | Existing methods for document-level novelty detection are limited and do not require manual feature engineering. |
| Approach: | They propose a deep Convolutional Neural Networks based model to classify a document as novel or redundant on the basis of documents already seen by the system. |
| Outcome: | The proposed model outperforms the state-of-the-art on a document-level novelty detection dataset by a margin of 5% in terms of accuracy. |
Copied to clipboard
| Challenge: | Existing methods to predict text quality include estimating subjective aspects of text, like structure, clarity, etc. |
| Approach: | They propose to capture gaze behaviour to help predict text quality by reporting improvements obtained by adding gaze features to traditional textual features for score prediction. |
| Outcome: | The proposed model shows that capturing gaze behaviour improves the accuracy of score prediction when the reader has fully understood the text. |
Copied to clipboard
| Challenge: | a recent study shows that large language models perform well in low-resource languages . a vast majority of languages don't have comparable data as compared to English . |
| Approach: | They propose to use Translationese as synthetic data for pre-training language models for low-resource languages. |
| Outcome: | The proposed method reduces performance of LMs trained on clean data in Indian languages . the proposed model performs better in English than in other languages, but is not comparable to English. |
Copied to clipboard
| Challenge: | Unsupervised style transfer has been explored in text. |
| Approach: | They propose a system where aspect-level sentiments can be controlled at the output . they propose to use unsupervised techniques such as ABSA masked-language-modelling . |
| Outcome: | The proposed system is successful in controlling aspect-level sentiments. |
Copied to clipboard
| Challenge: | Emotion and sentiment classification in dialogues has gained popularity in recent times . a number of datasets are imbalanced in representing different emotions and consist of an only single emotion. |
| Approach: | They propose to use a dataset to analyze emotions and sentiments in dialogues . they use text, audio and video to identify the correct emotions with the appropriate intensity and sentiment in an utterance of a dialogue . |
| Outcome: | The proposed datasets are balanced in representing different emotions and consist of only one emotion. |
Copied to clipboard
| Challenge: | Existing research on disfluency correction has primarily focused on English due to the unavailability of large-scale open-source datasets. |
| Approach: | They propose to use an annotated human-annotated corpus to analyze disfluency correction in four important Indo-European languages to demonstrate the benefits. |
| Outcome: | The proposed model improves BLEU scores by 5.65 points when used with a state-of-the-art machine translation system. |
Copied to clipboard
| Challenge: | Existing methods of identifying ADRs are reliable but time-consuming and offer a limited amount of ADR relevant information. |
| Approach: | They propose a neural network-inspired multi-task learning framework that can simultaneously extract ADRs from various sources. |
| Outcome: | The proposed framework achieves state-of-the-art performance on three publicly available real-world benchmark pharmacovigilance datasets, a Twitter dataset from PSB 2016 Social Me- dia Shared Task, CADEC corpus and Medline ADR corpus. |
Copied to clipboard
| Challenge: | Existing multimodal dialogue systems are based on unimodal sources, capturing information from text and image. |
| Approach: | They propose a position and attribute aware attention mechanism to learn enhanced image representation conditioned on the user utterance. |
| Outcome: | The proposed model outperforms the state-of-the-art models on text similarity metrics. |
Copied to clipboard
| Challenge: | Large language models (LLMs) produce state-of-the-art performance on natural language to code generation for resource-rich general-purpose languages like C++, Java, and Python. |
| Approach: | They propose a framework that breaks the NL-to-Code generation task into two steps . they use library documentation to detect the correct libraries and schema rules extracted from the documentation to constrain the decoding . |
| Outcome: | The proposed framework improves different sized language models across all six evaluation metrics, reducing syntactic and semantic errors in structured code. |
Copied to clipboard
| Challenge: | Open-source LLMs often depend on large proprietary models, which introduce serious privacy concerns. |
| Approach: | They propose a plug-and-play framework that improves SQL generation for smaller LLMs . they propose to apply question decomposition at the schema linking stage rather than during SQL generation . |
| Outcome: | The proposed framework improves schema linking recall by 25.1% and execution accuracy by 8.2% on the BIRD benchmark. |
Copied to clipboard
| Challenge: | Using prepositions, noun compounds are interpreted in two ways: labelling and paraphrasing. |
| Approach: | They propose to paraphrase noun compounds using prepositions by using parallelly aligned sequences of words. |
| Outcome: | The proposed approach performs well on datasets manually annotated with prepositions. |
Copied to clipboard
| Challenge: | Existing models of disease diagnosis using AI do not use knowledge infusion. |
| Approach: | They propose a transformer-based, knowledge-infused multi-modal medical dialogue generation framework . they propose 'discourse-aware' image identifier that recognizes signs and their severity . |
| Outcome: | The proposed model outperforms state-of-the-art models by 7.84% in the english language. |
Copied to clipboard
| Challenge: | Noun compound interpretation is the task of uncovering the semantic relation between the components of a noun compound. |
| Approach: | They propose an unsupervised method for identifying such relations between the components of a noun compound using pre-trained contextualized language models. |
| Outcome: | The proposed method outperforms supervised approaches for free paraphrasing and prepositional paraphrases using pre-trained language models to uncover ‘missing’ words. |
Copied to clipboard
| Challenge: | Empirical evaluation shows our model to outperform the single-hop question generation models on both automatic evaluation metrics such as BLEU, METEOR, and ROUGE and human evaluation metrics for quality and coverage of the generated questions. |
| Approach: | They propose a question-aware reward function to maximize the utilization of supporting facts in the context. |
| Outcome: | The proposed model outperforms single-hop neural question generation models on automatic evaluation metrics and human evaluation metrics for quality and coverage of the generated questions. |
Copied to clipboard
| Challenge: | Existing approaches to synthetic APE data generation use source (src) sentences in a parallel corpus to obtain translations (mt) through an MT system and treat corresponding reference (ref) sentences as post-edits (pe). |
| Approach: | They propose a reference-focused synthetic APE data generation technique that uses ‘ref’ instead of src’ sentences to obtain corrupted translations. |
| Outcome: | The proposed technique improves on English-German, English-Russian, English -Marathi, English and Hindi language pairs. |
Copied to clipboard
| Challenge: | Cross-domain sentiment analysis (CDSA) is a well-known problem in text analysis, but sufficient datasets may not be available for a domain to be trained. |
| Approach: | They propose to use 11 similarity metrics to facilitate cross-domain sentiment analysis to identify the best domains for CDSA for a given target domain. |
| Outcome: | The proposed approach performs better on 20 domain pairs and is validated by 11 similarity metrics. |
Copied to clipboard
| Challenge: | Current state of the art approaches for unsupervised neural machine translation (NMT) use only monolingual data for training. |
| Approach: | They propose an approach to filter back-translated data as part of the training process of unsupervised neural machine translation (NMT) they propose a weight component based on the quality of pseudo parallel sentence pairs generated in back-translation phase. |
| Outcome: | The proposed approach improves the training performance of unsupervised neural machine translation systems by giving weight to good pseudo parallel sentence pairs in the back-translation phase. |
Copied to clipboard
| Challenge: | In-car AI assistants struggle with multi-turn conversations and fail to handle cognitively complex follow-up questions. |
| Approach: | They propose a framework that leverages Bloom's Taxonomy to generate follow-up questions with increasing cognitive complexity and a Gricean-inspired evaluation framework to assess their Logical Consistency, Informativeness, Relevance, and Clarity. |
| Outcome: | The proposed framework validates both LLM-based and human evaluations and identifies the specific cognitive complexity level at which in-car AI assistants begin to falter information. |
Copied to clipboard
| Challenge: | Existing work on multi-domain, multi-lingual question answering is limited to the same language. |
| Approach: | They curate 500 articles in six different domains from the web and create question-answer pairs . they develop a deep learning based model for classifying an input question into coarse and finer categories . |
| Outcome: | The proposed model accuracies 90.12% and 80.30% for coarse and finer classes . the proposed model is the first attempt to create multi-domain, multi-lingual question answering evaluation involving English and Hindi. |
Copied to clipboard
| Challenge: | Qualitative and quantitative analysis shows that our proposed model can converse in both the languages and the information shared between the languages helps in improving the performance of the overall system. |
| Approach: | They propose a deep learning framework that can handle different languages and incorporate courteous behaviour in generic customer care responses in a multi-lingual scenario. |
| Outcome: | The proposed model can converse in both languages and the information shared between the languages helps in improving the overall performance of the system. |
Copied to clipboard
| Challenge: | Existing techniques for visual question answering focus on English questions, but many applications require a multilingual module. |
| Approach: | They propose a deep learning framework for multilingual and code- mixed visual question answering . they create Hindi and Code-mixed VQA datasets by exploiting linguistic properties of these languages . |
| Outcome: | The proposed model is capable of predicting answers from the questions in Hindi, English or Code- mixed (Hindi-English) languages. |
Copied to clipboard
| Challenge: | Almost 50% of depression patients face the risk of going into relapse. |
| Approach: | They propose to validate a social media dataset on depression relapse using cognitive theories of depression. |
| Outcome: | The first clinically validated social media dataset focused on depression relapse comprises 204 Reddit users annotated by mental health professionals. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated potential in code generation and natural language understanding, but they struggle with code constraints. |
| Approach: | They propose to use Large Language Models to handle constraints represented in code . they use JSON, YAML, XML, Python, and natural language to test their effectiveness . |
| Outcome: | The proposed benchmark shows that LLMs can handle code constraints better than natural language . the results suggest that conscious choice of representations can lead to optimal use of LLM in enterprise use cases involving code constraints. |
Copied to clipboard
| Challenge: | Existing approaches to cognate detection use orthographic, phonetic and semantic similarity based features sets. |
| Approach: | They propose a method for enriching feature sets with cognitive features extracted from gaze behaviour data from human readers’ gaze behaviour. |
| Outcome: | The proposed method improves cognate detection performance by 10% and 12% over existing methods. |
Copied to clipboard
| Challenge: | sarcasm and emotion are often used in conversational systems to generate the right response. |
| Approach: | They use a sarcastic expression dataset pre-annotated with 9 emotions to detect emotion . they identify and correct 343 incorrect emotion labels and label each sarkastic utterance with one of four sarcasm types. |
| Outcome: | The proposed model outperforms state-of-the-art sarcasm detection methods by using a multimodal sarcastic detection dataset. |
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is a lowerlevel task that aims to provide class labels like Person, Location, Organisation, Time, and Number to words in free text. |
| Approach: | They propose to use a standard-abiding Hindi NER dataset to analyze the annotations of a class of naming entities in free text. |
| Outcome: | The proposed dataset achieves a weighted F1 score of 88.78 with all the tags and 92.22 when we collapse the tag-set. |
Copied to clipboard
| Challenge: | Recent research has tackled this task using neural generative methods by augmenting emotion classes with the input sequences. |
| Approach: | They propose to use a self-attention based encoder and a decoder with dot product attention mechanism to generate a viable response with a specified emotion. |
| Outcome: | The proposed model outperforms baselines on automatic evaluation measures such as F1 and BLEU scores, thus resulting in more fluent and adequate responses. |
Copied to clipboard
| Challenge: | a corpus of 1.49 million parallel segments is available in the public domain . the corpus is the largest publicly available English-Hindi parallel corpus . |
| Approach: | They present the IIT Bombay English-Hindi Parallel Corpus . they present a compilation of public and private parallel corpora . |
| Outcome: | The corpus contains 1.49 million parallel segments, of which 694k were not previously available in the public domain. |
Copied to clipboard
| Challenge: | Mental health is a critical component of the United Nations’ Sustainable Development Goals (SDGs), particularly Goal 3 which aims to provide “good health and well-being”. |
| Approach: | They propose a task of detecting emotional reasoning and accompanying emotions in conversations that is manually annotated at the utterance level. |
| Outcome: | The proposed model achieves 6% accuracy and 4.62% accuracy on the emotion detection task and 3.56% accuracy, and 3.31% F1 on the ER detection task, compared to the existing state-of-the-art model. |
Copied to clipboard
| Challenge: | Using gaze behaviour to solve automatic essay grading tasks is costly in terms of time and money. |
| Approach: | They propose to collect gaze behaviour from 48 essays and learn gaze behaviour for the rest of the essays using a multi-task learning framework. |
| Outcome: | The proposed approach achieves a statistically significant improvement over the state-of-the-art system for the essay sets where gaze data is available. |
Copied to clipboard
| Challenge: | Existing methods to improve the reasoning capabilities of VQA systems are limited due to complexity of graph neural networks and end-to-end training. |
| Approach: | They propose a method to integrate Dense Passage Retrievers with Vision Language Models to boost the reasoning capabilities of VQA systems. |
| Outcome: | The proposed method outperforms human accuracy and GPT-4 in the ScienceQA dataset. |
Copied to clipboard
| Challenge: | Modern natural language processing systems thrive when given access to large datasets, but a large fraction of the world’s languages are not privy to such benefits due to sparse documentation and inadequate digital representation. |
| Approach: | They propose a parallel part-of-speech evaluation dataset for Angika, Magahi, Bhojpuri and Hindi. |
| Outcome: | The proposed approach improves F1 scores by up to 8% on Angika, Magahi, Bhojpuri and Hindi while ignoring the tokenization challenge. |
Copied to clipboard
| Challenge: | Quality Estimation (QE) is the task of evaluating the quality of a translation when reference translation is unavailable. |
| Approach: | They propose a Quality Estimation based Filtering approach to extract high-quality parallel data from the pseudo-parallel corpus. |
| Outcome: | The proposed approach improves the machine translation system performance by up to 1.8 BLEU points over the baseline model. |
Copied to clipboard
| Challenge: | Social media platforms such as Twitter and Facebook are a new channel of information dissemination for many negative groups for recruitment. |
| Approach: | They propose to use a social media sentiment analysis corpus annotated with the sentiment classes positive, negative and neutral to investigate the polarity of user-expressed opinions. |
| Outcome: | The proposed model is based on a set of benchmark datasets for sentiment analysis across a range of domains and languages. |
Copied to clipboard
| Challenge: | Existing datasets for read speech for Hindi lack expressiveness and character voice consistency. |
| Approach: | They propose to use a Hindi text-to-speech (TTS) dataset to train a multi-speaker model on the single-sector data and propose to improve expressiveness and character voice consistency. |
| Outcome: | The proposed model improves expressiveness and character voice consistency compared to the baseline single-speaker model. |
Copied to clipboard
| Challenge: | Timely generation of radiology reports and diagnoses is a challenge worldwide due to the enormous number of cases and shortage of radiologists. |
| Approach: | They propose a Knowledge Graph Augmented Vision Language BART model that takes two chest X-ray images and outputs a report with patient-specific findings. |
| Outcome: | The proposed model outperforms state-of-the-art transformer-based models on scoring metrics. |
Copied to clipboard
| Challenge: | Temporal sense detection of any word is an important aspect for detecting temporality at the sentence level. |
| Approach: | They build a temporal resource based on a semi-supervised learning approach . they use past, present, future, neutral and atemporal senses to tag sentences . |
| Outcome: | The proposed resource is based on a semi-supervised learning approach . it is used to tag sentences with past, present and future temporal senses . |
Copied to clipboard
| Challenge: | Word embeddings do not capture affective dimensions of valence, arousal, and dominance . valency, valance, and adolescence are present in words, but are not represented in text . |
| Approach: | They propose a method for updating word embeddings for affective meaning . they use a non-linear transformation function that maps pre-trained embedders to an affective vector space . |
| Outcome: | The proposed method improves inter-cluster and intra-c cluster distances for emotion-bearing words. |
Copied to clipboard
| Challenge: | Existing studies on content moderation of toxic memes focus on text-based content . current research neglects the widespread influence of multimodal content like memes . |
| Approach: | They propose a framework leveraging Large Language Models and Visual Language Model (VLMs) for meme intervention. |
| Outcome: | The proposed framework enables users to generate relevant and effective responses to toxic memes. |
Copied to clipboard
| Challenge: | Software Requirement Specification documents provide natural language descriptions of the core functional requirements as a set of use-cases. |
| Approach: | They propose a linguistic knowledge-based approach to extract software requirements from use-cases using a textual representation of the core functional requirements. |
| Outcome: | The proposed method performs better than existing techniques and improves performance. |
Copied to clipboard
| Challenge: | a study of 14 Indian languages shows that cognates can be detected by word embeddings . cognates are variants of the same lexical form across languages . |
| Approach: | They propose to use cross-lingual word embeddings to detect cognates among 14 Indian languages . they then evaluate the impact of their method on neural machine translation . |
| Outcome: | The proposed method improves on a dataset of 12 Indian languages . it also improves quality of the extracted cognates by up to 2.76 BLEU . |
Copied to clipboard
| Challenge: | Movies reflect society and also hold power to transform opinions. |
| Approach: | They propose to annotate movie scripts for identity bias using a dataset that is annotated for gender, race/ethnicity, religion, age, occupation, LGBTQ, and other . |
| Outcome: | The proposed dataset contains dialogue turns annotated for gender, race/ethnicity, religion, age, occupation, LGBTQ, and other, which contains biases like body shaming, personality bias, etc. |
Copied to clipboard
| Challenge: | BIStereo is a suite of language models that uncover body image stereotypes in language models. |
| Approach: | They propose a metric, TriSentBias, that captures the biased preferences of LMs towards a certain body type over others. |
| Outcome: | The proposed metric captures biased preferences of LMs towards a certain body type over others. |
Copied to clipboard
| Challenge: | In order to ensure customer satisfaction and retention, it is imperative for customer care agents and chatbots to be cordial and emphatic to the customer. |
| Approach: | They propose a deep learning framework that automatically transforms neutral customer care responses into courteous replies by stylistic transfer. |
| Outcome: | The proposed model can generate courteous expressions consistent with the emotional state of the customer while preserving the content. |
Copied to clipboard
| Challenge: | Pragmatics understanding is not well studied in LLMs, but their understanding of pragmatics is lacking. |
| Approach: | They propose to use a dataset to measure LLMs' understanding of pragmatics to evaluate their models. |
| Outcome: | The proposed dataset includes 14 tasks in four pragmatics phenomena, namely; Implicature, Presupposition, Reference, and Deixis. |
Copied to clipboard
| Challenge: | Existing APE and QE combination strategies have not shown significant performance gains in the field of automatic post-editing (APE). |
| Approach: | They propose to train a model on APE and QE tasks to improve the APE performance by using a multi-task learning methodology that treats both tasks as a 'bargaining game' they also investigate various existing combination strategies and show that their approach achieves state-of-the-art performance for a ‘distant’ language pair, viz., English-Marathi. |
| Outcome: | The proposed model improves on two different language pairs, viz., English-Marathi and English-German. |
Copied to clipboard
| Challenge: | conventional radiology workflows involve dictating diagnosis to transcriptionists, which is prone to delay and error. |
| Approach: | They propose to generate a set of knowledge graphs from a large collection of free-text radiology reports and use them to generate automatic radiology report generation. |
| Outcome: | The proposed model improves the reported BLEU-3, ROUGE-L, METEOR, and CIDEr scores by 2%, 4%, 2% and 2% respectively. |
Copied to clipboard
| Challenge: | Accurately grounding visual and textual elements within mobile user interfaces remains a challenge for Vision-Language Models (VLMs). |
| Approach: | They propose a mobile UI understanding model trained on a dataset specifically tailored for mobile screen understanding and grounding. |
| Outcome: | The proposed model achieves significant gains in accuracy across all perception tasks and on reasoning benchmarks. |
Copied to clipboard
| Challenge: | Detecting novelty of an entire document is an AI frontier problem . present state-of-the-art text matching techniques are unable to process such redundancy. |
| Approach: | They propose a document-level novelty detection resource that can be used to benchmark techniques . they crawl news documents across several domains and use it to find out whether they contain new information . |
| Outcome: | The proposed dataset is compared with a standard system for document novelty detection . the proposed system can detect elements that have not appeared before, or new or original . |
Copied to clipboard
| Challenge: | a huge amount of content is being generated every day due to the pervasiveness of social media. |
| Approach: | They firstly create a multi-domain tweet sentiment corpora and then establish a deep neural network based baseline framework to address the above mentioned issues. |
| Outcome: | The proposed dataset achieves 84.65% accuracy for sentiment analysis using a neural network, long short term memory, and gated recurrent unit (GRU). |
Copied to clipboard
| Challenge: | Existing benchmark datasets focus on English language and the Western context, leaving a void for a reliable dataset that encapsulates India’s unique socio-cultural nuances. |
| Approach: | They propose to use CrowS-Pairs to create a benchmark dataset that captures and evaluates social biases in Large Language Models (LLMs). |
| Outcome: | The proposed dataset is available in English and Hindi and leverages LLMs ChatGPT and InstructGPT to augment the existing dataset with diverse societal biases and stereotypes prevalent in India. |
Copied to clipboard
| Challenge: | Noun compounds are interesting constructs in Natural Language Processing . lack of standardized set of relation inventories and annotated datasets hinders interpretation . |
| Approach: | They propose a dataset that uses FrameNet as its semantic relation inventory to examine noun compounds. |
| Outcome: | The proposed dataset is linguistically grounded and uses FrameNet as its semantic relation inventory. |
Copied to clipboard
| Challenge: | Suicide continues to be one of the significant causes of death worldwide . EMotion-assisted personality subtyping is a novel approach to identify personality traits from suicide notes . |
| Approach: | They propose to use a PERSONAlity Detection Framework to identify personality traits from suicide notes and annotate them using a benchmark dataset. |
| Outcome: | The proposed method outperforms baselines on comprehensive evaluation using multiple state-of-the-art systems. |
Copied to clipboard
| Challenge: | Recent advances in large language models have significantly enhanced their ability to understand both natural language and code, but are prone to hallucinations. |
| Approach: | They propose a first-of-its-kind dataset, CodeSumEval, with 10K samples, curated specifically for hallucination detection in code summarisation. |
| Outcome: | The proposed framework has a 73% F1 score and is curated specifically for detection of hallucinations in code summarisation. |
Copied to clipboard
| Challenge: | Identifying distinct and independent participants in a narrative is crucial for many NLP applications. |
| Approach: | They propose an approach based on linguistic knowledge for identification of aliases mentioned using proper nouns, pronouns or noun phrases with common noun headword. |
| Outcome: | The proposed approach performs better than the state-of-the-art approach on four diverse history narratives of varying complexity. |
Copied to clipboard
| Challenge: | sarcasm detection depends on content spoken, tonality, facial expressions, context, and personal traits like language proficiency and cognitive capabilities. |
| Approach: | They propose to use synthetic gaze data to improve sarcasm detection in conversational context . they collect gaze features for 20% of data instances and use them to predict gaze features . |
| Outcome: | The proposed model improves performance on a conversational dataset using gaze features . it achieves a gain of 6.6% points on the complete dataset with only predicted gaze features. |
Copied to clipboard
| Challenge: | Existing approaches to improve NER performance add training data from one or more assisting languages to the primary language. |
| Approach: | They propose a metric based on symmetric KL divergence to filter out highly divergent training instances in the assisting language. |
| Outcome: | The proposed method improves NER performance in many languages, including those with limited training data. |
Copied to clipboard
| Challenge: | Existing methods to generate opinion summarization without supervised training data are limited due to the lack of additional sources. |
| Approach: | They propose a synthetic dataset creation strategy that leverages reviews and additional sources to generate a pseudo-summary. |
| Outcome: | The proposed approach achieves 14.5% improvement in ROUGE-1 F1 over existing models. |
Copied to clipboard
| Challenge: | In-context mixing is a prompting technique for effective in-contact learning with multilingual large language models. |
| Approach: | They propose a prompting technique called in-context mixing for effective in-constext learning with multilingual large language models. |
| Outcome: | The proposed prompts perform better with multilingual large language models. |
Copied to clipboard
| Challenge: | Currently, the majority of social bias datasets available are in English and this inhibits progress on social bias detection in low-resource languages. |
| Approach: | They propose a dataset for social bias detection in Hindi and investigate multilingual transfer learning using publicly available English, Italian, and Korean datasets. |
| Outcome: | The proposed dataset is compared with a dataset available in English, Italian, and Korean using multilingual models. |
Copied to clipboard
| Challenge: | Efficient word representations play an important role in solving various problems related to NLP, data mining, text mining etc. |
| Approach: | They propose to leverage bilingual word embeddings learned through a parallel corpus to minimize the effect of data sparsity. |
| Outcome: | The proposed model is tested against state-of-the-art methods in two experimental setups. |
Copied to clipboard
| Challenge: | Conventional approaches to QE involve training separate models at different levels of granularity viz., word-level, sentence-level and document-level . |
| Approach: | They propose to train a single model for sentence-level and word-level QE tasks in a multi-task learning framework and compare them to baseline models. |
| Outcome: | The proposed model improves on the single-pair, multi-patch, and zero-shot settings. |
Copied to clipboard
| Challenge: | Cross-domain sentiment classification is challenging due to polarity orientation and significance differences . supervised learning algorithms have to be re-trained on every new domain . |
| Approach: | They propose that words that do not change their polarity and significance represent transferable information across domains for cross-domain sentiment classification. |
| Outcome: | The proposed method improves cross-domain sentiment classification performance by identifying polarity-preserving significant words across domains. |
Copied to clipboard
| Challenge: | Past work on sarcasm detection has focused on identifying the sarcasm target of ridicule in a sarkastic text. |
| Approach: | They propose a task of extracting the sarcastic target of ridicule from a sarcastical text using a manually annotated dataset and an automatic approach. |
| Outcome: | The proposed approach establishes the viability of sarcasm target identification and will serve as a baseline for future work. |
Copied to clipboard
| Challenge: | Existing Question Answering systems for commercial aviation use a large number of documents . a Knowledge Graph (KG) guided Deep Learning (DL) based system can be used to query the documents based on accident reports . |
| Approach: | They propose a Knowledge Graph (KG) guided Deep Learning (DL) based Question Answering system to cater to these requirements. |
| Outcome: | The proposed system achieves 7% and 40% increase in accuracy over existing systems. |
Copied to clipboard
| Challenge: | Existing research has not explored the joint task of emotion detection and explanatory span identification in e-commerce reviews. |
| Approach: | They propose a joint task unifying Emotion detection and Opinion Trigger extraction (EOT) which explicitly models the relationship between causal text spans (opinion triggers) and affective dimensions (emotion categories). |
| Outcome: | The proposed framework surpasses zero-shot and chain-of-thought techniques across e-commerce domains. |
Copied to clipboard
| Challenge: | Existing systems for sarcasm detection are limited by the use of sarcasm . sarasm is often used to convey thinly veiled disapproval humorously. |
| Approach: | They propose a multi-task deep learning framework to solve sarcasm problems simultaneously . they manually annotate a sarcsm dataset with sentiment and emotion classes . |
| Outcome: | The proposed framework is able to solve sarcasm, sentiment and emotion problems in a multi-modal conversational scenario. |
Copied to clipboard
| Challenge: | Clinical practice frequently uses medical imaging for diagnosis and treatment. |
| Approach: | They propose a template-based approach to generate radiology reports from radiographs . they use multilabel image classifiers to generate tags, pathological descriptions from tags . |
| Outcome: | The proposed method improves on the most popular radiology report datasets. |
Copied to clipboard
| Challenge: | Disfluency correction models can help alleviate this problem, but the unavailability of labeled data in low-resource languages impairs progress. |
| Approach: | They propose to use a pretrained multilingual model to detect zero-shot disfluency in Indian languages. |
| Outcome: | The proposed model achieves F1 scores of 75 and higher on five disfluency types across four languages. |
Copied to clipboard
| Challenge: | a huge number of people use social media to express and exchange information in their own languages. |
| Approach: | They propose to use a code-mixed environment to extract higher level features from text . they use 'gadget' algorithm that automatically discovers higher level feature from text. |
| Outcome: | The proposed approach is generic and does not make use of handcrafted features or rules. |
Copied to clipboard
| Challenge: | Multi-modal analysis is a field emerging in the fields of natural language processing, computer vision and speech processing . multimodal analysis uses a variety of information from multiple sources to build efficient systems . acoustic and visual information can provide better information for classification decisions . |
| Approach: | They propose a recurrent neural network based approach for multi-modal sentiment and emotion analysis . they employ a context-aware attention module to exploit the correspondence among neighboring utterances . |
| Outcome: | The proposed model learns inter-modal interaction among participating modalities through auto-encoder mechanism . it is compared with existing state-of-the-art models on five standard multi-modal affect analysis datasets . |
Copied to clipboard
| Challenge: | a study conducted by the pew Internet & American Life Project 1 shows that almost 80 percent of Internet users have explored health-related topic online. |
| Approach: | They propose to crawl medical forums with opinions about medical condition self narrated by users. |
| Outcome: | The proposed system is based on opinions about medical condition self-narrated by users on medical forums. |
Copied to clipboard
| Challenge: | a significant gap exists in understanding code-mixed languages and the need for explainability in this context. |
| Approach: | They propose to annotate posts with four labels to identify bullies in code-mixed languages . they propose to use a generative framework to reimagine the multitask problem as a text-to-text generation task. |
| Outcome: | The proposed model outperforms baseline models and state-of-the-art models on the BullyExplain dataset. |
Copied to clipboard
| Challenge: | Statistical Machine Translation fails to handle the rich morphology when translating into morphologically rich language. |
| Approach: | They propose a method to generate unseen morphological forms from the parallel corpus . they propose morphology injection method to enrich the corpus with generated morphologies . |
| Outcome: | The proposed method improves the quality of the translation in English-Malayalam. |
Copied to clipboard
| Challenge: | Existing studies on MRC on scholarly articles have focused on general domain datasets of news articles and elementary school-level storybooks. |
| Approach: | They propose to generate automatic questions from span-of-word-based scholarly articles’ Reading Comprehension dataset with approximately 10K manually checked passage-question-answer instances. |
| Outcome: | The proposed model yields the F1 score of 37.31% and is useful for building Question-Answering (QA) systems on scientific articles. |
Copied to clipboard
| Challenge: | Event Extraction is an important task in the widespread field of NLP, but there is no benchmark setup in Hindi. |
| Approach: | They propose an Event Extraction framework for Hindi language and develop deep learning based models to set as the baselines. |
| Outcome: | The proposed framework crawls more than seventeen hundred disaster related Hindi news articles from various news sources. |
Copied to clipboard
| Challenge: | Existing methods to improve automatic post-editing (APE) systems struggle with over-correction, despite the principle of minimal editing. |
| Approach: | They propose a method that incorporates word-level Quality Estimation (QE) information during the decoding process. |
| Outcome: | The proposed method improves on English-German, English-Hindi, and English-Marathi language pairs, with TER gains of 0.65, 1.86, and 1.44 points, respectively. |
Copied to clipboard
| Challenge: | Inference-based scripts such as Abjad are difficult for cross-lingual models to learn in extremely low resource scenarios. |
| Approach: | They evaluate cross-lingual approaches for low resource languages and compare their performance against other models using different linguistic factors. |
| Outcome: | The proposed model on six low resource languages from two different families is compared with monolingual models on morphologically rich Indian languages. |
Copied to clipboard
| Challenge: | a new study addresses bias and stereotypes in language models by exploring how learning them together improves performance. |
| Approach: | They propose a dataset for bias and stereotype detection that integrates religion, gender, socio-economic status, race, profession, and others. |
| Outcome: | The proposed dataset compares encoder-only models and fine-tuned decoder- only models . the results show that learning stereotypes together improves bias detection . |
Copied to clipboard
| Challenge: | Human-machine interactions have increased rapidly assisting humans in their everyday lives. |
| Approach: | They propose to automatically identify the sentiment of the user and transform the neutral responses into polite responses conforming to the sentiment and the conversational history. |
| Outcome: | The proposed approach achieves superior performance compared to baseline models. |
Copied to clipboard
| Challenge: | Existing frameworks for sentiment and emotion analysis are not efficient for inter-task learning. |
| Approach: | They propose a multi-task learning framework that performs sentiment and emotion analysis together. |
| Outcome: | The proposed framework improves on a CMU-MOSEI dataset for sentiment and emotion analysis. |
Copied to clipboard
| Challenge: | Mental health disorders are one of the primary causes of disability worldwide . lack of qualified and competent mental health professionals is a major problem . we propose a virtual assistant that can act as the first point of contact and comfort for mental health patients. |
| Approach: | They propose a virtual assistant that can act as the first point of contact and comfort for mental health patients. |
| Outcome: | The proposed system outperforms baselines in the evaluation of 7k dyadic conversations from a peer-to-peer support platform. |
Copied to clipboard
| Challenge: | Mental models are "basic units of coherently structured knowledge" but stu-dents do not always construct coherent mental models, which can limit conceptual understanding. |
| Approach: | They propose an approach that infers the quality of students’ mental models from their multimodal responses using concept graphs as an analytical framework. |
| Outcome: | The proposed model infers the quality of students’ mental models from their multimodal responses using concept graphs as an analytical framework. |
Copied to clipboard
| Challenge: | Non-profit industry needs a system for accurately matching fund-seekers with fund-givers aligned in cause and target beneficiary group. |
| Approach: | They propose a search system that takes a fund-giver’s mission description as input and returns a ranked list of fund-seekers as output. |
| Outcome: | The proposed system improves on the non-profit evaluation dataset and the state-of-the-art model. |
Copied to clipboard
| Challenge: | This tutorial provides a comprehensive overview of two critical aspects of Large Language Models: bias and hallucination. |
| Approach: | This tutorial provides an overview of two critical aspects of Large Language Models: bias and hallucination. |
| Outcome: | This tutorial delves into the complex dimensions of Large Language Models (LLMs) it outlines ethical considerations pertinent to their development and discusses hallucination, a prevalent issue in generative AI systems such as LLMs. |
Copied to clipboard
| Challenge: | Temporal orientation refers to an individual’s tendency to connect to the psychological concepts of past, present or future and affects personality, motivation, emotion, decision making and stress coping processes. |
| Approach: | They propose to use a minimally supervised method to classify tweets in one of three temporal categories, past, present, and future, and a deep bi-directional long-term memory (BLSTM) to measure correlation between sentiment view of temporal orientation and different psycho-demographic factors. |
| Outcome: | The proposed method achieves 78.27% accuracy on a manually created test set. |
Copied to clipboard
| Challenge: | Existing QA systems that answer factual questions with short answers are rare in practice. |
| Approach: | They propose a proposed two-layered taxonomy technique for semantic question matching . they augment state-of-the-art deep learning models with question classes from a deep learning based question classifier . |
| Outcome: | The proposed technique achieves state-of-the-art on an open-domain dataset. |
Copied to clipboard
| Challenge: | Empirically, we show that the optimisation of multi-modal DAC, SA and ER tasks produces better results compared to its different counterparts. |
| Approach: | They propose a dual attention mechanism that integrates sentiment tags into a multi-modal conversational framework that integrate modal attentions and multiple loss optimization. |
| Outcome: | The proposed framework integrates sentiment tags for each utterance and learns generalized features across multiple tasks. |
Copied to clipboard
| Challenge: | Existing benchmark datasets suffer from leakage or evidence incompleteness, limiting the realism of current evaluations. |
| Approach: | They propose an agentic framework that iteratively generates and answers sub-questions to verify different aspects of the claim before finally generating the label. |
| Outcome: | The proposed system outperforms existing methods by 57.5% on Politi-Fact-Only and 3.05% on the widely used HOVER datasets. |