Papers by Ramaneswaran S

7 papers
CoDa: Constrained Generation based Data Augmentation for Low-Resource NLP (2024.findings-naacl)

Copied to clipboard

Challenge: a low-resource dataset is limited in training data, so generating task-specific data is challenging.
Approach: They propose a data augmentation technique that prompts off-the-shelf instruction-following Large Language Models to generate augmentations.
Outcome: The proposed technique outperforms baselines on 11 datasets spanning 3 tasks and 3 low-resource settings.
ABEX: Data Augmentation for Low-Resource NLU via Expanding Abstract Descriptions (2024.acl-long)

Copied to clipboard

Challenge: ABEX is a novel and effective generative data augmentation methodology for low-resource Natural Language Understanding (NLU) tasks.
Approach: They propose a novel generative data augmentation methodology for low-resource Natural Language Understanding (NLU) tasks based on a paradigm for generating diverse forms of an input document .
Outcome: The proposed method outperforms all baselines qualitatively with improvements of 0.04% - 38.8%.
From Multilingual Complexity to Emotional Clarity: Leveraging Commonsense to Unveil Emotions in Code-Mixed Dialogues (2023.emnlp-main)

Copied to clipboard

Challenge: Understanding emotions during conversation is a fundamental aspect of human communication.
Approach: They propose an approach that integrates commonsense information with dialogue context to facilitate a deeper understanding of emotions.
Outcome: The proposed approach improves ERC for code-mixed conversations by integrating commonsense with dialogue context.
EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning (2024.emnlp-main)

Copied to clipboard

Challenge: EH-MAM is a self-supervised learning approach for speech representation learning . prior methods used random masking schemes to learn speech representations .
Approach: They propose a self-supervised approach that automatically selects hard regions during SSL training and introduces them to the model for reconstruction.
Outcome: The proposed approach outperforms state-of-the-art models across low-resource speech recognition and SUPERB benchmarks by 5%-10%.
MEMEX: Detecting Explanatory Evidence for Memes via Knowledge-Enriched Contextualization (2023.acl-long)

Copied to clipboard

Challenge: Besides digital archiving of memes and their metadata, there is no efficient way to deduce a meme’s context dynamically.
Approach: They propose a task to mine the context that succinctly explains the background of a meme and a related document to capture cross-modal semantic dependencies between the meme and the context.
Outcome: The proposed dataset outperforms existing systems and shows that it can capture cross-modal semantic dependencies between the meme and the context.
ACLM: A Selective-Denoising based Generative Data Augmentation Approach for Low-Resource Complex NER (2023.acl-long)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a task of detecting linguistically complex named entities in low-context text.
Approach: They propose a keyword-based augmentation approach to address the context-entity mismatch issue in complex name recognition (NER) they use selective masking to retain the named entities and certain keywords in the input sentence that provide contextually relevant additional knowledge or hints about the named entity.
Outcome: The proposed approach outperforms baseline methods on monolingual, cross-lingual, and multilingual complex NER in various low-resource settings.
DALE: Generative Data Augmentation for Low-Resource Legal NLP (2023.emnlp-main)

Copied to clipboard

Challenge: DALE addresses the challenges existing frameworks pose in generating effective data augmentations of legal documents.
Approach: They propose a generative Data Augmentation framework for low-resource legal NLP that exploits domain-specific language characteristics of templated legal documents to mask collocated spans of text.
Outcome: The proposed framework outperforms baseline frameworks on 13 datasets and 4 low-resource settings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations