Papers by Priyankoo Sarmah

5 papers
Lexical Tone Recognition in Mizo using Acoustic-Prosodic Features (2020.lrec-1)

Copied to clipboard

Challenge: Mizo is an under-studied Tibeto-Burman tonal language of the Northeast of India.
Approach: They propose to use acoustic-prosodic parameters to automatically recognize four phonological tones in Mizo using a set of features computed from Fundamental Frequency contours.
Outcome: The proposed model performs better than the existing classifiers in recognizing four phonological tones in Mizo using acoustic-prosodic parameters.
AssameseBackTranslit: Back Transliteration of Romanized Assamese Social Media Text (2024.lrec-main)

Copied to clipboard

Challenge: a novel dataset capturing native text composed in the Roman/Latin script is presented . the dataset comprises 60,312 Roman-native parallel transliterated sentences .
Approach: They propose a back transliteration dataset capturing native text composed in the Roman/Latin script and its corresponding representation in the native Assamese script.
Outcome: The proposed dataset outperforms baseline models in terms of word-level transliteration evaluation benchmarks and performance assessments.
TEAM: A multitask learning based Taxonomy Expansion approach for Attach and Merge (2022.findings-naacl)

Copied to clipboard

Challenge: Existing methods for automating taxonomy expansion are attach and merge . elucidating the problem of limited coverage of WordNets is presented .
Approach: They propose a multitask learning-based deep learning method that performs both merge and attach operations in a single model.
Outcome: The proposed method outperforms state-of-the-art models on three WordNet taxonomies . it performs both merge and attach operations and also provides encouraging performance for merge operation .
Evaluating Performance of Pre-trained Word Embeddings on Assamese, a Low-resource Language (2024.lrec-main)

Copied to clipboard

Challenge: Word embeddings are not explored in high-resource languages such as Assamese, where resources are limited.
Approach: They propose to use assamese pre-trained word embeddings for sequence labeling tasks such as Parts-of-speech and Named Entity Recognition to evaluate their performance.
Outcome: The proposed embeddings outperform the existing methods on Parts-of-speech and Named Entity Recognition tasks.
AsNER - Annotated Dataset and Baseline for Assamese Named Entity recognition (2022.lrec-1)

Copied to clipboard

Challenge: Named entity recognition (NER) is a type of annotation that classifies text into predefined classes such as person, location, organization etc.
Approach: They propose to use a named entity annotation dataset for low resource Assamese language with a baseline NER model.
Outcome: The proposed dataset is likely to be significant resource for deep neural based Assamese language processing.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations