Papers by Priyankoo Sarmah
Lexical Tone Recognition in Mizo using Acoustic-Prosodic Features (2020.lrec-1)
Copied to clipboard
| Challenge: | Mizo is an under-studied Tibeto-Burman tonal language of the Northeast of India. |
| Approach: | They propose to use acoustic-prosodic parameters to automatically recognize four phonological tones in Mizo using a set of features computed from Fundamental Frequency contours. |
| Outcome: | The proposed model performs better than the existing classifiers in recognizing four phonological tones in Mizo using acoustic-prosodic parameters. |
AssameseBackTranslit: Back Transliteration of Romanized Assamese Social Media Text (2024.lrec-main)
Copied to clipboard
| Challenge: | a novel dataset capturing native text composed in the Roman/Latin script is presented . the dataset comprises 60,312 Roman-native parallel transliterated sentences . |
| Approach: | They propose a back transliteration dataset capturing native text composed in the Roman/Latin script and its corresponding representation in the native Assamese script. |
| Outcome: | The proposed dataset outperforms baseline models in terms of word-level transliteration evaluation benchmarks and performance assessments. |
TEAM: A multitask learning based Taxonomy Expansion approach for Attach and Merge (2022.findings-naacl)
Copied to clipboard
| Challenge: | Existing methods for automating taxonomy expansion are attach and merge . elucidating the problem of limited coverage of WordNets is presented . |
| Approach: | They propose a multitask learning-based deep learning method that performs both merge and attach operations in a single model. |
| Outcome: | The proposed method outperforms state-of-the-art models on three WordNet taxonomies . it performs both merge and attach operations and also provides encouraging performance for merge operation . |
Evaluating Performance of Pre-trained Word Embeddings on Assamese, a Low-resource Language (2024.lrec-main)
Copied to clipboard
| Challenge: | Word embeddings are not explored in high-resource languages such as Assamese, where resources are limited. |
| Approach: | They propose to use assamese pre-trained word embeddings for sequence labeling tasks such as Parts-of-speech and Named Entity Recognition to evaluate their performance. |
| Outcome: | The proposed embeddings outperform the existing methods on Parts-of-speech and Named Entity Recognition tasks. |
AsNER - Annotated Dataset and Baseline for Assamese Named Entity recognition (2022.lrec-1)
Copied to clipboard
| Challenge: | Named entity recognition (NER) is a type of annotation that classifies text into predefined classes such as person, location, organization etc. |
| Approach: | They propose to use a named entity annotation dataset for low resource Assamese language with a baseline NER model. |
| Outcome: | The proposed dataset is likely to be significant resource for deep neural based Assamese language processing. |