Papers by Hemanta Baruah
AssameseBackTranslit: Back Transliteration of Romanized Assamese Social Media Text (2024.lrec-main)
Copied to clipboard
| Challenge: | a novel dataset capturing native text composed in the Roman/Latin script is presented . the dataset comprises 60,312 Roman-native parallel transliterated sentences . |
| Approach: | They propose a back transliteration dataset capturing native text composed in the Roman/Latin script and its corresponding representation in the native Assamese script. |
| Outcome: | The proposed dataset outperforms baseline models in terms of word-level transliteration evaluation benchmarks and performance assessments. |