Papers by Arif Shahriar

1 papers
Improving Bengali and Hindi Large Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Bengali and Hindi are low-resource languages, and the state-of-the-art tokenization methods fail to separate roots from affixes.
Approach: They used BERT and Wordpiece tokenizers to train a wordpiece tokenization system for Bengali and Hindi to model fine-grained character-level information.
Outcome: The proposed tokenizers outperform the state-of-the-art and Wordpiece tokenizer for modeling Bengali and Hindi.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations