Papers by Arif Shahriar
Improving Bengali and Hindi Large Language Models (2024.lrec-main)
Copied to clipboard
| Challenge: | Bengali and Hindi are low-resource languages, and the state-of-the-art tokenization methods fail to separate roots from affixes. |
| Approach: | They used BERT and Wordpiece tokenizers to train a wordpiece tokenization system for Bengali and Hindi to model fine-grained character-level information. |
| Outcome: | The proposed tokenizers outperform the state-of-the-art and Wordpiece tokenizer for modeling Bengali and Hindi. |