Papers by Rajat Verma
Multilingual Tokenization through the Lens of Indian Languages: Challenges and Insights (2026.findings-acl)
Copied to clipboard
Maharaj Brahma, N J Karthika, Rajat Verma, Nagasai Saketh Naidu, Rohit Saluja, Maunendra Sankar Desarkar, Ganesh Ramakrishnan
| Challenge: | Existing tokenizers are often skewed towards high-resource languages limiting their effectiveness for linguistically diverse and morphologically rich languages. |
| Approach: | They evaluate multilingual tokenization across 17 Indic languages spanning 11 scripts and two language families. |
| Outcome: | The proposed method improves tokenization quality and vocabulary size in 17 languages . poor tokenization can lead to increase in sequence lengths, fragment meaningful units, weaken model's ability to capture linguistic structure and semantics. |
TEEMIL : Towards Educational MCQ Difficulty Estimation in Indic Languages (2025.coling-main)
Copied to clipboard
| Challenge: | Traditionally, educators manually create and calibrate MCQs, a process that is time-consuming and subjective. |
| Approach: | They propose to use TEEMIL-H and TEIMEL-K to create a dataset with manually annotated difficulty labels for MCQs in Hindi and Kannada. |
| Outcome: | The proposed datasets contain 4689 and 4215 MCQs with manually annotated difficulty labels. |