Papers by Rajat Verma

2 papers
Multilingual Tokenization through the Lens of Indian Languages: Challenges and Insights (2026.findings-acl)

Copied to clipboard

Challenge: Existing tokenizers are often skewed towards high-resource languages limiting their effectiveness for linguistically diverse and morphologically rich languages.
Approach: They evaluate multilingual tokenization across 17 Indic languages spanning 11 scripts and two language families.
Outcome: The proposed method improves tokenization quality and vocabulary size in 17 languages . poor tokenization can lead to increase in sequence lengths, fragment meaningful units, weaken model's ability to capture linguistic structure and semantics.
TEEMIL : Towards Educational MCQ Difficulty Estimation in Indic Languages (2025.coling-main)

Copied to clipboard

Challenge: Traditionally, educators manually create and calibrate MCQs, a process that is time-consuming and subjective.
Approach: They propose to use TEEMIL-H and TEIMEL-K to create a dataset with manually annotated difficulty labels for MCQs in Hindi and Kannada.
Outcome: The proposed datasets contain 4689 and 4215 MCQs with manually annotated difficulty labels.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations