Papers by Rituraj Joshi

2 papers
Instruction Tuning on Public Government and Cultural Data for Low-Resource Language: a Case Study in Kazakh (2025.acl-long)

Copied to clipboard

Challenge: Instruction tuning in low-resource languages remains underexplored due to limited text data, particularly in government and cultural domains.
Approach: They propose to open-source a large-scale instruction-following dataset covering key institutional and cultural knowledge relevant to Kazakhstan.
Outcome: The proposed dataset improves LLMs’ understanding of procedural, legal, and structural governance topics.
Nanda Family: Open-Weights Generative Large Language Models for Hindi (2026.eacl-long)

Copied to clipboard

Challenge: Large language models remain predominantly English-centric, which limits their utility for underrepresented languages.
Approach: They propose to extend Llama’s vocabulary with 20% Hindi-specific tokens, thus halving Hindi tokenization fertility while preserving English efficiency.
Outcome: The proposed models outperform open-weight models of comparable size on a 65B-token corpus and bilingual instruction and safety alignment on . a culturally grounded dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations