Papers by Mukund Choudhary

2 papers
Do LLMs model human linguistic variation? A case study in Hindi-English Verb code-mixing (2026.findings-eacl)

Copied to clipboard

Challenge: Existing large language models (LLMs) do not reliably classify verb language preferences to match native speaker judgments.
Approach: They investigate whether large language models (LLMs) model linguistic variation by comparing Hindi-English verb code-mixing with English verb karna.
Outcome: The proposed models do not reliably classify verb language preferences to match native speaker judgments, but with specific supervision, some models do predict human preference to an extent.
Nanda Family: Open-Weights Generative Large Language Models for Hindi (2026.eacl-long)

Copied to clipboard

Challenge: Large language models remain predominantly English-centric, which limits their utility for underrepresented languages.
Approach: They propose to extend Llama’s vocabulary with 20% Hindi-specific tokens, thus halving Hindi tokenization fertility while preserving English efficiency.
Outcome: The proposed models outperform open-weight models of comparable size on a 65B-token corpus and bilingual instruction and safety alignment on . a culturally grounded dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations