Papers by Sanjeev Kumar
SrcMix: Mixing of Related Source Languages Benefits Extremely Low-resource Machine Translation (2026.findings-eacl)
Copied to clipboard
| Challenge: | Multilingual models are widely used for machine translation, but their effectiveness for extremely low-resource languages (ELRLs) is dependent on how related languages are incorporated during fine-tuning. |
| Approach: | They propose a source-side mixing strategy that combines related ELRLs during fine-tuning while constraining the decoder to a single target language. |
| Outcome: | The proposed approach improves performance in high-resource to ELRL translations and in mid-resourced to MT translations. |
Part-of-speech Tagging for Extremely Low-resource Indian Languages (2024.findings-acl)
Copied to clipboard
| Challenge: | Modern natural language processing systems thrive when given access to large datasets, but a large fraction of the world’s languages are not privy to such benefits due to sparse documentation and inadequate digital representation. |
| Approach: | They propose a parallel part-of-speech evaluation dataset for Angika, Magahi, Bhojpuri and Hindi. |
| Outcome: | The proposed approach improves F1 scores by up to 8% on Angika, Magahi, Bhojpuri and Hindi while ignoring the tokenization challenge. |