Papers with Bodo
RELIC: Enhancing Reward Model Generalization for Low-Resource Indic Languages with Few-Shot Examples (2025.findings-emnlp)
Copied to clipboard
Soumya Suvra Ghosal, Vaibhav Singh, Akash Ghosh, Soumyabrata Pal, Subhadip Baidya, Sriparna Saha, Dinesh Manocha
| Challenge: | a new reward model for low-resource Indic languages is proposed . a preference-based training approach is prohibitively expensive, authors say . |
| Approach: | a new in-context learning framework is proposed to train a retriever to select in-constext examples from low-resource Indic languages. |
| Outcome: | a new in-context learning framework for reward modeling in low-resource Indic languages is developed . the proposed framework outperforms existing examples on three preference datasets . |
Generating Monolingual Dataset for Low Resource Language Bodo from old books using Google Keep (2022.lrec-1)
Copied to clipboard
Sanjib Narzary, Maharaj Brahma, Mwnthai Narzary, Gwmsrang Muchahary, Pranav Kumar Singh, Apurbalal Senapati, Sukumar Nandi, Bidisha Som
| Challenge: | Bodo is a scheduled Indian language spoken largely by the Boda community in Assam and other northeastern Indian states. |
| Approach: | They propose to generate a monolingual Bodo corpus from different books using Google Keep for OCR. |
| Outcome: | The proposed method generates a monolingual Bodo corpus from different books using free, accessible, and daily-usable applications. |