Papers by Bhargav Shandilya

2 papers
Boosting the Capabilities of Compact Models in Low-Data Contexts with Large Language Models and Retrieval-Augmented Generation (2025.coling-main)

Copied to clipboard

Challenge: Existing language models lack data and computation power, but they are extremely parameter-heavy and difficult to train.
Approach: They propose a retrieval augmented generation framework backed by a large language model to correct the output of a smaller model for morphological glossing.
Outcome: The proposed model is highly effective in data-scarce settings and offers a state-of-the-art for morphological glossing.
BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data (2026.eacl-long)

Copied to clipboard

Challenge: prevailing trend in language modeling research is to prioritize scaling, authors say . from infancy to maturity, English learners acquire language through exposure to less than 100M words .
Approach: They propose a multilingual collection of datasets modeling the language a person observes from birth until they acquire a native language.
Outcome: The proposed models outperform models trained on a fixed, developmentally plausible English corpus on various benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations