Papers by Ashutosh Sathe

9 papers
Efficient Training of Language Models with Compact and Consistent Next Token Distributions (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods to train language models have focused on maximizing the likelihood of the next token . however, the construction and querying of such n-grams can be costly and impede training speed.
Approach: They propose a method to train language models faster by pre-aggregating corpus with collapsed n-gram distribution.
Outcome: The proposed model improves model quality and convergence rate while reducing variance across mini-batches compared to the standard next-token loss method.
MAFIA: Multi-Adapter Fused Inclusive Language Models (2024.eacl-long)

Copied to clipboard

Challenge: Pretrained Language Models (PLMs) are widely used in NLP for various tasks.
Approach: They propose to modularly debias a pre-trained language model across multiple bias dimensions using structured knowledge and a large generative model.
Outcome: The proposed model is able to debias a pre-trained language model across multiple bias dimensions in a semi-automated way.
MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks (2024.naacl-long)

Copied to clipboard

Challenge: Several new LLMs have been introduced necessitating their evaluation on non-English languages.
Approach: They perform a thorough evaluation of the non-English capabilities of SoTA LLMs by comparing them on the same set of multilingual datasets.
Outcome: The proposed model outperforms models on multilingual datasets on 22 languages including low-resource African languages.
Benchmarking and Improving Text-to-SQL Generation under Ambiguity (2023.emnlp-main)

Copied to clipboard

Challenge: Existing decoding algorithms treat SQL queries as a string and produce unhelpful token-level diversity in the top-k.
Approach: They propose a benchmarking algorithm that generates all SQLs in top-k ranked outputs . they use plan-based template generation and constrained infilling to bridge this gap .
Outcome: The proposed algorithm is 2.5 times more effective than state-of-the-art models at generating all candidate SQLs in the top-k ranked outputs.
MAPLE: Multilingual Evaluation of Parameter Efficient Finetuning of Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Prior work on multilingual evaluation has shown that there is a large gap between the performance of Large Language Models on English and other languages.
Approach: They propose to finetune Llama-2 and Mistral models on two datasets to determine their effect on model performance on six downstream tasks covering forty one languages.
Outcome: The proposed model can improve on six multilingual tasks while degrading on high-resource languages.
Improving Consistency in LLM Inference using Probabilistic Tokenization (2025.findings-naacl)

Copied to clipboard

Challenge: Prior work has shown that probabilistic tokenizations can generate multiple tokenization of the same input string.
Approach: They propose a method to leverage the multiple tokenization capabilities of modern LLM tokenizers.
Outcome: The proposed method improves the self-consistency of large language models by generating multiple tokenizations.
A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have highlighted the existence of social biases within large vision and language models.
Approach: They propose a framework for systematically evaluating gender, race, and age biases in vision-language models with respect to professions.
Outcome: The proposed framework covers all supported inference modes of the recent vision-language models, including image-to-text, text-to image, and image- to-image.
Improving Cross Lingual Transfer by Pretraining with Active Forgetting (2025.emnlp-main)

Copied to clipboard

Challenge: Prior work has shown that encoder-only LLMs show impressive cross lingual transfer of their capabilities from English to other languages.
Approach: They propose a pretraining strategy that uses active forgetting to achieve similar cross lingual transfer in decoder-only LLMs.
Outcome: The proposed model improves cross lingual transfer capabilities on non-English languages despite being trained on English data.
Diverse Parallel Data Synthesis for Cross-Database Adaptation of Text-to-SQL Parsers (2022.emnlp-main)

Copied to clipboard

Challenge: Adapting Text-to-SQL parsers to new database schemas is a challenging task owing to a vast diversity of schemas and zero availability of natural language queries in new schemas.
Approach: They propose a framework for synthesizing parallel datasets for adapting Text-to-SQL parsers.
Outcome: The proposed framework outperforms existing methods on databases with diverse schemas and zero availability of natural language queries.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations