Papers by Sarath Chandar

22 papers
Exploring Quantization for Efficient Pre-Training of Transformer Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Quantization has proven to be effective after pre-training and during fine-tuning, but its effects on pre-trainer performance have remained unexplored.
Approach: They propose a linear quantization strategy to be applied during the pre-training of Transformers to improve model efficiency and stability.
Outcome: The proposed method improves model efficiency, stability, and performance while maintaining language modeling ability.
MLMLM: Link Prediction with Mean Likelihood Masked Language Model (2021.findings-acl)

Copied to clipboard

Challenge: Knowledge Bases (KBs) are easy to query, verifiable, and interpretable. however, they scale with man-hours and high-quality data.
Approach: They propose to commit the knowledge embedded in MLMs to a KB, making it interpretable . they propose to use a mean likelihood Masked Language Model to compare the likelihood of generating different entities to perform link prediction in a tractable manner.
Outcome: The proposed approach compares the likelihood of generating different entities to perform link prediction in a tractable manner.
Investigating the Multilingual Calibration Effects of Language Model Instruction Tuning (2026.eacl-short)

Copied to clipboard

Challenge: despite advances in foundation model research, the relationship between large language models and their calibration remains an open area of research.
Approach: They examine a gap in the calibration of large language models within multilingual settings to better understand how data scarcity can potentially lead to different calibration effects.
Outcome: The proposed calibration gap is found in two multilingual benchmarks over 29 and 42 languages.
Detecting Languages Unintelligible to Multilingual Models through Local Structure Probes (2022.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in multilingual pretrained models have proven effective at zero-shot transfer to a wide variety of languages, but this transfer is not universal, with many languages not currently understood by multilingual approaches.
Approach: They propose a general approach that requires only unlabelled text to detect which languages are not well understood by a cross-lingual model.
Outcome: The proposed model can detect which languages are not well understood by a multilingual model on 350 low-resource languages.
Do Large Language Models Know How Much They Know? (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models are highly capable systems, but their capabilities and limitations are unclear.
Approach: They develop a benchmark that challenges LLMs to recall all information they possess on specific topics.
Outcome: The proposed model can recall excessive, insufficient, or the precise amount of information they possess on a given topic, indicating their awareness of how much they know about the given topic.
Local Structure Matters Most: Perturbation Study in NLU (2022.findings-acl)

Copied to clipboard

Challenge: Recent research shows that neural models are insensitive to word-order perturbations, but other studies suggest that models learn some abstract notion of syntax.
Approach: They develop order-altering perturbations on the order of words, subwords, and characters to analyze their effect on neural models’ performance on language understanding tasks.
Outcome: The proposed models are insensitive to word-order perturbations while the local ordering remains relatively unperturbed.
Towards Lossless Encoding of Sentences (P19-1)

Copied to clipboard

Challenge: Existing methods for encoding text into lossless representations focus on performing well on downstream tasks and are unable to reconstruct original sequence from learned embedding.
Approach: They propose a lossless method for encoding long sequences of texts into feature rich representations by recursive autoencoding.
Outcome: The proposed method performs well on sentiment analysis and sentiment classification tasks.
Structure Learning for Neural Module Networks (D19-64)

Copied to clipboard

Challenge: Neural Module Networks are a class of neural networks that involve human-specified neural modules . current models only learn the parameters of the modules and/or the order of their execution .
Approach: They propose to learn internal structure and sequence without extra supervisory signals . they use dynamically composable modules which are then assembled into a layout .
Outcome: The proposed model performs comparable to models using hand-designed modules.
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have a tendency to hallucinate false or misleading information, limiting their reliability.
Approach: They examine how architecture-based inductive biases affect the propensity to hallucinate . they find that the models are more reliable and more reliable than traditional models .
Outcome: The proposed models can be used to train and train large language models that are factual or able to explain themselves through their knowledge.
Self-Influence Guided Data Reweighting for Language Model Pre-training (2023.emnlp-main)

Copied to clipboard

Challenge: Language Models (LMs) pre-trained with selfsupervision on large text data are the default starting point for developing models for various downstream tasks.
Approach: They propose a method for jointly reweighting samples by leveraging self-influence scores as an indicator of sample importance and pre-training.
Outcome: The proposed method promotes novelty and stability for model pre-training.
MVP: Minimal Viable Phrase for Long Text Understanding (2024.lrec-main)

Copied to clipboard

Challenge: Renewed interest in understanding long texts has sparked interest in benchmarks based on length of input text .
Approach: They propose a new metric that determines the shortest average text length that needs to be preserved to execute the task with limited performance degradation.
Outcome: The proposed benchmarks show that models outperform the previous generation on the QuALITY task due to their limited understanding of long-range dependencies.
Measuring the Knowledge Acquisition-Utilization Gap in Pretrained Language Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent research has demonstrated that pre-trained language models acquire a broad range of knowledge about linguistic structures, encyclopedic relations, levels of commonsense, and even coding and reasoning rules.
Approach: They propose a systematic framework to measure parametric knowledge utilization in pre-trained language models by extracting parametric information from a PLM and constructing a downstream task around this extracted knowledge.
Outcome: The proposed framework extracts parametric knowledge from a PLM and constructs a downstream task around this extracted knowledge.
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are prohibitive to use under resource constraints due to their high latency and high latex.
Approach: They propose to use a contextual bandit to help choose a model based on a context to improve performance.
Outcome: The proposed model can be used to improve performance on multiple domains even without prior knowledge of the model.
Local Structure Matters Most in Most Languages (2022.aacl-short)

Copied to clipboard

Challenge: Recent perturbation studies have found unintuitive results on what does and does not matter when performing Natural Language Understanding (NLU) tasks in English.
Approach: They replicate a study on the importance of local structure and relative unimportance of global structure in a multilingual setting.
Outcome: The proposed model replicates a study on the importance of local structure and relative unimportance of global structure in a multilingual setting.
Do Neural Dialog Systems Use the Conversation History Effectively? An Empirical Study (P19-1)

Copied to clipboard

Challenge: Neural generative models are becoming more popular when building conversational agents.
Approach: They propose to study the sensitivity of neural dialog models to unnatural perturbations . they experiment with 10 different types of perturbations on 4 multi-turn dialog datasets .
Outcome: The proposed model is sensitive to unnatural changes or perturbations on 4 multi-turn dialog datasets.
Combining Domain and Alignment Vectors Provides Better Knowledge-Safety Trade-offs in LLMs (2025.acl-short)

Copied to clipboard

Challenge: Large language models (LLMs) excel in specific technical fields, but are not explicitly trained to be safe.
Approach: They propose a model merging-based alignment method that allows for safer domain-specific models that preserve their utility.
Outcome: The proposed method improves safety alignment on LLMs with minimal degradation on domain-specific benchmarks.
Why Don’t Prompt-Based Fairness Metrics Correlate? (2024.acl-long)

Copied to clipboard

Challenge: Existing methods to assess fairness using prompts have low correlations between fairness metrics.
Approach: They propose a method to enhance the correlation between fairness metrics by using pre-trained language models.
Outcome: The proposed method improves the correlation between fairness metrics by using pre-trained language models.
EpiK-Eval: Evaluation for Language Models as Epistemic Models (2023.emnlp-main)

Copied to clipboard

Challenge: Developing systems that can reason through language understanding has been a cornerstone in natural language processing research.
Approach: They propose a question-answering benchmark to evaluate LLMs' ability to combine knowledge from different training documents within their parameter space.
Outcome: The proposed benchmark aims to evaluate LLMs' ability to combine knowledge from different training documents within their parameter space.
A Deep Dive into the Trade-Offs of Parameter-Efficient Preference Alignment Techniques (2024.acl-long)

Copied to clipboard

Challenge: Large language models are pre-trained on trillions of tokens and instruction-tuned or aligned to specific preferences.
Approach: They propose guidelines to help researchers perform more effective parameter-efficient LLM alignment.
Outcome: The proposed methods outperform preference optimization and outperformed pre-trained models on three key axes.
Small Encoders Can Rival Large Decoders in Detecting Groundedness (2025.findings-acl)

Copied to clipboard

Challenge: Large language models struggle to answer queries reliably when the provided context lacks information, often resorting to ungrounded speculation or internal knowledge.
Approach: They propose to detect whether a given query is grounded in a document provided in context before LLMs generate answers.
Outcome: The proposed model can generate answers that are grounded in the document provided in context while reducing inference latency by orders of magnitude.
A Survey of Data Augmentation Approaches for NLP (2021.findings-acl)

Copied to clipboard

Challenge: Data augmentation is a field of research that has been underexplored due to the discrete nature of language data.
Approach: They present a comprehensive survey of data augmentation for NLP by summarizing the literature in a structured manner.
Outcome: The proposed methods are used for popular NLP applications and tasks and highlight current challenges and directions for future research.
Are self-explanations from Large Language Models faithful? (2024.findings-acl)

Copied to clipboard

Challenge: Instruction-tuned Large Language Models excel at many tasks and will explain their reasoning, so-called self-explanations.
Approach: They propose to employ self-consistency checks to measure faithfulness to LLMs to determine if they are model-dependent and if their reasoning is convincing and wrong.
Outcome: The proposed measures show that self-explanations are explanation, model, and task-dependent and should not be trusted in general.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations