Papers by Vedant Nanda

    1 papers
    The Impact of Inference Acceleration on Bias of LLMs (2025.naacl-long)

    Copied to clipboard

    Challenge: Recent work suggests strategies to increase inference efficiency with LLMs . however, these strategies may inadvertently lead to some side-effects.
    Approach: They propose to optimize inference acceleration strategies such as quantization, pruning, and caching to reduce inference cost and latency while maintaining predictive performance.
    Outcome: The proposed strategies reduce cost and latency while maintaining predictive performance while preserving the model size.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations