Papers by Simone Scardapane

2 papers
A Simple and Effective L_2 Norm-Based Strategy for KV Cache Compression (2024.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to reduce the KV cache size involve fine-tuning the model to learn a compression strategy or leveraging attention scores to reduce sequence length.
Approach: They find a correlation between the L2 norm and attention scores over cached KV pairs . they compress the KV cache based on the L1 norm of key embeddings .
Outcome: The proposed approach reduces the KV cache size by 50% on language modelling and needle-in-a-haystack tasks and 90% on passkey retrieval tasks without losing accuracy.
Attention Sinks in Diffusion Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Masked Diffusion Language Models (DLMs) employ transformer encoders with bidirectional attention, enabling parallel token generation while maintaining competitive performance.
Approach: They conduct an empirical analysis of DLM attention patterns focusing on the attention sinking phenomenon . they find that DLMs also exhibit attention sinks, but with distinct characteristics .
Outcome: The proposed models employ transformer encoders with bidirectional attention, enabling parallel token generation while maintaining competitive performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations