Papers by Jeffrey Olmo
Features that Make a Difference: Leveraging Gradients for Improved Dictionary Learning (2025.findings-naacl)
Copied to clipboard
| Challenge: | Sparse Autoencoders (SAEs) are a promising approach for extracting neural network representations by learning a sparse and overcomplete decomposition of the network’s internal activations. |
| Approach: | They propose a method that learns a sparse and overcomplete decomposition of the network's internal activations and a gradient approach to learn latents. |
| Outcome: | The proposed algorithms improve the performance of the k-sparse autoencoder and the ability to learn latent features. |