Papers by Anton Korznikov
Out of Distribution, Out of Luck: Process Rewards Misguide Reasoning Models (2026.eacl-short)
Copied to clipboard
| Challenge: | 80% of reasoning model outputs respond to formatting artifacts rather than mathematical content. |
| Approach: | They evaluate process reward models that provide step-level feedback during inference . they identify distinct reward prediction patterns that differentiate reasoning from non-reasoning model outputs . |
| Outcome: | The proposed model fails to enhance and sometimes degrade reasoning model performance. |
Feature Drift: How Fine-Tuning Repurposes Representations in LLMs (2026.findings-eacl)
Copied to clipboard
| Challenge: | Sparse autoencoders (SAEs) are a powerful tool for interpreting neural networks by extracting concepts (features) represented in their activations. |
| Approach: | They propose to use Sparse Autoencoders to extract concepts from their activations to explain how fine-tuning changes model capabilities. |
| Outcome: | The proposed model recombines existing concepts rather than learning new ones, and shows that it is a better explanation for how fine-tuning changes model capabilities. |