Layer Importance for Mathematical Reasoning is Forged in Pre-Training and Invariant after Post-Training

Abstract

Layer ablations compare base models with instruction-tuned, reinforcement-learned, and distilled variants. The study finds a small set of layers important for mathematical tasks whose importance remains stable across the examined post-training methods, alongside changes in token representations.

Publication
NeurIPS 2025 MATH-AI: The 5th Workshop on Mathematical Reasoning and AI
Ravid Shwartz Ziv
Ravid Shwartz Ziv
AI Researcher

AI researcher at Meta MSL with a background in information theory and computational neuroscience, working on world models, memory, and compression. Former Assistant Professor and Faculty Fellow at NYU. My industry experience spans Wand AI, Intel, Google AI, and Wikipedia.

Related