Abstract
We connect attention sinks and compression valleys to large activations in the residual stream. Theory and ablations motivate a mix-compress-refine account of computation across model depth, helping explain why intermediate representations can suit embedding tasks while generation still uses later layers.
Publication
International Conference on Learning Representations

AI Researcher
AI researcher at Meta MSL with a background in information theory and computational neuroscience, working on world models, memory, and compression. Former Assistant Professor and Faculty Fellow at NYU, collaborating on research across academia and industry.