Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin

Abstract

We connect attention sinks and compression valleys to large activations in the residual stream. Theory and ablations motivate a mix-compress-refine account of computation across model depth, helping explain why intermediate representations can suit embedding tasks while generation still uses later layers.

Publication
International Conference on Learning Representations
Ravid Shwartz-Ziv
Ravid Shwartz-Ziv
AI Researcher

AI researcher at Meta MSL with a background in information theory and computational neuroscience, working on world models, memory, and compression. Former Assistant Professor and Faculty Fellow at NYU, collaborating on research across academia and industry.

Related