Seq-VCR: Preventing Collapse in Intermediate Transformer Representations for Enhanced Reasoning

Abstract

Seq-VCR applies variance and covariance regularization to intermediate transformer representations. Experiments with arithmetic and sequential reasoning tasks study how preventing representation collapse, together with pause tokens, can improve reasoning without explicit chain-of-thought supervision.

Publication
International Conference on Learning Representations
Ravid Shwartz Ziv
Ravid Shwartz Ziv
AI Researcher

AI researcher at Meta MSL with a background in information theory and computational neuroscience, working on world models, memory, and compression. Former Assistant Professor and Faculty Fellow at NYU. My industry experience spans Wand AI, Intel, Google AI, and Wikipedia.

Related