Neural Networks

When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models

Inheritune builds smaller language models by reusing useful layers and reducing attention collapse, while retaining or improving performance.

Seq-VCR: Preventing Collapse in Intermediate Transformer Representations for Enhanced Reasoning

Regularizing intermediate representations to reduce collapse and support sequential reasoning.

An Information Theory Perspective on Variance-Invariance-Covariance Regularization

We provide an information-theoretic analysis of VICReg, deriving theoretical foundations for deterministic networks and introducing new SSL methods based on these insights.

Back to Basics: Revisiting Standard Deep Learning Components for Class Imbalance

We show that carefully tuning standard deep learning components can achieve state-of-the-art performance on class-imbalanced datasets without specialized techniques.

To Compress or Not to Compress--Self-Supervised Learning and Information Theory: A Review

We present a comprehensive review of self-supervised learning through the lens of information theory, introducing a unified framework that encompasses existing approaches and highlighting the interplay between compression and information preservation in deep neural networks.