Representation Learning

World Models and Predictive Representations

Ravid Shwartz Ziv's research on world models, predictive representations, learned dynamics, and training agents in imagination.

HP-JEPA: Hierarchical Partitioning for Multi-Resolution Graph Joint-Embedding Predictive Learning

Learning graph representations at multiple scales through latent prediction over coarse and fine partitions.

S-JEPA: Soft Clustering Anchors for Self-Supervised Speech Representation Learning

Learning speech representations with soft predictive targets, preserving acoustic ambiguity while avoiding repeated offline reclustering.

Soft Clustering Anchors for Self-Supervised Speech Representation Learning in Joint Embedding Prediction Architectures

An earlier speech JEPA study using fixed soft clustering targets to stabilize representation learning.

JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention

Connecting JEPA representations to speech compression through density-adaptive attention and compact audio tokens.

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations

Compressing an LLM for the task it actually needs to perform, allocating precision to the layers that matter.

Exploring Human-AI Conceptual Alignment through the Prism of Chess

Using chess and Chess960, we investigate how strategic concepts appear across model layers and how stronger play can diverge from human conceptual understanding.

Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin

Connecting attention sinks to representation compression to explain how information changes across LLM layers.

Layer by Layer: Uncovering Hidden Representations in Language Models

Finding useful embeddings inside language models. We study why intermediate layers can outperform the final layer and how to identify them using information, geometry, and downstream evaluation.

Layer Importance for Mathematical Reasoning is Forged in Pre-Training and Invariant after Post-Training

Studying which layers support mathematical reasoning and how their roles persist after post-training.