Speech

S-JEPA: Soft Clustering Anchors for Self-Supervised Speech Representation Learning

Learning speech representations with soft predictive targets, preserving acoustic ambiguity while avoiding repeated offline reclustering.

Soft Clustering Anchors for Self-Supervised Speech Representation Learning in Joint Embedding Prediction Architectures

An earlier speech JEPA study using fixed soft clustering targets to stabilize representation learning.

JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention

Connecting JEPA representations to speech compression through density-adaptive attention and compact audio tokens.