S-JEPA: Soft Clustering Anchors for Self-Supervised Speech Representation Learning

Abstract

S-JEPA learns speech representations by predicting soft Gaussian-mixture targets at masked positions. Training moves from acoustic features to online targets built from learned representations, avoiding repeated offline reclustering and preserving uncertainty where speech categories overlap.

Publication
arXiv preprint

Related work: the earlier GMM-Anchored JEPA study uses fixed soft acoustic clusters to support speech representation learning.

Ravid Shwartz-Ziv
Ravid Shwartz-Ziv
AI Researcher

AI researcher at Meta MSL with a background in information theory and computational neuroscience, working on world models, memory, and compression. Former Assistant Professor and Faculty Fellow at NYU, collaborating on research across academia and industry.

Related