S-JEPA learns speech representations by predicting soft Gaussian-mixture targets at masked positions. Training moves from acoustic features to online targets built from learned representations, avoiding repeated offline reclustering and preserving uncertainty where speech categories overlap.
Related work: the earlier GMM-Anchored JEPA study uses fixed soft acoustic clusters to support speech representation learning.