GMM-Anchored JEPA fits soft clusters to acoustic features and uses them as auxiliary targets during self-supervised speech training. A changing supervision weight balances this grounding signal with latent prediction, avoiding iterative offline reclustering.
Related work: S-JEPA is the later speech study with a continuous two-phase training procedure and online clustering targets.