Soft Clustering Anchors for Self-Supervised Speech Representation Learning in Joint Embedding Prediction Architectures

Abstract

GMM-Anchored JEPA fits soft clusters to acoustic features and uses them as auxiliary targets during self-supervised speech training. A changing supervision weight balances this grounding signal with latent prediction, avoiding iterative offline reclustering.

Publication
arXiv preprint

Related work: S-JEPA is the later speech study with a continuous two-phase training procedure and online clustering targets.

Ravid Shwartz Ziv
Ravid Shwartz Ziv
AI Researcher

AI researcher at Meta MSL with a background in information theory and computational neuroscience, working on world models, memory, and compression. Former Assistant Professor and Faculty Fellow at NYU. My industry experience spans Wand AI, Intel, Google AI, and Wikipedia.

Related