Learning graph representations at multiple scales through latent prediction over coarse and fine partitions.
Learning speech representations with soft predictive targets, preserving acoustic ambiguity while avoiding repeated offline reclustering.
An earlier speech JEPA study using fixed soft clustering targets to stabilize representation learning.
Connecting JEPA representations to speech compression through density-adaptive attention and compact audio tokens.