JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention

Abstract

We combine JEPA-based masked prediction with density-adaptive attention to learn speech features without waveform reconstruction in the first stage. A second stage quantizes those features into compact tokens and reconstructs audio, connecting predictive representations with speech compression.

Publication
UniReps: Unifying Representations in Neural Models (NeurIPS 2025 Workshop)
Ravid Shwartz Ziv
Ravid Shwartz Ziv
AI Researcher

AI researcher at Meta MSL with a background in information theory and computational neuroscience, working on world models, memory, and compression. Former Assistant Professor and Faculty Fellow at NYU. My industry experience spans Wand AI, Intel, Google AI, and Wikipedia.

Related