JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention

Abstract

We combine JEPA-based masked prediction with density-adaptive attention to learn speech features without waveform reconstruction in the first stage. A second stage quantizes those features into compact tokens and reconstructs audio, connecting predictive representations with speech compression.

Publication
UniReps: Unifying Representations in Neural Models (NeurIPS 2025 Workshop)
Ravid Shwartz-Ziv
Ravid Shwartz-Ziv
AI Researcher

AI researcher at Meta MSL with a background in information theory and computational neuroscience, working on world models, memory, and compression. Former Assistant Professor and Faculty Fellow at NYU, collaborating on research across academia and industry.

Related