Abstract
We study policies trained on imagined trajectories from learned dynamics and reward models. The analysis connects model error to return estimates and policy optimization, and examines how to allocate data and rollout budgets when reward signals vary in cost and noise.
Publication
arXiv preprint

AI Researcher
AI researcher at Meta MSL with a background in information theory and computational neuroscience, working on world models, memory, and compression. Former Assistant Professor and Faculty Fellow at NYU, collaborating on research across academia and industry.