Abstract
Intermediate language-model layers can provide stronger representations than the final layer. We study this using information theory, geometry, and downstream evaluation.
Publication
Proceedings of the 42nd International Conference on Machine Learning

AI Researcher
AI researcher at Meta MSL with a background in information theory and computational neuroscience, working on world models, memory, and compression. Former Assistant Professor and Faculty Fellow at NYU, collaborating on research across academia and industry.