From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning

Abstract

Using the information bottleneck framework, we compare human conceptual categories with language-model embeddings. The study examines the tension between compression and semantic nuance, differences between encoder and decoder representations, and how conceptual information moves across layers during training.

Publication
International Conference on Learning Representations
Ravid Shwartz-Ziv
Ravid Shwartz-Ziv
AI Researcher

AI researcher at Meta MSL with a background in information theory and computational neuroscience, working on world models, memory, and compression. Former Assistant Professor and Faculty Fellow at NYU, collaborating on research across academia and industry.

Related