Using chess and Chess960, we investigate how strategic concepts appear across model layers and how stronger play can diverge from human conceptual understanding.
Finding useful embeddings inside language models. We study why intermediate layers can outperform the final layer and how to identify them using information, geometry, and downstream evaluation.