Interpretability

Mirage Probes: How Vision Models Fake Visual Understanding

Probing whether vision-language models answer from visual evidence, language priors, or spurious visual representations.

Exploring Human-AI Conceptual Alignment through the Prism of Chess

Using chess and Chess960, we investigate how strategic concepts appear across model layers and how stronger play can diverge from human conceptual understanding.

Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin

Connecting attention sinks to representation compression to explain how information changes across LLM layers.

Layer Importance for Mathematical Reasoning is Forged in Pre-Training and Invariant after Post-Training

Studying which layers support mathematical reasoning and how their roles persist after post-training.