1

Mirage Probes: How Vision Models Fake Visual Understanding

Probing whether vision-language models answer from visual evidence, language priors, or spurious visual representations.

JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention

Connecting JEPA representations to speech compression through density-adaptive attention and compact audio tokens.

A superpersuasive autonomous policy debating system

A research system that combines retrieval and specialized agents for structured policy debate.

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations

Compressing an LLM for the task it actually needs to perform, allocating precision to the layers that matter.

Exploring Human-AI Conceptual Alignment through the Prism of Chess

Using chess and Chess960, we investigate how strategic concepts appear across model layers and how stronger play can diverge from human conceptual understanding.

Antislop: A Comprehensive Framework for Identifying and Eliminating Repetitive Patterns in Language Models

Detecting and reducing repetitive language patterns through sampling and targeted fine-tuning.

Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin

Connecting attention sinks to representation compression to explain how information changes across LLM layers.

The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs

Reassessing hallucination detection with evaluation metrics that better reflect meaning.

Layer by Layer: Uncovering Hidden Representations in Language Models

Finding useful embeddings inside language models. We study why intermediate layers can outperform the final layer and how to identify them using information, geometry, and downstream evaluation.

Layer Importance for Mathematical Reasoning is Forged in Pre-Training and Invariant after Post-Training

Studying which layers support mathematical reasoning and how their roles persist after post-training.