Evaluation

Mirage Probes: How Vision Models Fake Visual Understanding

Probing whether vision-language models answer from visual evidence, language priors, or spurious visual representations.

Latent Transfer Attack: Adversarial Examples via Generative Latent Spaces

Using generative latent spaces to study adversarial transfer and the robustness of vision models.

Do Multi-Agents Dream of Electric Screens? Achieving Perfect Accuracy on AndroidWorld Through Task Decomposition

We study a mobile AI agent that separates planning, execution, verification, and reflection to improve reliability on AndroidWorld tasks.

The Illusion of AI Expertise Under Uncertainty: Navigating Elusive Ground Truth via a Probabilistic Paradigm

Accounting for ambiguous ground truth when comparing AI systems with experts.

A superpersuasive autonomous policy debating system

A research system that combines retrieval and specialized agents for structured policy debate.

Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests

We develop a framework for evaluating AI behavior with realistic situational judgment tests and structured personas, connecting AI evaluation with psychometrics.

The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs

Reassessing hallucination detection with evaluation metrics that better reflect meaning.

LiveBench: A Challenging, Contamination-Limited LLM Benchmark

An evolving language-model benchmark with recent questions and objective scoring, designed to limit test contamination.