Probing whether vision-language models answer from visual evidence, language priors, or spurious visual representations.
Using generative latent spaces to study adversarial transfer and the robustness of vision models.
We study a mobile AI agent that separates planning, execution, verification, and reflection to improve reliability on AndroidWorld tasks.
Accounting for ambiguous ground truth when comparing AI systems with experts.
A research system that combines retrieval and specialized agents for structured policy debate.
We develop a framework for evaluating AI behavior with realistic situational judgment tests and structured personas, connecting AI evaluation with psychometrics.
Reassessing hallucination detection with evaluation metrics that better reflect meaning.
An evolving language-model benchmark with recent questions and objective scoring, designed to limit test contamination.