Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests

Abstract

We develop a framework for evaluating AI behavior with realistic situational judgment tests and structured personas, connecting AI evaluation with psychometrics.

Publication
arXiv preprint
Ravid Shwartz-Ziv
Ravid Shwartz-Ziv
AI Researcher

AI researcher at Meta MSL with a background in information theory and computational neuroscience, working on world models, memory, and compression. Former Assistant Professor and Faculty Fellow at NYU, collaborating on research across academia and industry.

Related