Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests

Abstract

We develop a framework for evaluating AI behavior with realistic situational judgment tests and structured personas, connecting AI evaluation with psychometrics.

Publication
arXiv preprint
Ravid Shwartz Ziv
Ravid Shwartz Ziv
AI Researcher

AI researcher at Meta MSL with a background in information theory and computational neuroscience, working on world models, memory, and compression. Former Assistant Professor and Faculty Fellow at NYU. My industry experience spans Wand AI, Intel, Google AI, and Wikipedia.

Related