The Illusion of AI Expertise Under Uncertainty: Navigating Elusive Ground Truth via a Probabilistic Paradigm

Abstract

This paper studies how disagreement in reference labels can obscure differences between expert and non-expert performance. Expected accuracy and F1 scores provide a probabilistic view of evaluation, motivating comparisons stratified by the certainty of the reference answers.

Publication
arXiv preprint
Ravid Shwartz Ziv
Ravid Shwartz Ziv
AI Researcher

AI researcher at Meta MSL with a background in information theory and computational neuroscience, working on world models, memory, and compression. Former Assistant Professor and Faculty Fellow at NYU. My industry experience spans Wand AI, Intel, Google AI, and Wikipedia.

Related