Abstract
This paper studies how disagreement in reference labels can obscure differences between expert and non-expert performance. Expected accuracy and F1 scores provide a probabilistic view of evaluation, motivating comparisons stratified by the certainty of the reference answers.
Publication
arXiv preprint

AI Researcher
AI researcher at Meta MSL with a background in information theory and computational neuroscience, working on world models, memory, and compression. Former Assistant Professor and Faculty Fellow at NYU, collaborating on research across academia and industry.