Abstract
The paper compares lexical-overlap evaluation with human judgments and semantic assessment of hallucinations. It finds that apparent detector performance can depend strongly on the evaluation metric and on simple response-length effects, motivating more careful reliability benchmarks.
Publication
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing

AI Researcher
AI researcher at Meta MSL with a background in information theory and computational neuroscience, working on world models, memory, and compression. Former Assistant Professor and Faculty Fellow at NYU. My industry experience spans Wand AI, Intel, Google AI, and Wikipedia.