Abstract
The paper compares lexical-overlap evaluation with human judgments and semantic assessment of hallucinations. It finds that apparent detector performance can depend strongly on the evaluation metric and on simple response-length effects, motivating more careful reliability benchmarks.
Publication
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing

AI Researcher
AI researcher at Meta MSL with a background in information theory and computational neuroscience, working on world models, memory, and compression. Former Assistant Professor and Faculty Fellow at NYU, collaborating on research across academia and industry.