Abstract
LiveBench evaluates language models using frequently refreshed questions and objective scoring across tasks such as reasoning, coding, mathematics, and instruction following. Updating the test material is intended to limit contamination and make comparisons more informative as models change.
Publication
International Conference on Learning Representations

AI Researcher
AI researcher at Meta MSL with a background in information theory and computational neuroscience, working on world models, memory, and compression. Former Assistant Professor and Faculty Fellow at NYU. My industry experience spans Wand AI, Intel, Google AI, and Wikipedia.