You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations

Abstract

Task-Aware Quantization uses hidden representations and output sensitivity to allocate mixed precision across transformer layers under a fixed bit budget. Calibration uses a small set of task prompts, with no weight training, and the study measures accuracy, memory use, throughput, and latency.

Publication
ICML 2026 Workshop on AdaptFM: Resource-Adaptive Foundation Model Inference

The workshop version is identified in the arXiv record. The original preprint was released in November 2025; the latest revision was submitted in June 2026.

Ravid Shwartz Ziv
Ravid Shwartz Ziv
AI Researcher

AI researcher at Meta MSL with a background in information theory and computational neuroscience, working on world models, memory, and compression. Former Assistant Professor and Faculty Fellow at NYU. My industry experience spans Wand AI, Intel, Google AI, and Wikipedia.

Related