Task-Aware Quantization uses hidden representations and output sensitivity to allocate mixed precision across transformer layers under a fixed bit budget. Calibration uses a small set of task prompts, with no weight training, and the study measures accuracy, memory use, throughput, and latency.
The workshop version is identified in the arXiv record. The original preprint was released in November 2025; the latest revision was submitted in June 2026.