Quantization

Model Compression and Efficient AI

Ravid Shwartz-Ziv's research on model compression, task-aware quantization, efficient representations, and reducing AI memory and computation.

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations

Compressing an LLM for the task it actually needs to perform, allocating precision to the layers that matter.