Skip to content

Quantization

Reducing the numeric precision of model weights (e.g. to 4-bit) to shrink size and run models on less hardware.

Related resources