본문으로 건너뛰기

Quantization

Reducing the numeric precision of model weights (e.g. to 4-bit) to shrink size and run models on less hardware.

관련 리소스