Nunchaku Lite: 4-bit Inference of Diffusion Models in Diffusers
Nunchaku is a quantization method SVDQuant (W4A4) for diffusion transformers. The new Nunchaku Lite integration allows loading quantized checkpoints directly via from_pretrained() in Diffusers, without a separate engine. Generation speed increases up to 1.8x, peak VRAM consumption drops up to 50%.
Hugging Face
Baidu
Hugging Face blog24.07 · 02:04

