Summary
- Qwen3.8-27B quantized from 55.6GB to 19.7GB (65% smaller) using QAT (Quantization-Aware Training) + distillation to NVFP4 format
- Nearly 100% of original reasoning & math ability retained
- Beats traditional PTQ (Post-Training Quantization) by a clear margin
- Native vLLM/SGLang support — plug and play
- 4-bit quantization aiming for near-original performance
- 16GB GPU: 50-100+ tok/s, 70-75% MMLU score (Gavin Nunns)
- RTX 5060 16GB: 35-50 tps decode, 500-700 tps prefill with 64k context using UD_IQ4_XS format (Aleksei Martemianov)
- One commenter noted IQ4_NL models may hold up better than QAT in some tests (Donn Lasher)