Facebook (Md Ismail Sojal)
local-llmai-tools
Original source

Qwen3.8-27B QUASAR-QAT NVFP4 — 65% Smaller, Near-Original Performance

Summary

  • Qwen3.8-27B quantized from 55.6GB to 19.7GB (65% smaller) using QAT (Quantization-Aware Training) + distillation to NVFP4 format
  • Nearly 100% of original reasoning & math ability retained
  • Beats traditional PTQ (Post-Training Quantization) by a clear margin
  • Native vLLM/SGLang support — plug and play
  • 4-bit quantization aiming for near-original performance

Community Benchmarks (from comments)

  • 16GB GPU: 50-100+ tok/s, 70-75% MMLU score (Gavin Nunns)
  • RTX 5060 16GB: 35-50 tps decode, 500-700 tps prefill with 64k context using UD_IQ4_XS format (Aleksei Martemianov)
  • One commenter noted IQ4_NL models may hold up better than QAT in some tests (Donn Lasher)