Recently re-quantized with the latest version of advanced-gguf-quantizer tool.

MTP included.

Qwen3.6-27B RSF NVFP4 GGUF v4

This latest (6-June-2026) NVFP4 quant of Qwen3.6-27B keeps quality very high, while improving KLD, tail KLD metrics, top-token stability, and probability error compared to the earlier releases. It is faster and smaller in size. It applies my RSF scale fitting technique to the Q_K quants and uses a different tensor mix layout compared to previously. 6-June-2026: changed the MTP tensors to be NVFP4, ~130tk/s tg is now possible in some configurations on 5090.

I am experimenting with a smaller verison of this model designed to be used on machines with 16GB of VRAM:
Qwen3.6-27B-NVFP4-SMALL-MTP-GGUF.
The SMALL version will be slower than this model but should maintain similar quality, although I am still evaluating it.

Feedback to improve this is appreciated.

Results below on RTX 5090 with Wikitest:
Model Size GiB pp512 tg128 pg32768,256
RSF NVFP4 15.27 5174.31 ± 1.30 76.61 ± 0.18 2751.93 ± 23.73

Quality Metric Earlier NVFP4 Improved RSF NVFP4 Change
Mean PPL(Q) 7.128133 ± 0.047613 7.030348 ± 0.046636 1.37% lower
Mean PPL(base) 6.900856 ± 0.045374 6.900856 ± 0.045374 (Same Base)
PPL gap vs base 0.227277 ± 0.008252 0.129492 ± 0.006913 43.0% smaller gap
Mean PPL(Q) / PPL(base) 1.032935 ± 0.001175 1.018765 ± 0.000996 43.0% lower overhead
Mean ln(PPL(Q) / PPL(base)) 0.032404 ± 0.001137 0.018591 ± 0.000978 42.6% lower
Correlation ln(PPL) 98.54% 98.91% +0.37 pp

KLD Metric Earlier v2 Latest RSF NVFP4 Change
Mean KLD 0.058781 ± 0.000935 0.044590 ± 0.000832 24.1% lower
99.9% KLD 4.869176 3.726478 23.5% lower
99.0% KLD 0.591435 0.422199 28.6% lower
95.0% KLD 0.165489 0.122647 25.9% lower
90.0% KLD 0.096133 0.070556 26.6% lower
Median KLD 0.019355 0.013642 29.5% lower
Maximum KLD 20.671190 24.703529 higher

Token Probability Metric Earlier NVFP4 Improved RSF NVFP4 Change
RMS Δp 6.613 ± 0.060% 5.774 ± 0.060% 12.7% lower
Same-top probability 90.450 ± 0.076% 91.924 ± 0.071% +1.474 pp
Top-flip weight 0.009715 0.007193 26.0% lower
Top-prob RMSE 0.080946 0.069932 13.6% lower
Entropy RMSE 0.248197 0.212250 14.5% lower
Downloads last month
43,643
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for michaelw9999/Qwen3.6-27B-NVFP4-MTP-GGUF

Base model

Qwen/Qwen3.6-27B
Quantized
(664)
this model

Collection including michaelw9999/Qwen3.6-27B-NVFP4-MTP-GGUF