Qwen3.6-35B-A3B Heretic — oQ3 (3-bit, MLX)

A sensitivity-guided ~3-bit (oQ3) quant of the abliterated Qwen3.6-35B-A3B "Heretic" (Native-MTP-Preserved) model, built on-device with omlx's mixed-precision oQ quantizer. MoE arch qwen3_5_moe (35B total / ~3B active). Apple-Silicon MLX format.

  • Effective precision: ~3.6 bpw (≈16 GB weights) — oQ keeps sensitive layers higher-bit, so it punches above its nominal bit-width.
  • Abliterated / uncensored. Use responsibly; you are accountable for your outputs.

Why this quant

Despite being the smallest/fastest quant of the family, it matched the higher-bit builds on every benchmark tried (M4 Max, thinking modes as noted):

Test Score
Hard reasoning + code (8 tasks) 8/8
Harder quality (multi-digit math, DP) (6) 6/6
General knowledge (20 facts) 20/20
Multi-step agentic tool-use (5) 5/5
Agentic "gauntlet" — flaky-tool retry, traps, branch (7, thinking-OFF) 7/7
Throughput ~100+ tok/s

Run it

Serve with omlx (or any MLX-LM runtime) on Apple Silicon:

omlx serve --port 8000
# then request model "Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-oQ3"

Best as an agent default with thinking OFF (cleanest tool-discipline); flip thinking ON for hard multi-step reasoning.

Private quant for personal use. License inherits from the base model.

Downloads last month
1,519
Safetensors
Model size
5B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

3-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support