Parable

🪶 Parable-Qwen3-4B — trained on genuine Claude Fable 5 agent traces

A small local model with agent instincts — planning, tool use, and terminal habits distilled from real agent sessions, not synthetic Q&A. v2: every prompt gets an answer, eval-gated before publish.

~4 GB of RAM is all you need. Laptop, old GPU, modest desktop — the Q4 build runs anywhere. One command and you have a private, offline reasoning model on your machine:

ollama run hf.co/AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF:Q4_K_M

v2 is here (2026-07-22)

Every prompt now gets an answer. v1 inherited Qwen3's runaway thinking: on 21% of ordinary prompts it spent its entire token budget reasoning and returned nothing. v2 fixes that — 34/34 prompts answered, answers roughly half as long, coding ability held.

Same repo, same links, same commands. Pull again and you have it. v1 remains in the version history below.

Full family. This 4B sits between its Parable siblings — a 3B and two 8Bs. Browse the full collection.


Pick your size

File Size Fits in Notes
Q4_K_M 2.5 GB ~4 GB RAM/VRAM Recommended — best size/quality balance
Q5_K_M 2.9 GB ~4.5 GB Higher quality
Q6_K 3.3 GB ~5 GB Near-lossless
Q8_0 4.3 GB ~6 GB Maximum quality

Full-precision safetensors (vLLM, transformers, further fine-tuning): Parable-Qwen3-4B-Claude-Fable-5

How to run it

Ollama (shorter command via the Parable namespace, or pull straight from this repo):

ollama run parable/qwen3-fable:4b
# or straight from this repo:
ollama run hf.co/AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF:Q4_K_M

llama.cpp:

llama-cli -m Parable-Qwen3-4B-Claude-Fable-5-GGUF-Q4_K_M.gguf --jinja \
  -p "Write a bash one-liner to find the 10 largest files in a directory tree."

LM Studio: lms get parable/qwen3-fable, or search "parable" in-app (parable on LM Studio Hub).

Python (llama-cpp-python):

from llama_cpp import Llama

llm = Llama.from_pretrained(
    repo_id="AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF",
    filename="*Q4_K_M.gguf", n_ctx=8192,
)
out = llm.create_chat_completion(
    messages=[{"role": "user", "content": "Write a Python function that retries an HTTP request with exponential backoff."}],
    max_tokens=3000, temperature=0.3,
)
print(out["choices"][0]["message"]["content"])

Answers first

v2 responds directly — no <think> preamble to strip, no reasoning budget to babysit. The planning instincts from the agent traces are still there; they just don't get narrated at you first.

Sampling: temperature 0.3–0.7, max_tokens 800–1500 is plenty for most prompts. (v1 needed ≥2500 and could still time out mid-thought. That's over.)


The numbers

Two builds, one harness, greedy decoding, HumanEval+ scored with EvalPlus on identical hardware:

Qwen3-4B base Parable v1 Parable v2
Prompts answered (34-prompt suite) 27/34 34/34
HumanEval 76.8 72.0 73.2
HumanEval+ 71.3 68.3 67.1
Held-out trace loss 2.846 1.876 (−34%)

The headline is the first row. On 7 of 34 ordinary prompts — "write a retry function with exponential backoff", "write a bash script that watches a log file" — v1 burned its whole budget thinking and returned an empty response. v2 returns 0 empty responses, using 140× less reasoning text. Coding scores are unchanged within noise, so this is a reliability fix, not a capability trade.

Against the base model, v2 trails by 3.6 HumanEval points — the cost of specializing on agent traces, and about a third of the regression the same recipe produces at 3B. Trace loss drops 34% on held-out sessions: it learned the agent style without paying the usual forgetting tax.

Reproducibility: two training seeds land within 0.001 trace loss of each other (1.877 / 1.876); we ship the second and publish both in parable-v2-artifacts. Benchmarks run with thinking disabled on both base and fine-tune — the base is a hybrid-thinking model and scores near zero in default mode because it never emits a final answer. Comparing against that would have flattered us.

Training data

v2 trains on corpus v2 — 10.6k curated traces (13× v1) with completion-only loss masking, a 30% general-instruction replay mix, and span-filtered truncation. Every example passed a quality gate (schema validation, secrets scrub, length filtering). QLoRA fine-tune (TRL), quantized with llama.cpp.

Good to know

  • Answers directly rather than reasoning aloud. If you want visible chain-of-thought, prompt for it explicitly ("think step by step, show your work").
  • Tuned hard toward agentic coding behavior; that focus trades some general-knowledge breadth, as with any specialized fine-tune in this class.
  • Verify critical output. Small models over-commit to plausible specifics; treat generated commands and code as drafts to review.
  • Inherits Qwen3-4B's base limitations and knowledge cutoff.

Base & license

Weights: Apache-2.0 (inherited from Qwen/Qwen3-4B). Training data: Fable-5-traces AGPL-3.0, gpt5.5-terminal MIT — since those traces originate from third-party assistants, the providers' terms may apply to downstream training and distillation; confirm your use aligns with them before building on this model commercially.

Get Parable

Platform
Ollama ollama run parable/qwen3-fable:4b · parable namespace
Ollama (family flagship, best per size) ollama run parable/fable
Hugging Face full collection — GGUF quants, full weights, eval reports
LM Studio lms get parable/qwen3-fable · parable on LM Studio Hub

Acknowledgements

Glint-Research & Roman1111111 for the open trace data · Qwen team for the base · empero-ai whose Qwable recipe this release follows · mlx-lm & llama.cpp


Four gigabytes of RAM. Real Fable 5 reasoning. Yours, offline, right now.

ollama run hf.co/AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF:Q4_K_M

Version history

  • v2 (2026-07-22) — rebuilt recipe: 10.6k-trace corpus v2, completion-only loss masking, 30% replay mix, span-filtered truncation. Fixes v1's empty-response failure mode (7/34 → 0/34). Coding parity with v1, trace loss −34% vs base.
  • v1 (2026-07-11) — initial release: QLoRA on 4.4k Fable-5 traces via mlx-lm.

More on the Parable models: ankitaglawe.com/parable

Downloads last month
3,796
GGUF
Model size
4B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF

Finetuned
Qwen/Qwen3-4B
Finetuned
(1030)
this model

Datasets used to train AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF

Collection including AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF