The base comparison is the part that holds up. arc/c 0.647 to 0.711 on ARC-C's 1172 test items is about 4.8 standard errors. That is a real move, not harness noise.
The 700 line is where I would not put weight. mxfp8 0.711 and mxfp4 0.701 is 1.0pp, roughly 12 questions out of 1172. A single ARC-C score at p near 0.71 carries a standard error of about 1.32pp, so the threshold the model is named after sits inside the gap between your own two quants. Both quants also answered the same 1172 items, so the honest test there is paired, McNemar on the discordant ones, not two numbers compared by eye. Do you have per-item outputs to run it?
The other thing I cannot untangle from the table: it compares fine tune plus merge plus heretic against plain base, so no stage gets credit for anything. The heretic step is the strange one. Refusals go 99/100 to 4/100 at KL 0.0469 and all seven benchmarks go up. If the harness scores by loglikelihood over the fixed options, refusal direction should be close to invisible, because the model never gets to emit a refusal in the first place. That points the gain at the fine tune stages rather than the ablation.
Is there a bench of base plus heretic only, no fine tunes? That one arm separates them, and it is cheap.