
ArticleAI CodingLocal AI
MXFP8 vs Q8: 10x the weight error, 1% the perplexity
I quantised real Qwen3-Coder weights both ways on an M3 Ultra. MXFP8 reconstructs them about 10x worse than 8-bit affine at identical size — and then costs only 1% perplexity end to end. Both numbers are true, and the gap between them is the interesting part.