Real Atlassian solutions to real problems — no fluff, no SEO spam.
NVFP4 has a second scaling level. mlx-lm never sets it, so NVFP4 converts to the worst 4-bit result on the table — worse than a format using fewer bits. Set it and NVFP4 wins. I nearly published this as a defect in NVIDIA's spec; it was my own unit error.
Metal will not give you the RAM on the box, and the number it does give is not the 75% everyone repeats. I measured the three ceilings on a 96 GB Mac Studio, then measured what modern hybrid-attention models actually spend against them — including a Gemma 4 cache that quietly holds three times its own sliding window.
I quantised real Qwen3-Coder weights both ways on an M3 Ultra. MXFP8 reconstructs them about 10x worse than 8-bit affine at identical size — and then costs only 1% perplexity end to end. Both numbers are true, and the gap between them is the interesting part.