qwopus3.6-27b: MLX vs GGUF on Apple Silicon
The first benchmark run with goose local-edition and its swarm-gym harness. We had a local-model swarm build three real command-line apps, five times each, with two builds of the same model — qwopus3.6-27b as the original GGUF and as the MLX translation — on a 3-node Apple Silicon fleet via LM Studio. Identical frozen prompts, identical binary, identical flags. 30 app builds in total, every one graded by running it against golden values. Here is everything the runs produced.
The apps that got built
Three archetypes, chosen to stress different failure modes — data modelling, algorithms, and stateful logic. Each was built 5× per model build (10× total) and checked by running real commands against it:
Archetype | The app | What it does | A golden check (verified by running) |
|---|---|---|---|
crud-multiformat | invtrack | inventory store with JSON + table output and value aggregation | add 4 × $5 → “Total inventory value: $20.00” |
compute-parser | rdcalc | a recursive-descent arithmetic calculator | 2^3^2 = 512 (right-assoc); 2+3*4 = 14 |
transaction | txkvbench | a key/value store with nested BEGIN/ROLLBACK | nested rollback prints “2 1 1”; COUNT |
When a build passed, the app genuinely worked — we ran these commands on the finished binaries and got the golden answers back. When it failed, it failed one of these checks, no matter what the model claimed.
Speed
Median build time per run, per archetype — lower is faster:
Spec | GGUF median | MLX median | Faster |
|---|---|---|---|
compute (rdcalc) | 1860.5s | 2192.6s | GGUF, by 15% |
transaction (txkvbench) | 1945.8s | 1722.9s | MLX, by 13% |
crud (invtrack) | 2239.2s | 2342.9s | ~tie |
overall | 1945.8s | 1923.9s | ~tie |
But the medians hide the real story, which is consistency:
Consistency | GGUF | MLX |
|---|---|---|
Std deviation | 750.9s | 580.4s (tighter) |
p90 (slow tail) | 3600.7s | 3049.7s |
Range (min–max) | 1210s – 3601s | 1001s – 3093s |
Runs that hit the 60-min cap | 2 | 0 |
GGUF has a higher ceiling (its fastest run, a rdcalc build at 1210s, beat everything MLX did) but a much longer tail. MLX never once ran away.
Quality — did the apps actually work
Share of builds that passed every golden check, per archetype:
Spec | GGUF checks-pass | MLX checks-pass |
|---|---|---|
compute (rdcalc) | 5 / 5 (100%) | 4 / 5 (80%) |
crud (invtrack) | 4 / 5 (80%) | 5 / 5 (100%) |
transaction (txkvbench) | 3 / 5 (60%) | 4 / 5 (80%) |
overall | 12 / 15 (80%) | 13 / 15 (87%) |
MLX leads on raw checks, 87% to 80% — but two of GGUF's three misses were a correct app the harness mis-scored (a run capped mid-cleanup, and one transient false-partial that ran fine when we re-ran it). Judged on whether the software works, the two are even — ~13 of 15 each.
What worked, and what didn't
compute — GGUF's archetype
GGUF was flawless here: 5/5, and the fastest spec on the board (median 1860s). MLX went 4/5 — its one miss was a scheduler deadlock: the run stalled with two tasks stuck, the calculator's entry point never got written, and the app simply couldn't run. Not a parsing mistake, a swarm stall. Every other compute build, both formats, nailed right-associativity and precedence.
crud — MLX's archetype
MLX went 5/5 with the tightest variance of any cell (stdev 442s) — steady, correct inventory stores every time. GGUF built five working apps too (we verified the totals by running them), but two of the five hit the 60-minute cap and one scored a false-partial. Those caps are why GGUF's crud variance is nearly 2× MLX's — and, as it turns out, they weren't the model's fault (see below).
transaction — the hard one, for both
This was the weakest archetype for both formats, and it's the most interesting. The nested-rollback logic was usually right; the multi-command exec path is where it broke. MLX had one build whose exec GET printed nothing. GGUF had it worse: one build cascaded (3 of 4 tasks failed) and one simply never implemented the COUNT command — yet its own unit tests passed, so the model shipped it green. Same weakness, both formats: a weak-model completeness gap on the hardest spec, and the single clearest reason to grade by running instead of trusting the model's tests.
Three things running the apps revealed
1. GGUF's caps were a swarm bug, not the model
Both of GGUF's capped crud runs had already built a correct app — the harness caught the swarm's own integrate-verify step churning after the fact, with no wall-clock budget to stop it. That's a scheduler bug we've since fixed (an env-gated cap on that step) and are re-benchmarking. So GGUF's crud numbers are penalised for something that isn't the model.
2. Passing tests, broken feature — on both builds
The transaction misses shared a signature across GGUF and MLX: the model's generated unit tests passed while a required feature was broken or missing. Because it showed up on both formats, it's a model limit, not a runtime difference — and the harness's golden checks caught every instance the model's tests missed.
3. Raw pass-rate understates GGUF
If we'd stopped at the checks column we'd have called MLX the quality winner by 7 points. Running the apps erased that lead. Report what the software does, not what the grader's first pass said.
The verdict
Remarkably close, and the folklore was wrong — MLX is not a downgrade. It ties GGUF on whether the app works, and it's the steadier build: lower variance, and it never blew the cap. GGUF is faster on compute-heavy work and has a higher peak, at the cost of a long tail. On Apple Silicon, MLX is the safe default; GGUF has the edge if your workload is compute-bound and you can tolerate the spikes.
Method & raw data
swarm-gym medium tier — 5 runs each of 3 archetypes per variant, 30 runs total. Flags: smoke + split + pre-review + contracts. Deterministic golden-value grading. Fleet: three Apple Silicon nodes serving qwopus3.6-27b via LM Studio. The per-run CSV (wall time, checks, task counts, exit codes for all 30 runs) is emitted by the harness. Reproduce it:
1python -m harness bench --tier medium --variant mlx
2python -m harness bench --tier medium --variant gguf
3python -m harness bench-report # writes BENCHMARK.md + the CSVs