Agentic Benchmark

qwopus3.6-27b: MLX vs GGUF on Apple Silicon

July 2, 2026by LeanZero

The first benchmark run with goose local-edition and its swarm-gym harness. We had a local-model swarm build three real command-line apps, five times each, with two builds of the same model — qwopus3.6-27b as the original GGUF and as the MLX translation — on a 3-node Apple Silicon fleet via LM Studio. Identical frozen prompts, identical binary, identical flags. 30 app builds in total, every one graded by running it against golden values. Here is everything the runs produced.

The apps that got built

Three archetypes, chosen to stress different failure modes — data modelling, algorithms, and stateful logic. Each was built 5× per model build (10× total) and checked by running real commands against it:

Archetype

The app

What it does

A golden check (verified by running)

crud-multiformat

invtrack

inventory store with JSON + table output and value aggregation

add 4 × $5 → “Total inventory value: $20.00”

compute-parser

rdcalc

a recursive-descent arithmetic calculator

2^3^2 = 512 (right-assoc); 2+3*4 = 14

transaction

txkvbench

a key/value store with nested BEGIN/ROLLBACK

nested rollback prints “2 1 1”; COUNT

4 rows × 4 columnsHeader row enabled

When a build passed, the app genuinely worked — we ran these commands on the finished binaries and got the golden answers back. When it failed, it failed one of these checks, no matter what the model claimed.

Speed

Median build time per run, per archetype — lower is faster:

Spec

GGUF median

MLX median

Faster

compute (rdcalc)

1860.5s

2192.6s

GGUF, by 15%

transaction (txkvbench)

1945.8s

1722.9s

MLX, by 13%

crud (invtrack)

2239.2s

2342.9s

~tie

overall

1945.8s

1923.9s

~tie

5 rows × 4 columnsHeader row enabled

But the medians hide the real story, which is consistency:

Consistency

GGUF

MLX

Std deviation

750.9s

580.4s (tighter)

p90 (slow tail)

3600.7s

3049.7s

Range (min–max)

1210s – 3601s

1001s – 3093s

Runs that hit the 60-min cap

2

0

5 rows × 3 columnsHeader row enabled

GGUF has a higher ceiling (its fastest run, a rdcalc build at 1210s, beat everything MLX did) but a much longer tail. MLX never once ran away.

Quality — did the apps actually work

Share of builds that passed every golden check, per archetype:

Spec

GGUF checks-pass

MLX checks-pass

compute (rdcalc)

5 / 5 (100%)

4 / 5 (80%)

crud (invtrack)

4 / 5 (80%)

5 / 5 (100%)

transaction (txkvbench)

3 / 5 (60%)

4 / 5 (80%)

overall

12 / 15 (80%)

13 / 15 (87%)

5 rows × 3 columnsHeader row enabled

MLX leads on raw checks, 87% to 80% — but two of GGUF's three misses were a correct app the harness mis-scored (a run capped mid-cleanup, and one transient false-partial that ran fine when we re-ran it). Judged on whether the software works, the two are even — ~13 of 15 each.

What worked, and what didn't

compute — GGUF's archetype

GGUF was flawless here: 5/5, and the fastest spec on the board (median 1860s). MLX went 4/5 — its one miss was a scheduler deadlock: the run stalled with two tasks stuck, the calculator's entry point never got written, and the app simply couldn't run. Not a parsing mistake, a swarm stall. Every other compute build, both formats, nailed right-associativity and precedence.

crud — MLX's archetype

MLX went 5/5 with the tightest variance of any cell (stdev 442s) — steady, correct inventory stores every time. GGUF built five working apps too (we verified the totals by running them), but two of the five hit the 60-minute cap and one scored a false-partial. Those caps are why GGUF's crud variance is nearly 2× MLX's — and, as it turns out, they weren't the model's fault (see below).

transaction — the hard one, for both

This was the weakest archetype for both formats, and it's the most interesting. The nested-rollback logic was usually right; the multi-command exec path is where it broke. MLX had one build whose exec GET printed nothing. GGUF had it worse: one build cascaded (3 of 4 tasks failed) and one simply never implemented the COUNT command — yet its own unit tests passed, so the model shipped it green. Same weakness, both formats: a weak-model completeness gap on the hardest spec, and the single clearest reason to grade by running instead of trusting the model's tests.

Three things running the apps revealed

1. GGUF's caps were a swarm bug, not the model

Both of GGUF's capped crud runs had already built a correct app — the harness caught the swarm's own integrate-verify step churning after the fact, with no wall-clock budget to stop it. That's a scheduler bug we've since fixed (an env-gated cap on that step) and are re-benchmarking. So GGUF's crud numbers are penalised for something that isn't the model.

2. Passing tests, broken feature — on both builds

The transaction misses shared a signature across GGUF and MLX: the model's generated unit tests passed while a required feature was broken or missing. Because it showed up on both formats, it's a model limit, not a runtime difference — and the harness's golden checks caught every instance the model's tests missed.

3. Raw pass-rate understates GGUF

If we'd stopped at the checks column we'd have called MLX the quality winner by 7 points. Running the apps erased that lead. Report what the software does, not what the grader's first pass said.

The verdict

Remarkably close, and the folklore was wrong — MLX is not a downgrade. It ties GGUF on whether the app works, and it's the steadier build: lower variance, and it never blew the cap. GGUF is faster on compute-heavy work and has a higher peak, at the cost of a long tail. On Apple Silicon, MLX is the safe default; GGUF has the edge if your workload is compute-bound and you can tolerate the spikes.

Method & raw data

swarm-gym medium tier — 5 runs each of 3 archetypes per variant, 30 runs total. Flags: smoke + split + pre-review + contracts. Deterministic golden-value grading. Fleet: three Apple Silicon nodes serving qwopus3.6-27b via LM Studio. The per-run CSV (wall time, checks, task counts, exit codes for all 30 runs) is emitted by the harness. Reproduce it:

bash
1python -m harness bench --tier medium --variant mlx
2python -m harness bench --tier medium --variant gguf
3python -m harness bench-report   # writes BENCHMARK.md + the CSVs
ShareXLinkedIn