Benchmark run

qwen3.6-27b 3-node swarm — 0.9300 on Gauntlet 5.3

qwen3.6-27b-fable-fusion-711-uncensored-heretic-nm-dau-neo-max-mtp · 3 nodes · scorer sb-5.3

mighty-crane-54f23-node fleet — full fix-chain run3-node local fleetAug 18, 2026
Overall
0.9300
Hard
0.83
Excellent
yes
Benchmark
Gauntlet 5.3
Recorded duration
189m 55s
Nodes
3

Tier breakdown

A · structure & runtime1.00
B · behaviour1.00
C · vendor contract1.00
D · finesse0.94

Scoring detail

How this score was built

Per-check scorer output, as posted by the app. Expand a section to inspect its checks.

ComponentWeightEarnedPoints
Core (tiers A–D)60%0.9880.593
A · Structure25% of core1.000.150
B · Behaviour30% of core1.000.180
C · Vendor contract25% of core1.000.150
D · Finesse20% of core0.940.112
Journey (J)15%0.7500.112
Visual (V)10%0.9170.092
Performance (P)5%1.0000.050
Hard blocks10%0.8330.083
Total0.9300

Findings that held

Raised by the verify gate, still open when the run ended.

1

`pytest -q` failed — the generated tests exercise runtime paths that `--help`/`--collect-only` never invoke: client = MeridianClient("http ... [middle elided — head + tail shown] ... ") > self.assertEqual(captured_req[0].headers["Idempotency-Key"], "key-abc") ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ E KeyError: 'Idempotency-Key' tests/test_merid

2

POST /api/sync is not CHEAP on a repeat run — the second sync re-fetched 247 row(s) it already had. FIX: make the client send If-None-Match per page — ... [middle elided — head + tail shown] ... lay THAT page's ETag on the matching request; treat 304 as 'page unchanged, keep local rows'. One ETag replayed on every page never matches and re-fetches everything. Rows are correct, so this is not a d

3

GET /api/summary answered while the app was EMPTY and returns NOTHING AT ALL once it holds rows — same process, same endpoint, one sync in between, and its sibling endpoints still answer. That pair rules out both a slow endpoint and a dead server, and points at the code path that reads STORED rows: an

4

the page renders but the browser console carries 4 error(s) in normal use (first: Failed to load resource: net::ERR_EMPTY_RESPONSE) — fix the JS errors; users hit them as broken interactions.

Repair progression

R0·14R1·4

One chip per verify round: the number of open findings at that round. The run ends when a round reports zero findings or the round budget is exhausted.

Screenshots

Recorded views of the built application. Captions identify the available captures; the number of images does not imply a number of repair rounds.

1First render — before repairs
2Final render
3Mobile · 375px
4After sync

Run details

Model
qwen3.6-27b-fable-fusion-711-uncensored-heretic-nm-dau-neo-max-mtp
Engine events
801
Repair rounds
1
Started
Aug 18, 2026, 10:36 AM
Finished
Aug 18, 2026, 01:46 PM

Fleet nodes

NodeModel
gabeegabee-qwen3.6-27b-fable-fusion-711-uncensored-heretic-nm-dau-neo-max-mtp
mihaimihai-qwen3.6-27b-fable-fusion-711-uncensored-heretic-nm-dau-neo-max-mtp
workhorseworkhorse-qwen3.6-27b-fable-fusion-711-uncensored-heretic-nm-dau-neo-max-mtp