Benchmark run

qwen3.8-27b 3-node swarm — 0.4616 on Gauntlet 7.0

qwen/qwen3.8-27b (3-node LM Studio fleet) · 3 nodes · scorer sb-7.0

3-node local fleet (qwen3.8)Sep 2, 2026
Overall
0.4616
Excellent
no
Benchmark
Gauntlet 7.0
Recorded duration
507m 21s
Nodes
3
Prompt tok
11.2M
Gen tok
622.3k

Tier breakdown

A · structure & runtime1.00
B · behaviour1.00
C · vendor contract0.95
D · finesse0.80

Scoring detail

How this score was built

Per-check scorer output, as posted by the app. Expand a section to inspect its checks.

The number, exactly

(0.88 × 0.8252 core + 0.12 × 0.74 gate × 0.49 excellence) × 0.6000 critical = 0.4616

Critical defects compound a multiplier on the whole score (pre-severity 0.7693):

j_workflow_journey 0.00 → ×0.60

The core (88% of the total) is the weighted mean of the ten measured tiers. The last 12% is the excellence slice: it unlocks in proportion to the perfection conditions below (11 of 16 met here), then pays out at the excellence tier's own measured mean. Core 0.7262 + excellence 0.0431.

j_first_use 0.25j_workflow_journey 0.00j_error_state 0.30j_empty_state 1.00console_clean 0.00v_responsive_375 1.00v_dates_readable 1.00t_scene_binding 0.21x_conservation_residual 1.00r_no_row_loss 1.00p_drag_frames 1.00p_idle_flatness 1.00p_stream_apply 0.00p_under_stream 1.00p_api_latency 1.00p_sync_wall 1.00

Run notes and corrections

Local swarm run r6h (engine 393a99351, 2026-09-02): hermetic sb-7 score 0.4616, inner 0.8252. Tier means: structure 1.00, spec behaviour 1.00, contracts 0.95, persistence 0.80, journeys 0.51, visual 1.00, streaming 0.83, 3D field 0.69, cross-checks 0.83, resilience 0.90, excellence 0.36. One CRITICAL check multiplies the result by 0.6: the maker/checker approval journey cannot complete through the UI (the backend flow works over the API; the console page does not drive it). The frontend renders 50 rows with zero console errors and a live notifications feed, and for the first time the 3D field draws and answers picks: 12,290 instanced columns bound to the data, camera math exact, 4 of 6 decisive picks agree three ways; it misses the digest tolerance, the rest-pixel colour and the pick triplet. Run shape: 8.45 h wall, UNCAPPED — no wall-clock or volume limit anywhere in the engine. 10 of 10 planned tasks completed, none failed, none retried, and the judge never intervened (supervision is evidence-only and found no repeat, no degenerate answer, no stall). Phases: open 66 min, research 8, synthesis 12, split 20, build 319, integrate 46, repair 36. The engine measured the 3D module as too fat for one lane (12 spec sections in one file) and split it into three shards built in parallel on the three nodes, then assembled the 11 pieces by code and had the merger write glue only (81 KB from 79 KB assembled). Repair promoted two fixes to the page markup and shipped one finding as a known active bug: the app carries no executable test suite. The pre-repair tree scored 0.4748, so repair cost 0.013. Sampling was left to LM Studio’s model defaults. The score is hermetic: a fresh clone, the run’s own fixture seed (8ee514993063ff59), the advertised port, scored serially. Fleet: 3× qwen/qwen3.8-27b on LM Studio (workhorse/mihai/gabee). Token rates are engine-measured per call (median prefill 2811.8 tok/s, decode 15.1 tok/s).

Token rates

Measured by the engine itself, one record per completed model call: prefill rate is prompt tokens over time-to-first-token, generation rate is completion tokens over the decode window. Medians per node.

NodeCallsPrompt tokGen tokPrefill tok/sGen tok/s
gabee432,442,182193,8551732.611.6
mihai513,028,073116,5581855.212.8
workhorse1185,759,466311,8844081.817.2
fleet21211,229,721622,2972811.815.1

Screenshots

Recorded views of the built application. Captions identify the available captures; the number of images does not imply a number of repair rounds.

1loaded
2mobile
3loaded
4mobile

Run details

Model
qwen/qwen3.8-27b (3-node LM Studio fleet)

Fleet nodes

NodeModel
workhorseworkhorse-qwen/qwen3.8-27b
mihaimihai-qwen/qwen3.8-27b
gabeegabee-qwen/qwen3.8-27b