Benchmark run

qwen3.6-27b 3-node swarm — 0.0154 on Gauntlet 7.0

qwen3.6-27b-fable-fusion-711 (3-node LM Studio fleet) · 3 nodes · scorer sb-7.0

3-node local fleetAug 21, 2026
Overall
0.0154
Excellent
no
Benchmark
Gauntlet 7.0
Recorded duration
180m 0s
Nodes
3
Prompt tok
4.9M
Gen tok
214.5k

Tier breakdown

A · structure & runtime0.38
B · behaviour0.00
C · vendor contract0.00
D · finesse0.13

Scoring detail

How this score was built

Per-check scorer output, as posted by the app. Expand a section to inspect its checks.

The number, exactly

(0.88 × 0.0292 core + 0.12 × 0.00 gate × 0.00 excellence) × 0.6000 critical = 0.0154

Critical defects compound a multiplier on the whole score (pre-severity 0.0257):

server_runs 0.00 → ×0.60

The core (88% of the total) is the weighted mean of the ten measured tiers. The last 12% is the excellence slice: it unlocks in proportion to the perfection conditions below (0 of 16 met here), then pays out at the excellence tier's own measured mean. Core 0.0257 + excellence 0.0000.

j_first_use 0.00j_workflow_journey 0.00j_error_state 0.00j_empty_state 0.00console_clean not measuredv_responsive_375 0.00v_dates_readable 0.00t_scene_binding 0.00x_conservation_residual 0.00r_no_row_loss 0.00p_drag_frames 0.00p_idle_flatness 0.00p_stream_apply 0.00p_under_stream 0.00p_api_latency 0.00p_sync_wall 0.00

Run notes and corrections

The local fleet’s best sb-7 run so far (its 4th attempt), still an honest FLOOR: the app now owns every advertised boot entry and validates its inputs exactly as the spec documents, but neither service ever starts listening — the process runs and never binds — so every runtime tier stays at zero and the score reflects structure and effort, not served behaviour. Two engine defects this run exposed (missing package entries in earlier plans; the boot probe handing the app an empty tokens file and repairing a defect that did not exist) are fixed in the engine for the next attempt. Fleet: 3× qwen3.6-27b on LM Studio (workhorse/mihai/gabee). Token rates in this entry are engine-measured per call (median prefill 364.5 tok/s, decode 12.9 tok/s).

Token rates

Measured by the engine itself, one record per completed model call: prefill rate is prompt tokens over time-to-first-token, generation rate is completion tokens over the decode window. Medians per node.

NodeCallsPrompt tokGen tokPrefill tok/sGen tok/s
gabee631,334,32948,6311414.28.7
mihai571,295,61654,9731172.510.4
workhorse1592,229,139110,880331.613.6
fleet2794,859,084214,484364.512.9

Run details

Model
qwen3.6-27b-fable-fusion-711 (3-node LM Studio fleet)

Fleet nodes

NodeModel
workhorseworkhorse-qwen3.6-27b-fable-fusion-711
mihaimihai-qwen3.6-27b-fable-fusion-711
gabeegabee-qwen3.6-27b-fable-fusion-711