Benchmark run
3-node local fleet (qwen3.8) — 0.4616 on sb-7.0
qwen/qwen3.8-27b (3-node LM Studio fleet) · 3 nodes · scorer sb-7.0
- Overall
- 0.4616
- Excellent
- no
- Scorer
- sb-7.0
- Wall clock
- 507 min
- Nodes
- 3
- Prompt tok
- 11.2M
- Gen tok
- 622.3k
Tier breakdown
Scoring detail
How this score was built
Per-check scorer output, as posted by the app.
The number, exactly
(0.88 × 0.8252 core + 0.12 × 0.74 gate × 0.49 excellence) × 0.6000 critical = 0.4616
Critical defects compound a multiplier on the whole score (pre-severity 0.7693):
The core (88% of the total) is the weighted mean of the ten measured tiers. The last 12% is the excellence slice: it unlocks in proportion to the perfection conditions below (11 of 16 met here), then pays out at the excellence tier's own measured mean. Core 0.7262 + excellence 0.0431.
a package layout
1.00The submitted tree contains every deliverable the spec names by path — app/__main__.py, DECISIONS.md, and the four web files (index.html, styles.css, app.js, viz.js) — plus a bootable module for each of the two services.
6/6 named files, 2/2 bootable service modules
server runs
1.00Both services boot as real processes and report healthy over HTTP within the spec's 10-second budget — the tool runs at all.
both services healthy, ledgerd in 0.3s
a combined entrypoint
1.00The documented single-command form — python -m app — really boots BOTH services, not just the two individual service commands the harness normally uses.
python -m app: boot=True ledger=200 notifier=200
serves page
1.00The backend itself serves the frontend: GET / returns the page and styles.css, app.js, viz.js come back with correct content types, as a real browser would need.
page 200, 3/3 assets served with correct content types
a asset budget
1.00The whole frontend fits the 150 KB source budget and ships zero external code — the budget exists so a vendored 3D library cannot, and the page must work fully offline.
119 KB of 150 KB (4/4 files), 0 external ref(s)
a db ownership
1.00Each service owns exactly its own SQLite file under --db-dir — ledgerd's ledger.db and notifierd's notifier.db — the one-file-per-service contract.
ledger.db=True notifier.db=True
sync completeness
1.00After the self-driven full sync the app's local store holds every one of the vendor's 12,288 fixture payments — an incomplete sync silently loses money rows.
12290/12288 payments after the self-driven full sync
b total field
1.00GET /api/payments reports the correct collection total — a wrong total breaks every caller's paging math.
total=12289 (want 12289)
b row shape
1.00A payment row on the wire carries exactly the 10 documented keys (id, amount_minor, currency, created_at, settled_at, status, version, note, counterparty_name, country) — the vendor's nested counterparty object flattened into the last two.
10/10 documented keys
b chronological order
1.00The default payments page is ordered by created_at as a parsed instant, not as a string — mixed UTC offsets make string ordering wrong.
49/49 adjacent pairs ordered by instant
b summary shape
1.00GET /api/summary carries the documented per-currency money blocks: by_currency and the reversals block, both sorted ascending by currency code — the summary is the money surface.
by_currency 4 rows (sorted=True), reversals present (4 rows)
b money rendered
1.00Rendered money is arithmetically correct per currency — right digits and right decimal exponent, with the zero-decimal JPY and three-decimal KWD traps weighted separately — and no cross-currency sum exists anywhere in the summary.
1.00 of 12 cells exponent-correct; JPY/KWD traps 1.00
b buckets dst
1.00GET /api/buckets — the 3D field's aggregate — buckets payments into Europe/Berlin calendar days per status, correct across the seeded DST transition (the data-correctness trap).
384/384 cells exact, DST-window 1.00, tz=Europe/Berlin
b viz records
1.00GET /api/viz/records — the 3D field's sanctioned full fetch — serves the whole collection columnar with the server-computed Berlin day, so the frontend never recomputes days in UTC.
7/7 columns, count=12289, order 1.00, server Berlin day 1.00
b events log
1.00The append-only event ledger is contiguous from seq 1 with the frozen type and source vocabularies — a seq gap is evidence of a lost write and is graded as one.
16000 events, contiguous=True, vocab 1.00/1.00
b error envelope
1.00Error responses carry the documented structured envelope — error.code and error.message plus field_errors[] with dot-and-[index] paths where validation detail is expected.
envelope 1.00, field paths 1.00 over 7 cases
b json shapes
1.00Every sampled API response — success and error paths alike — is parseable JSON; an HTML error page mid-API breaks every client that trusted the contract.
6/6 responses parse as JSON
c paged walk
1.00The first sync walks the vendor's server-fixed 64-per-page protocol to completion — all 192 pages, documented parameters only, no page fetched twice.
192/192 pages served, 0 undocumented-param requests, 0 duplicate pages
c b1 drop resume
1.00A connection dropped mid-page-stream during the first walk costs a documented resume — not a full restart of committed work, and not a hole in the collection.
{'armed': True, 'resumed': True, 'no_full_restart': True, 'walk_completed': True}
c b2 retry after
1.00A seeded 500 with Retry-After during the walk costs exactly one documented retry after the advertised wait — never a fresh unconditional restart of committed work.
{'armed': True, 'single_retry': True, 'waited': True, 'continued': True}
c b5 generation 304
0.67The collection-generation rule: a 304 whose X-Collection-Generation disagrees with the stored generation is a cache miss — drop the validator and refetch unconditionally exactly once, without looping and without serving stale data as fresh.
unconditional=0 conditional-repeats=2 propagated=True
c conditional resync
1.00Later syncs are cheap: every re-sync request carries a validator (If-None-Match) except the documented unconditional cases — an unconditional full re-walk on every sync is the expensive-client defect.
279/279 conditional (1 documented unconditional), 202 x 304 across post-mutation syncs
c webhook discipline
1.00The app accepts the vendor's signed push traffic — every delivery acknowledged 2xx within budget — and bounces the one forged signature with 401, state untouched.
53/53 deliveries 2xx-acked, forged -> 401
c send idempotency
1.00Approved drafts become real vendor payments through POST /v3/payments with a stored Idempotency-Key, and the kill-interrupted send's retry reuses that key — a fresh key per retry is the seeded duplicate-payment bug.
2/2 sends carried Idempotency-Key, retry_reused=True
d content types
0.50API responses declare a JSON content type and the SSE stream declares text/event-stream — a wrong content type breaks strict clients and kills EventSource.
api json 1.00, /api/stream ctype ''
d validation
1.00Invalid input and wrong-role requests are rejected with the documented status codes — bad limit/offset/sort/status 400, unknown paths 404, missing tokens 401, wrong roles 403, self-approval 403 approval_forbidden.
14/14 status codes correct
d client timeouts
1.00The app is provably resilient to an unresponsive vendor: it boots, binds and serves local data while the vendor refuses connections — one hung vendor call must never hang the tool.
ledgerd bound and served local state while the vendor refused connections (8s window)
d peer absence
1.00Neither service crashes on the other's absence: with notifierd down the proxy answers 502 with the documented envelope code and ledgerd keeps running; with ledgerd killed notifierd keeps running.
proxy -> 502 code=True, ledgerd-alive=True, notifierd-alive=True
d decisions doc
0.50The three deliberately unstated corners — D1 brush survival on streamed mutation, D2 rejected-draft terminality, D3 pre-first-sync table state — are decided, documented under the frozen headings in DECISIONS.md, and consistent with observed behavior.
D1=0.5, D2=0.5, D3=0.5
j loads data
1.00Whether the Meridian Payments Console renders any payment rows at all in a real browser, and whether the total it claims on the page matches the collection truth at probe time.
50 rows rendered, DOM claims 12290 (want 12290)
j console clean
1.00Whether normal use of the console — loading it and running a sync — produces JavaScript console errors or uncaught page errors.
0 console error(s) across load+sync
j first use
0.25The first-visit experience as one property: real data appears quickly, the on-page total agrees with the collection truth, and the console stays clean.
ttfd=4926ms reconcile=1.0 consoleErrs=0
j sync journey
1.00The headline interactive flow: a user clicks Sync now and the app visibly starts, indicates progress, finishes, and refreshes the view.
button_found, in_flight_state, completed, state_stays_correct
j workflow journey
0.00The maker/checker approval workflow driven end to end through the UI alone: token in, draft created, submitted, approved, listed, notified, and the vendor-created payment landing in the payments table.
roleTokenAccepted
j workflow reject
0.00The checker's other half of the workflow: rejecting a submitted draft completes through the UI, the rejected state shows in the draft list, and the rejection notification appears in the feed.
rejection never completed
j notifications feed
1.00The notifications feed visibly degrades while notifierd is down and heals itself — without a reload — once it returns.
partition state=degraded, heal state=live, live in 0.3s
j error state
0.30What a user sees when the backend is unreachable: a visible, actionable error state instead of a blank or silently broken page.
error indication only
j empty state
1.00What a fresh install looks like before the first sync completes: an honest empty-or-progress state rather than phantom data or a blank page.
pre-sync state rendered: {'tablePresent': True, 'renderedRowCount': 0, 'progressText': 'No payments to render yet.', 'emptyWithProgress': True, 'blocked': False, 'timeline': [{'tMs': 10, 'rows': 0, 'emptyText': 'No payme
v dates readable
1.00Whether the Date column shows humans a readable date instead of a raw machine timestamp.
rendered dates like Aug 19, 2025, 01:01
v money presentation
1.00Whether rendered amounts identify their currency — the presentation half of money; the exponent-and-digits truth is graded separately by the critical b_money_rendered.
12/12 amount cells carry a recognizable currency token
v status badges
1.00Whether the four payment statuses are visually distinguishable at a glance and painted in the spec's frozen palette — the same four hexes the 3D field uses.
4 statuses on page 1, 4 at the frozen hex, 4 distinct styles
v responsive 375
1.00Whether the page survives a phone-width viewport: no horizontal scrolling and real content still rendered at 375 px.
no h-scroll, 50 rows (tap targets not measured by this probe)
v styling
1.00Whether the page is deliberately styled rather than browser-default: a real stylesheet, layered surface colors, a chosen font, and a branded header.
stylesheet=True, 7 backgrounds, font '-apple-system, "system-ui", "S', header=True
p drag frames
1.00The 3D field stays interactive under input: real frames rendered during the scripted 40-move budget drag at the full 12,288-instance count.
40 frames over the 40-move budget drag (12435.5ms)
p idle flatness
1.00Demand rendering, the frozen spec rule: at rest — no input, no coast, no pending stream batch — the scene draws nothing; a continuous rAF render loop fails by design.
0 default-FBO draws in the worst 500ms rest window (2 sampled)
p stream apply
0.00A live SSE batch becomes visible fast: the median time from receipt to applied — store, digest and pixels — against the spec's 250 ms budget.
no stream batch applied
p under stream
1.00The API stays fast AND correct while the SSE stream burst is landing — read p95 under proven concurrent load.
p95=6.3ms with readers during the stream burst (overlap 1.00)
p api latency
1.00The read endpoints answer within their spec budgets when idle: the worst p95 across the API latency battery.
worst idle p95 16.32483396679163 ms across ['payments', 'summary']
p sync wall
1.00Wall-clock time for the self-driven first sync — the full 192-page walk, seeded faults and documented waits included — against the spec's 120 s budget.
sync #1 wall 7542 ms (self-driven, faults included)
t context real
0.75The #viz3d panel is a real, drawing WebGL surface with the pinned context attributes and a correctly sized backing store — not a styled div, an image, or an unused canvas.
ctx=1 type=webgl2 nonBg=1/24 distinct=1 backingOk=True
t layout basis
1.00The locked layout basis vs7dbg.layout() reports — d0 (first day), D0 = 96 (the span), R0 (max in-day count at load) — matches the fixture truth and never moves when a streamed create arrives.
layout got={'d0': '2025-08-19', 'D0': 96, 'R0': 179} want={'d0': '2025-08-19', 'D0': 96, 'R0': 179} unmoved_after_stream=True
t scene binding
0.21The rendered scene actually encodes the payment data: vs7dbg.sceneDigest()'s seven statistical moments over all 12,288 instanced columns match an independent recomputation from the fixture.
3/7 digest moments within tolerance (count got=12290 want=12290)
t height pixels
0.33Column heights are true in rendered pixels — including the JPY (exponent 0) and KWD (exponent 3) instances whose heights expose a forgotten currency exponent.
2/6 instance tops within ±3 px (currencies ['EUR', 'JPY', 'KWD', 'USD'])
t draw budget
1.00The field renders 12,288 instances inside the draw budget — at most 8 default-framebuffer draw calls per rendered frame — forcing instanced draws or a merged buffer instead of per-column draws.
ΔD=40 ΔF=40 over M=40 moves (limits 320/384)
t pick buffer
0.60Picking is GPU truth: the offscreen pick buffer answers occlusion exactly as the depth buffer says, on click points constructed to kill CPU raycasts and last-drawn-wins shortcuts.
4/6 decisive picks agree three ways; constructions {'occluded-by-lower-n': False, 'occluded-by-higher-n': True, 'partial-occlusion': False, 'background-in-hull': True}
t pick real pass
1.00The pick buffer is real GPU work with the documented cost profile: a fresh offscreen pass after each scene invalidation, bounded at 4 offscreen draws, and never flashing ID colors on the visible canvas.
since invalidation: 2 offscreen draws, 2 readbacks; 0 default-FBO draws across 6 pick calls
t click semantics
0.67Canvas clicks follow the frozen §3.3 semantics: click an instance to toggle it into the brush, click it again to toggle it out, click background to clear the brush.
{'click_toggles_on': False, 'click_toggles_off': True, 'background_clears': True}
t camera math
0.86The orbit camera implements the documented math exactly: defaults yaw 30 / pitch 40 / distance 260, the 0.30 deg-per-px drag law, the exponential wheel law without page scroll, both clamps, double-click reset, and a projection that matches the printed formula.
projErr=5px pitch@clamp=85 defaults, drag_law, wheel_law, pitch_clamp, dist_clamp, dblclick_reset
t coast identity
1.00Post-release inertia obeys the closed-form τ = 0.4 s decay law: at any coasting instant the remaining travel equals v(t)·τ, a slow release starts no coast, and the coast settles inside the printed budget.
identity 1.00 over 2 mid-coast samples, slow=True (drift 0°), settle=True (1142.1ms of 1657ms)
t coast reality
1.00The coast is real motion on the canvas, not a camera() narrative: after a fast flick the scene provably keeps moving past release.
coasted 8.229° past release, direction=True, pixels moved 5/5, rest pixel ok=False
t labels culling
0.90The 12 highest-amount records get screen-space labels that are collision-culled deterministically THROUGH the app's own pick buffer — exact set, exact geometry, zero overlap, never floating over an occluded instance.
setMatch=True at the decisive pose (expected 6, dom 6), geom 1.00, overlapViolations=0
t brush link
0.33One brush set links the table and the 3D field in both directions: row clicks and instance clicks toggle the same set, non-members dim to the exact 0.30 pixel rule, the count readout tracks, and background click restores full color.
row_click_toggles, brush_count, background_clears
t stream diff
0.50SSE batches apply as true diffs: only the changed instances upload, no buffer realloc, the digest moves by exactly the change, and the changed instance's pixels show it.
1 batches: bytes 1.00, reallocs none, digest 0.00 (1 graded), changed pixel False
t vs7dbg truth
0.20The mandated window.vs7dbg instrumentation tells the truth: camera() agrees with the pixels, sceneDigest() with the recomputed data, frames() with the wrapper's counted draws, and pick() with pickPixel() with the analytic answer.
cameraYawErr=0 restPixel=False digest 0.0 framesAgree=True pickTriplet=False
x l1 no invented states
1.00Every (payment, version) the app ever applied or served exists in the vendor's committed history — an invented state means the app fabricated data.
14837 observations, 0 invented states
x l2 per key order
1.00Applied versions per payment strictly increase in event-log order — duplicate and stale webhook outcomes belong in the counters, never as events.
0 order violations over 10 same-key event pairs
x l3 monotonic reads
1.00The version served for a payment never decreases from one read to the next — a sync page landing after a webhook applied v+1 must not regress the row.
0 regressions over 2790 sampled read pairs
x l4 convergence
1.00At quiescence every payment's version and status equals the vendor's final committed state, and the row count equals the vendor's — the mid-walk create present exactly once.
24/26 touched rows at vendor-final state, total 12290 (want 12290)
x l5 group atomicity
0.00No read ever observes half a transaction group: the refunded payment and its reversal become visible together or not at all.
19 confirmed half-applied group observations over 938 samples
x m1 amount immutability
1.00No served row ever shows an amount_minor different from the vendor-committed amount — v3 never mutates amounts, only status/note/version.
2817 amount observations, 0 mutated
x m2 pair conservation
0.00At every observed instant, per currency, the summary's reversal totals equal the refunded rows' amounts — both halves of each refund visible, or neither.
918/938 snapshots pair-conserved, 19 confirmed half states
x m3 terminal conservation
1.00Terminal per-currency counts and totals — reversals included — equal vendor ground truth: fixture plus scripted mutations plus every payment the app created.
4/4 currency totals exact, 4/4 reversal totals exact at quiescence
x m4 no cross currency
1.00No field anywhere in the summary carries a cross-currency money sum — minor units are not a common denomination and summing them is wrong money.
no cross-currency money sum anywhere
x conservation residual
1.00CRITICAL — after every duplicate and loss is attributed, no minor units remain created or destroyed: the money conservation residual is zero in every currency.
residual 0 in every currency after attribution
x no lost write
1.00CRITICAL — every mutation the app acknowledged with a 2xx is present in the final state; an acked-then-vanished write is silent data loss.
23 acked webhook mutations checked, 0 lost
x ooo dup forged
1.00The three webhook trust traps land correctly: the out-of-order pair keeps v+2 (the late v+1 never overwrites), the forged-signature event changes nothing, and no duplicate delivery applies twice.
{'ooo_kept_v2': True, 'forged_untouched': True, 'dup_applied_once': True}
r b3 sigkill resync
1.00ledgerd SIGKILLed mid-sync restarts, converges, and duplicates nothing — a kill mid-walk costs a clean cursor restart, never dupes or holes.
kill after list #2: restart=True converge=True nodupes=True
r b4 vendor down boot
1.00The app boots with the vendor down: binds within 10 seconds anyway, serves local data, does not crash, and completes the first sync unprompted once the vendor returns.
{'armed': True, 'refusal_secs': 8, 'bound_in_10s': True, 'served_while_down': True, 'no_crash': True, 'recovered': True}
r b6 outbox atomic
1.00A SIGKILL between an outbox commit and its delivery loses nothing and doubles nothing — the exact window where commit-then-POST and POST-then-commit both fail.
{'pending_before_kill': True, 'resumed': True, 'exactly_once': True, 'none_lost': True}
r b7 partition
1.00A notifier partition degrades visibly and heals in order: writes never block, /api/outbox/status reports down with growing pending, the feed shows degraded, and catch-up is in seq order with the UI live again within 5 seconds of heal.
{'ledger_alive': True, 'writes_fast': True, 'status_down': True, 'ui_degraded': True, 'catchup_in_order': True, 'ui_live_5s': True}
r notifier exactly once
1.00The notifier's durable processed set proves exactly-once: every outbox-crossing event processed, no seq twice, surviving kills.
4061 processed (4061 unique), crossing coverage 1.00
r notification multiset
0.00Selective materialization is exact: draft.submitted, draft.approved, draft.rejected and reversal.created each produce exactly one notification row; payment.sent produces none.
got {'draft.approved': 1, 'draft.submitted': 1, 'reversal.created': 198} want {'reversal.created': 3955}
r no row loss
1.00CRITICAL — no committed row goes missing after any seeded kill: the row count after every SIGKILL+restart is at least the count before it.
4 kills; no rows lost
r no dupe effect
1.00CRITICAL — no ledger effect applies twice: one vendor payment per approved draft, no duplicate notification rows, no (type, payment, version) event effect repeated.
no duplicated effects
r cache truth
1.00CRITICAL — the lying-304 trap: a 304 whose X-Collection-Generation disagrees with the stored generation is a cache miss, and serving the stale cache as fresh is graded data loss.
post-304 state reflects the committed mutations
r workflow durability
1.00CRITICAL — submitted and approved are durable the moment their 200 is written: a SIGKILL immediately after either, including mid-send, must find the state intact after restart.
A1 -> submitted, A2 -> sent
e frames under drag
1.00Excellence-grade fluidity: frames rendered during the scripted 40-move drag at the full 12,288-instance count, with proof the frames actually drew.
40 frames over the 40-move drag at N=12288
e stream apply latency
0.00Excellence-grade streaming: the median SSE batch apply time against rungs 2.5× tighter than the P-tier budget.
no stream batch applied
e under load latency
1.00The read API's p95 under the proven stream-burst load, graded again in the excellence slice.
p95=6.3ms under the stream burst (overlap 1.00)
e optimistic paint
0.00The workflow UI paints optimistically: the submitted state appears while the write is provably still on the wire, then really saves.
no optimistic-paint exercise in the flow emit
e mastery
0.44Excellence includes the mechanisms, not only the surface: mastery of the entire 3D contract (T), consistency-and-money invariants (X) and resilience (R) tiers together.
T+X+R mean 0.793
Token rates
Measured by the engine itself, one record per completed model call: prefill rate is prompt tokens over time-to-first-token, generation rate is completion tokens over the decode window. Medians per node.
| Node | Calls | Prompt tok | Gen tok | Prefill tok/s | Gen tok/s |
|---|---|---|---|---|---|
| gabee | 43 | 2,442,182 | 193,855 | 1732.6 | 11.6 |
| mihai | 51 | 3,028,073 | 116,558 | 1855.2 | 12.8 |
| workhorse | 118 | 5,759,466 | 311,884 | 4081.8 | 17.2 |
| fleet | 212 | 11,229,721 | 622,297 | 2811.8 | 15.1 |
Screenshots
Captured by the render gate during the run's repair rounds. The first capture is the initial render; the remainder are from the final epoch.
Run details
- Model
- qwen/qwen3.8-27b (3-node LM Studio fleet)
Fleet nodes
| Node | Model |
|---|---|
| workhorse | workhorse-qwen/qwen3.8-27b |
| mihai | mihai-qwen/qwen3.8-27b |
| gabee | gabee-qwen/qwen3.8-27b |
Notes
Local swarm run r6h (engine 393a99351, 2026-09-02): hermetic sb-7 score 0.4616, inner 0.8252. Tier means: structure 1.00, spec behaviour 1.00, contracts 0.95, persistence 0.80, journeys 0.51, visual 1.00, streaming 0.83, 3D field 0.69, cross-checks 0.83, resilience 0.90, excellence 0.36. One CRITICAL check multiplies the result by 0.6: the maker/checker approval journey cannot complete through the UI (the backend flow works over the API; the console page does not drive it). The frontend renders 50 rows with zero console errors and a live notifications feed, and for the first time the 3D field draws and answers picks: 12,290 instanced columns bound to the data, camera math exact, 4 of 6 decisive picks agree three ways; it misses the digest tolerance, the rest-pixel colour and the pick triplet. Run shape: 8.45 h wall, UNCAPPED — no wall-clock or volume limit anywhere in the engine. 10 of 10 planned tasks completed, none failed, none retried, and the judge never intervened (supervision is evidence-only and found no repeat, no degenerate answer, no stall). Phases: open 66 min, research 8, synthesis 12, split 20, build 319, integrate 46, repair 36. The engine measured the 3D module as too fat for one lane (12 spec sections in one file) and split it into three shards built in parallel on the three nodes, then assembled the 11 pieces by code and had the merger write glue only (81 KB from 79 KB assembled). Repair promoted two fixes to the page markup and shipped one finding as a known active bug: the app carries no executable test suite. The pre-repair tree scored 0.4748, so repair cost 0.013. Sampling was left to LM Studio’s model defaults. The score is hermetic: a fresh clone, the run’s own fixture seed (8ee514993063ff59), the advertised port, scored serially. Fleet: 3× qwen/qwen3.8-27b on LM Studio (workhorse/mihai/gabee). Token rates are engine-measured per call (median prefill 2811.8 tok/s, decode 15.1 tok/s).