Local fleet baseline

Gemini 3.8 Flash · SB7.1 pilot0.6990 on sb-7.1-rc

gemini-3.8-flash · 1 nodes · scorer sb-7.1-rc

Sep 20, 2026
Overall
0.6990
Excellent
no
Scorer
sb-7.1-rc
Wall clock
9 min
Nodes
1

Tier breakdown

A · structure & runtime1.00
B · behaviour1.00
C · vendor contract1.00
D · finesse0.83

Scoring detail

How this score was built

Per-check scorer output, as posted by the app.

Earned credit before admission

(0.88 × 0.9263 core + 0.12 × 0.94 gate × 0.75 excellence) × 0.8857 critical = 0.7968

Critical defects compound a multiplier on the whole score (pre-severity 0.8996):

j_workflow_journey 0.71 → ×0.89

The core (88% of the total) is the weighted mean of the ten measured tiers. The last 12% is the excellence slice: it unlocks in proportion to the perfection conditions below (14 of 16 met here), then pays out at the excellence tier's own measured mean. Core 0.8151 + excellence 0.0845.

j_first_use 1.00j_workflow_journey 0.71j_error_state 0.30j_empty_state 1.00console_clean 0.00v_responsive_375 1.00v_dates_readable 1.00t_scene_binding 1.00x_conservation_residual 1.00r_no_row_loss 1.00p_drag_frames 1.00p_idle_flatness 1.00p_stream_apply 1.00p_under_stream 1.00p_api_latency 1.00p_sync_wall 1.00

Earned score 0.797 · Admission ceiling 0.699 · Final score 0.699

Final score is the lower of earned credit and the admission ceiling. Passing admission adds no points.

Visible: passedMatching: not metGood: not metVisual excellence: not metBackend excellence: not met

Required payment-tower structure or currency mapping is incomplete

Payment context, readability or interaction quality is incomplete

Scores above 0.9 require both verified event animation and backend recovery excellence

ABoot & deliverablesthe two services boot and bind, the spec-named files exist, the 150 KB asset budget holds, each service owns only its own database

a package layout

1.00

The submitted tree contains every deliverable the spec names by path — app/__main__.py, DECISIONS.md, and the four web files (index.html, styles.css, app.js, viz.js) — plus a bootable module for each of the two services.

6/6 named files, 2/2 bootable service modules

server runs

1.00

Both services boot as real processes and report healthy over HTTP within the spec's 10-second budget — the tool runs at all.

both services healthy, ledgerd in 0.3s

a combined entrypoint

1.00

The documented single-command form — python -m app — really boots BOTH services, not just the two individual service commands the harness normally uses.

python -m app: boot=True ledger=200 notifier=200

serves page

1.00

The backend itself serves the frontend: GET / returns the page and styles.css, app.js, viz.js come back with correct content types, as a real browser would need.

page 200, 3/3 assets served with correct content types

a asset budget

1.00

The whole frontend fits the 150 KB source budget and ships zero external code — the budget exists so a vendored 3D library cannot, and the page must work fully offline.

92 KB of 150 KB (4/4 files), 0 external ref(s)

a db ownership

1.00

Each service owns exactly its own SQLite file under --db-dir — ledgerd's ledger.db and notifierd's notifier.db — the one-file-per-service contract.

ledger.db=True notifier.db=True

BWire behaviourwhat the API serves on the wire: sync completeness (12,288 rows), money and Berlin-DST bucketing, the append-only event ledger, error envelopes

sync completeness

1.00

After the self-driven full sync the app's local store holds every one of the vendor's 12,288 fixture payments — an incomplete sync silently loses money rows.

12291/12288 payments after the self-driven full sync

b total field

1.00

GET /api/payments reports the correct collection total — a wrong total breaks every caller's paging math.

total=12289 (want 12289)

b row shape

1.00

A payment row on the wire carries exactly the 10 documented keys (id, amount_minor, currency, created_at, settled_at, status, version, note, counterparty_name, country) — the vendor's nested counterparty object flattened into the last two.

10/10 documented keys

b chronological order

1.00

The default payments page is ordered by created_at as a parsed instant, not as a string — mixed UTC offsets make string ordering wrong.

49/49 adjacent pairs ordered by instant

b summary shape

1.00

GET /api/summary carries the documented per-currency money blocks: by_currency and the reversals block, both sorted ascending by currency code — the summary is the money surface.

by_currency 4 rows (sorted=True), reversals present (4 rows)

b money rendered

1.00

Rendered money is arithmetically correct per currency — right digits and right decimal exponent, with the zero-decimal JPY and three-decimal KWD traps weighted separately — and no cross-currency sum exists anywhere in the summary.

1.00 of 12 cells exponent-correct; JPY/KWD traps 1.00

b buckets dst

1.00

GET /api/buckets — the 3D field's aggregate — buckets payments into Europe/Berlin calendar days per status, correct across the seeded DST transition (the data-correctness trap).

384/384 cells exact, DST-window 1.00, tz=Europe/Berlin

b viz records

1.00

GET /api/viz/records — the 3D field's sanctioned full fetch — serves the whole collection columnar with the server-computed Berlin day, so the frontend never recomputes days in UTC.

7/7 columns, count=12289, order 1.00, server Berlin day 1.00

b events log

1.00

The append-only event ledger is contiguous from seq 1 with the frozen type and source vocabularies — a seq gap is evidence of a lost write and is graded as one.

15307 events, contiguous=True, vocab 1.00/1.00

b error envelope

1.00

Error responses carry the documented structured envelope — error.code and error.message plus field_errors[] with dot-and-[index] paths where validation detail is expected.

envelope 1.00, field paths 1.00 over 7 cases

b json shapes

1.00

Every sampled API response — success and error paths alike — is parseable JSON; an HTML error page mid-API breaks every client that trusted the contract.

6/6 responses parse as JSON

CSync disciplinehow the vendor is consumed: the 192-page walk, drop resume, Retry-After, the collection-generation rule, webhook and idempotency-key discipline

c paged walk

1.00

The first sync walks the vendor's server-fixed 64-per-page protocol to completion — all 192 pages, documented parameters only, no page fetched twice.

299/192 pages served, 0 undocumented-param requests, 0 duplicate pages

c b1 drop resume

1.00

A connection dropped mid-page-stream during the first walk costs a documented resume — not a full restart of committed work, and not a hole in the collection.

{'armed': True, 'resumed': True, 'no_full_restart': True, 'walk_completed': True}

c b2 retry after

1.00

A seeded 500 with Retry-After during the walk costs exactly one documented retry after the advertised wait — never a fresh unconditional restart of committed work.

{'armed': True, 'single_retry': True, 'waited': True, 'continued': True}

c b5 generation 304

1.00

The collection-generation rule: a 304 whose X-Collection-Generation disagrees with the stored generation is a cache miss — drop the validator and refetch unconditionally exactly once, without looping and without serving stale data as fresh.

unconditional=1 conditional-repeats=0 propagated=True

c conditional resync

1.00

Later syncs are cheap: every re-sync request carries a validator (If-None-Match) except the documented unconditional cases — an unconditional full re-walk on every sync is the expensive-client defect.

197/197 conditional (0 documented unconditional), 4 x 304 across post-mutation syncs

c webhook discipline

1.00

The app accepts the vendor's signed push traffic — every delivery acknowledged 2xx within budget — and bounces the one forged signature with 401, state untouched.

54/54 deliveries 2xx-acked, forged -> 401

c send idempotency

1.00

Approved drafts become real vendor payments through POST /v3/payments with a stored Idempotency-Key, and the kill-interrupted send's retry reuses that key — a fresh key per retry is the seeded duplicate-payment bug.

3/3 sends carried Idempotency-Key, retry_reused=True

DValidation & docsinput validation, content types, client timeouts, behavior when the peer service is down, and the DECISIONS.md judgment corners

d content types

0.50

API responses declare a JSON content type and the SSE stream declares text/event-stream — a wrong content type breaks strict clients and kills EventSource.

api json 1.00, /api/stream ctype ''

d validation

1.00

Invalid input and wrong-role requests are rejected with the documented status codes — bad limit/offset/sort/status 400, unknown paths 404, missing tokens 401, wrong roles 403, self-approval 403 approval_forbidden.

14/14 status codes correct

d client timeouts

1.00

The app is provably resilient to an unresponsive vendor: it boots, binds and serves local data while the vendor refuses connections — one hung vendor call must never hang the tool.

ledgerd bound and served local state while the vendor refused connections (5s window)

d peer absence

1.00

Neither service crashes on the other's absence: with notifierd down the proxy answers 502 with the documented envelope code and ledgerd keeps running; with ledgerd killed notifierd keeps running.

proxy -> 502 code=True, ledgerd-alive=True, notifierd-alive=True

d decisions doc

0.67

The three deliberately unstated corners — D1 brush survival on streamed mutation, D2 rejected-draft terminality, D3 pre-first-sync table state — are decided, documented under the frozen headings in DECISIONS.md, and consistent with observed behavior.

D1=1.0, D2=0.5, D3=0.5

JJourneysreal user flows in a real browser: first use, the sync click, the maker/checker approval journey, notifications, error and empty states

j loads data

1.00

Whether the Meridian Payments Console renders any payment rows at all in a real browser, and whether the total it claims on the page matches the collection truth at probe time.

50 rows rendered, DOM claims 12290 (want 12290)

j console clean

1.00

Whether normal use of the console — loading it and running a sync — produces JavaScript console errors or uncaught page errors.

0 console error(s) across load+sync

j first use

1.00

The first-visit experience as one property: real data appears quickly, the on-page total agrees with the collection truth, and the console stays clean.

ttfd=43ms reconcile=1.0 consoleErrs=0

j sync journey

0.75

The headline interactive flow: a user clicks Sync now and the app visibly starts, indicates progress, finishes, and refreshes the view.

button_found, in_flight_state, completed

j workflow journey

0.71

The maker/checker approval workflow driven end to end through the UI alone: token in, draft created, submitted, approved, listed, notified, and the vendor-created payment landing in the payments table.

roleTokenAccepted, draftCreated, submitCausal, draftListStates, notificationSeen

j workflow reject

1.00

The checker's other half of the workflow: rejecting a submitted draft completes through the UI, the rejected state shows in the draft list, and the rejection notification appears in the feed.

reject_completed, rejected_state_listed, reject_notification

j notifications feed

1.00

The notifications feed visibly degrades while notifierd is down and heals itself — without a reload — once it returns.

partition state=degraded, heal state=live, live in 0.3s

j error state

0.30

What a user sees when the backend is unreachable: a visible, actionable error state instead of a blank or silently broken page.

error indication only

j empty state

1.00

What a fresh install looks like before the first sync completes: an honest empty-or-progress state rather than phantom data or a blank page.

pre-sync state rendered: {'tablePresent': True, 'renderedRowCount': 0, 'progressText': 'No payments match the active filter.', 'emptyWithProgress': True, 'blocked': False, 'timeline': [{'tMs': 10, 'rows': 0, 'emptyText': 'No payments match the active filter.'}]}

VRendered truthwhat the page visibly shows: human dates, exponent-correct money, status badges at the frozen hexes, 375 px layout, deliberate styling

v dates readable

1.00

Whether the Date column shows humans a readable date instead of a raw machine timestamp.

rendered dates like Oct 14, 2026, 01:12

v money presentation

1.00

Whether rendered amounts identify their currency — the presentation half of money; the exponent-and-digits truth is graded separately by the critical b_money_rendered.

12/12 amount cells carry a recognizable currency token

v status badges

1.00

Whether the four payment statuses are visually distinguishable at a glance and painted in the spec's frozen palette — the same four hexes the 3D field uses.

4 statuses on page 1, 4 at the frozen hex, 4 distinct styles

v responsive 375

1.00

Whether the page survives a phone-width viewport: no horizontal scrolling and real content still rendered at 375 px.

no h-scroll, 50 rows (tap targets not measured by this probe)

v styling

1.00

Whether the page is deliberately styled rather than browser-default: a real stylesheet, layered surface colors, a chosen font, and a branded header.

stylesheet=True, 6 backgrounds, font '-apple-system, "system-ui", "S', header=True

PPerformancethe spec's budgets, measured: frames under drag, idle flatness (demand rendering), stream apply, API latency under load, sync wall clock

p drag frames

1.00

The 3D field stays interactive under input: real frames rendered during the scripted 40-move budget drag at the full 12,288-instance count.

40 frames over the 40-move budget drag (13004.8ms)

p idle flatness

1.00

Demand rendering, the frozen spec rule: at rest — no input, no coast, no pending stream batch — the scene draws nothing; a continuous rAF render loop fails by design.

0 default-FBO draws in the worst 500ms rest window (2 sampled)

p stream apply

1.00

A live SSE batch becomes visible fast: the median time from receipt to applied — store, digest and pixels — against the spec's 250 ms budget.

batch visible in 146.2 ms

p under stream

1.00

The API stays fast AND correct while the SSE stream burst is landing — read p95 under proven concurrent load.

p95=19.1ms with readers during the stream burst (overlap 0.96)

p api latency

1.00

The read endpoints answer within their spec budgets when idle: the worst p95 across the API latency battery.

worst idle p95 3.5818328615278006 ms across ['payments', 'summary']

p sync wall

1.00

Wall-clock time for the self-driven first sync — the full 192-page walk, seeded faults and documented waits included — against the spec's 120 s budget.

sync #1 wall 7548 ms (self-driven, faults included)

T3D fieldthe instanced WebGL field: real context, scene math, GPU pick buffer, camera and coast physics, collision-culled labels, brush, streaming diffs

t context real

1.00

The #viz3d panel is a real, drawing WebGL surface with the pinned context attributes and a correctly sized backing store — not a styled div, an image, or an unused canvas.

ctx=1 type=webgl2 nonBg=77/77 distinct=11 backingOk=True

t layout basis

1.00

The locked layout basis vs7dbg.layout() reports — d0 (first day), D0 = 96 (the span), R0 (max in-day count at load) — matches the fixture truth and never moves when a streamed create arrives.

layout got={'d0': '2026-10-14', 'D0': 96, 'R0': 176} want={'d0': '2026-10-14', 'D0': 96, 'R0': 176} unmoved_after_stream=True

t scene binding

1.00

The rendered scene actually encodes the payment data: vs7dbg.sceneDigest()'s seven statistical moments over all 12,288 instanced columns match an independent recomputation from the fixture.

7/7 digest moments within tolerance (count got=12291 want=12291)

t height pixels

1.00

Column heights are true in rendered pixels — including the JPY (exponent 0) and KWD (exponent 3) instances whose heights expose a forgotten currency exponent.

6/6 instance tops within ±3 px (currencies ['EUR', 'JPY', 'KWD', 'USD'])

t draw budget

1.00

The field renders 12,288 instances inside the draw budget — at most 8 default-framebuffer draw calls per rendered frame — forcing instanced draws or a merged buffer instead of per-column draws.

ΔD=40 ΔF=40 over M=40 moves (limits 320/384)

t pick buffer

1.00

Picking is GPU truth: the offscreen pick buffer answers occlusion exactly as the depth buffer says, on click points constructed to kill CPU raycasts and last-drawn-wins shortcuts.

6/6 decisive picks agree three ways; constructions {'occluded-by-lower-n': True, 'occluded-by-higher-n': True, 'partial-occlusion': True, 'background-in-hull': True}

t pick real pass

1.00

The pick buffer is real GPU work with the documented cost profile: a fresh offscreen pass after each scene invalidation, bounded at 4 offscreen draws, and never flashing ID colors on the visible canvas.

since invalidation: 1 offscreen draws, 1 readbacks; 0 default-FBO draws across 6 pick calls

t click semantics

1.00

Canvas clicks follow the frozen §3.3 semantics: click an instance to toggle it into the brush, click it again to toggle it out, click background to clear the brush.

{'click_toggles_on': True, 'click_toggles_off': True, 'background_clears': True}

t camera math

1.00

The orbit camera implements the documented math exactly: defaults yaw 30 / pitch 40 / distance 260, the 0.30 deg-per-px drag law, the exponential wheel law without page scroll, both clamps, double-click reset, and a projection that matches the printed formula.

projErr=1px pitch@clamp=85 defaults, drag_law, wheel_law, pitch_clamp, dist_clamp, dblclick_reset, projection

t coast identity

0.25

Post-release inertia obeys the closed-form τ = 0.4 s decay law: at any coasting instant the remaining travel equals v(t)·τ, a slow release starts no coast, and the coast settles inside the printed budget.

identity 0.00 over 2 mid-coast samples, slow=True (drift 0°), settle=False (Nonems of 1635ms)

t coast reality

1.00

The coast is real motion on the canvas, not a camera() narrative: after a fast flick the scene provably keeps moving past release.

coasted 3.465° past release, direction=True, pixels moved 5/5, rest pixel ok=False

t labels culling

1.00

The 12 highest-amount records get screen-space labels that are collision-culled deterministically THROUGH the app's own pick buffer — exact set, exact geometry, zero overlap, never floating over an occluded instance.

setMatch=True at the decisive pose (expected 4, dom 4), geom 1.00, overlapViolations=0; SB7.1: zero label offset is valid evidence

t brush link

0.78

One brush set links the table and the 3D field in both directions: row clicks and instance clicks toggle the same set, non-members dim to the exact 0.30 pixel rule, the count readout tracks, and background click restores full color.

row_click_toggles, dim_pixels, member_pixels, brush_count, instance_click_toggles, background_clears, pixels_restored

t stream diff

1.00

SSE batches apply as true diffs: only the changed instances upload, no buffer realloc, the digest moves by exactly the change, and the changed instance's pixels show it.

1 batches: bytes 1.00, reallocs none, digest 1.00 (1 graded), changed pixel True

t vs7dbg truth

0.70

The mandated window.vs7dbg instrumentation tells the truth: camera() agrees with the pixels, sceneDigest() with the recomputed data, frames() with the wrapper's counted draws, and pick() with pickPixel() with the analytic answer.

cameraYawErr=0 restPixel=False digest 1.0 framesAgree=True pickTriplet=True

XConsistency ledgerthe live consistency contract: no invented states, per-key version order, monotonic reads, convergence, money conservation, dup/forgery handling

x l1 no invented states

1.00

Every (payment, version) the app ever applied or served exists in the vendor's committed history — an invented state means the app fabricated data.

14776 observations, 0 invented states

x l2 per key order

1.00

Applied versions per payment strictly increase in event-log order — duplicate and stale webhook outcomes belong in the counters, never as events.

0 order violations over 43 same-key event pairs

x l3 monotonic reads

1.00

The version served for a payment never decreases from one read to the next — a sync page landing after a webhook applied v+1 must not regress the row.

0 regressions over 2442 sampled read pairs

x l4 convergence

1.00

At quiescence every payment's version and status equals the vendor's final committed state, and the row count equals the vendor's — the mid-walk create present exactly once.

24/27 touched rows at vendor-final state, total 12291 (want 12291)

x l5 group atomicity

0.00

No read ever observes half a transaction group: the refunded payment and its reversal become visible together or not at all.

26 confirmed half-applied group observations over 815 samples

x m1 amount immutability

1.00

No served row ever shows an amount_minor different from the vendor-committed amount — v3 never mutates amounts, only status/note/version.

2469 amount observations, 0 mutated

x m2 pair conservation

0.00

At every observed instant, per currency, the summary's reversal totals equal the refunded rows' amounts — both halves of each refund visible, or neither.

788/815 snapshots pair-conserved, 26 confirmed half states

x m3 terminal conservation

1.00

Terminal per-currency counts and totals — reversals included — equal vendor ground truth: fixture plus scripted mutations plus every payment the app created.

4/4 currency totals exact, 4/4 reversal totals exact at quiescence

x m4 no cross currency

1.00

No field anywhere in the summary carries a cross-currency money sum — minor units are not a common denomination and summing them is wrong money.

no cross-currency money sum anywhere

x conservation residual

1.00

CRITICAL — after every duplicate and loss is attributed, no minor units remain created or destroyed: the money conservation residual is zero in every currency.

residual 0 in every currency after attribution

x no lost write

1.00

CRITICAL — every mutation the app acknowledged with a 2xx is present in the final state; an acked-then-vanished write is silent data loss.

23 acked webhook mutations checked, 0 lost

x ooo dup forged

1.00

The three webhook trust traps land correctly: the out-of-order pair keeps v+2 (the late v+1 never overwrites), the forged-signature event changes nothing, and no duplicate delivery applies twice.

{'ooo_kept_v2': True, 'forged_untouched': True, 'dup_applied_once': True}

RResiliencethe seeded fault schedule: SIGKILL mid-sync, vendor-down boot, outbox atomicity, partition catch-up, exactly-once effects, workflow durability

r b3 sigkill resync

1.00

ledgerd SIGKILLed mid-sync restarts, converges, and duplicates nothing — a kill mid-walk costs a clean cursor restart, never dupes or holes.

kill after list #3: restart=True converge=True nodupes=True

r b4 vendor down boot

1.00

The app boots with the vendor down: binds within 10 seconds anyway, serves local data, does not crash, and completes the first sync unprompted once the vendor returns.

{'armed': True, 'refusal_secs': 5, 'bound_in_10s': True, 'served_while_down': True, 'no_crash': True, 'recovered': True}

r b6 outbox atomic

1.00

A SIGKILL between an outbox commit and its delivery loses nothing and doubles nothing — the exact window where commit-then-POST and POST-then-commit both fail.

{'pending_before_kill': True, 'resumed': True, 'exactly_once': True, 'none_lost': True}

r b7 partition

1.00

A notifier partition degrades visibly and heals in order: writes never block, /api/outbox/status reports down with growing pending, the feed shows degraded, and catch-up is in seq order with the UI live again within 5 seconds of heal.

{'ledger_alive': True, 'writes_fast': True, 'status_down': True, 'ui_degraded': True, 'catchup_in_order': True, 'ui_live_5s': True}

r notifier exactly once

1.00

The notifier's durable processed set proves exactly-once: every outbox-crossing event processed, no seq twice, surviving kills.

2970 processed (2970 unique), crossing coverage 1.00

r notification multiset

0.75

Selective materialization is exact: draft.submitted, draft.approved, draft.rejected and reversal.created each produce exactly one notification row; payment.sent produces none.

got {'draft.rejected': 1, 'draft.submitted': 3, 'draft.approved': 2, 'reversal.created': 194} want {'reversal.created': 2962, 'draft.submitted': 3, 'draft.approved': 2, 'draft.rejected': 1}

r no row loss

1.00

CRITICAL — no committed row goes missing after any seeded kill: the row count after every SIGKILL+restart is at least the count before it.

4 kills; no rows lost

r no dupe effect

1.00

CRITICAL — no ledger effect applies twice: one vendor payment per approved draft, no duplicate notification rows, no (type, payment, version) event effect repeated.

no duplicated effects

r cache truth

1.00

CRITICAL — the lying-304 trap: a 304 whose X-Collection-Generation disagrees with the stored generation is a cache miss, and serving the stale cache as fresh is graded data loss.

post-304 state reflects the committed mutations

r workflow durability

1.00

CRITICAL — submitted and approved are durable the moment their 200 is written: a SIGKILL immediately after either, including mid-send, must find the state intact after restart.

A1 -> submitted, A2 -> sent

EExcellencethe last 12%: drag frames, stream-apply latency, latency under load, optimistic paint, and mastery of the T+X+R mechanisms

e frames under drag

1.00

Excellence-grade fluidity: frames rendered during the scripted 40-move drag at the full 12,288-instance count, with proof the frames actually drew.

40 frames over the 40-move drag at N=12288

e stream apply latency

0.75

Excellence-grade streaming: the median SSE batch apply time against rungs 2.5× tighter than the P-tier budget.

batch visible in 146.2 ms (excellence rungs)

e under load latency

1.00

The read API's p95 under the proven stream-burst load, graded again in the excellence slice.

p95=19.1ms under the stream burst (overlap 0.96)

e optimistic paint

0.00

The workflow UI paints optimistically: the submitted state appears while the write is provably still on the wire, then really saves.

state did not paint while the write was provably held

e mastery

1.00

Excellence includes the mechanisms, not only the surface: mastery of the entire 3D contract (T), consistency-and-money invariants (X) and resilience (R) tiers together.

T+X+R mean 0.905

S3D structurePayment scene geometry and data mapping

s visible surface

1.00

Visible canvas, unobscured samples and independently predicted seeded-payment pixels from this browser

s tower geometry

1.00

Four currencies: pedestal, shaft, cap and shoulder voids match independent stepped geometry

s currency collar

1.00

Currency-specific collar widths match projected pixel surfaces

QPresentationReadability and interaction

q payment context

1.00

Inspector identity, currency, money, status and version match the backend

q legible presentation

0.75

Visible child text: anatomy, payment context, controls and part callouts are readable

MAnimationPayment state transitions

m committed event replay

0.00

Actual vendor-backed note update triggers collar motion; replay reproduces it without camera movement

Graded browser recording

Actual Gemini pilot: payment inspection, live update and replay checks. Includes observed failures.

Recording SHA-256: ba9a26e2d27c36cd2b6197b624f963b3ddd0024b5ea672ebe064958058834fea

Screenshots

Captured by the render gate during the run's repair rounds. The first capture is the initial render; the remainder are from the final epoch.

1Gemini-built payments console after synchronization
2Gemini-built data-backed payment tower field
3Payment inspection with currency and status context
4Recorded committed-update inspection
5Gemini-built console at mobile width

Run details

Model
gemini-3.8-flash
goose build
3f6f7ac1b

Notes

SB7.1 pilot, not a stable release or an unrestricted model ranking. Actual model: gemini-3.8-flash. App-launched single-agent session cloud-b3c625d6-1c5e-4c3f-8b5b-e2b154d7ae59, 20 September 2026. Candidate source was preserved unchanged. Final external CLI scorer commit 3f6f7ac1b; fixture seed 5cd00e961d80464f. Canonical report SHA-256 cbb5c823f319011f84d13fb537429ca61d2a4769ef797fac11a18464e400ed14. Model build: 554 seconds (9m14s), from the isolated session timestamps. Final grading pass: 226.297 seconds (3m46s). Wall time shown on this card is model build time, excluding operator debugging and rejected grading attempts. Earlier scorer failures were corrected and the unchanged artifact was rescored. Visible payment towers, geometry and currency mapping passed. The 0.699 ceiling is caused by scene-truth/rest-pixel failures, independently traced to fixed 1/60-second camera coast steps instead of actual elapsed frame time. The generic admission reason below says structure/mapping; this is not a finding that tower parts were absent. Raw earned credit: 0.7968. Actual defects: a stale event visibly regresses v4 to v3, and atomicity/conservation probes observed 26 half-applied backend states in 815 samples. Presentation earned partial credit. Remaining measurement caveats: annotation exit checks wrongly demand the hidden attribute when empty content can also be invisible; RAF-to-draw trajectory timing needs further validation. Those M checks have zero weight in raw credit; stale-event regression independently prevents animation excellence. The candidate sandbox blocked its Chrome self-test. External grading succeeded with no missing infrastructure or unavailable probes, but this limited developer-tools profile must be disclosed. Shared-host load also limits performance comparisons. Screenshots and video are from the actual graded Gemini app, not the reference. Recorded usage: 9,596,417 input tokens including 5,702,519 cached input; 75,873 candidate output tokens; 9,731,669 total tokens. Provider output accounting omits thinking. Reconstructed cost is approximately $3.63–$3.86 at the standard introductory rates checked on 20 September 2026; not a verified invoice, not zero cost, and not a proven saving against SB7. Pricing source: https://ai.google.dev/gemini-api/docs/pricing .