Cloud baseline

GPT-6 Luna, single model via OpenRouter — 0.4711 on sb-7.2

openai/gpt-6-luna · scorer sb-7.2

openai/gpt-6-luna · single agentOct 2, 2026
Overall
0.4711
Excellent
no
Scorer
sb-7.2
Model build
33m 30s
Prompt tok
12.9M
Gen tok
171.3k

Tier breakdown

A · structure & runtime1.00
B · behaviour1.00
C · vendor contract0.93
D · finesse0.93

Scoring detail

How this score was built

Per-check scorer output, as posted by the app. Expand a section to inspect its checks.

Earned credit before admission

(0.88 × 0.8105 core + 0.12 × 0.92 gate × 0.65 excellence) × 0.6000 critical = 0.4711

Critical defects compound a multiplier on the whole score (pre-severity 0.7851):

j_workflow_journey 0.00 → ×0.60

The core (88% of the total) is the weighted mean of the thirteen measured tiers. The last 12% is the excellence slice: it unlocks in proportion to the perfection conditions below (14 of 16 met here), then pays out at the excellence tier's own measured mean. Core 0.7132 + excellence 0.0719.

j_first_use 1.00j_workflow_journey 0.00j_error_state 1.00j_empty_state 1.00console_clean 0.00v_responsive_375 1.00v_dates_readable 1.00t_scene_binding 1.00x_conservation_residual 1.00r_no_row_loss 1.00p_drag_frames 1.00p_idle_flatness 1.00p_stream_apply 1.00p_under_stream 0.75p_api_latency 1.00p_sync_wall 1.00

Earned score 0.471 · Admission ceiling 0.699 · Final score 0.471

Final score is the lower of earned credit and the admission ceiling. Passing admission adds no points.

Visible: passedMatching: not metGood: not metVisual excellence: not metBackend excellence: not met

Tower structure, currency mapping and scene truth: s_tower_geometry, s_currency_collar, t_vs7dbg_truth (maximum 0.699)

3D interaction and overview legibility: t_click_semantics, t_brush_link, q_overview_legibility, q_inspector_framing (maximum 0.799)

Event animation and backend recovery: m_committed_event_replay, r_notification_multiset (maximum 0.899)

Token rates

Measured by the engine itself, one record per completed model call: prefill rate is prompt tokens over time-to-first-token, generation rate is completion tokens over the decode window. Medians per node.

NodeCallsPrompt tokGen tokPrefill tok/sGen tok/s
gemini197251641.951.0
gpt15012,909,185171,31922735.07000.0
fleet15112,910,157171,32422734.07000.0

Graded browser recording

Loading duration…

Watch the full graded browser recording at its original speed: the payment field, structure inspection, committed updates and replay. The check results and screenshots on this page provide the wider evidence. Pause or scrub to inspect a view.

Each tower represents a payment: height encodes its amount, cap color its status, and the collar its currency. The selected detailed tower is expected to animate its collar after a committed backend update or replay. These are the task requirements; the check results show what this app actually achieved.

Ready to play0:00 / —:—
Open recording file

Drag the timeline to seek. Keyboard: arrow keys move 5 seconds; Home and End jump to the start and end.

Full graded browser recording: payment field, structure inspection, committed updates and replay
Recording integrity

SHA-256: b72871083fc1a5c5468fe0339167a270e4d9b3c1ab4974667e04820f3de52dcb

Screenshots

Captured from the built application during browser grading. Captions identify the recorded views and checks; these images do not represent repair rounds.

1Initial app view
2Error state
3Payment workflow
4First captured render
5Latest captured render

Run details

Model
openai/gpt-6-luna
Engine events
0
Repair rounds
0
Started
Oct 2, 2026, 03:19 PM
Finished
Oct 2, 2026, 04:02 PM