Cloud baseline

openai/gpt-6-luna-pro, single model via OpenRouter — 0.1342 on Gauntlet 7.2

openai/gpt-6-luna-pro · scorer sb-7.2

openai/gpt-6-luna-pro · single agentOct 3, 2026
Overall
0.1342
Excellent
no
Benchmark
Gauntlet 7.2
Model build
53m 36s

Tier breakdown

A · structure & runtime1.00
B · behaviour0.39
C · vendor contract0.16
D · finesse0.87

Scoring detail

How this score was built

Per-check scorer output, as posted by the app. Expand a section to inspect its checks.

Earned credit before admission

(0.88 × 0.3682 core + 0.12 × 0.58 gate × 0.69 excellence) × 0.3606 critical = 0.1342

Critical defects compound a multiplier on the whole score (pre-severity 0.3721):

sync_completeness 0.00 → ×0.60j_workflow_journey 0.00 → ×0.60

b_money_rendered: no additional penalty (vacuous:sync_completeness); wrong money — wrong exponent/digits or a cross-currency sum.

b_buckets_dst: no additional penalty (vacuous:sync_completeness); Incorrect daily bucket counts in the incomplete initial dataset (0/12289 payments).

x_conservation_residual: no additional penalty (vacuous:sync_completeness); wrong money — unexplained minor units created/destroyed after dupe/loss attribution.

x_no_lost_write: no additional penalty (vacuous:sync_completeness); wrong money — an acknowledged mutation absent from final state.

r_no_dupe_effect: no additional penalty (vacuous:sync_completeness); wrong money — a ledger effect applied twice.

r_cache_truth: no additional penalty (vacuous:sync_completeness); data loss — 304-vs-cache mismatch served as fresh.

The core (88% of the total) is the weighted mean of the thirteen measured tiers. The last 12% is the excellence slice: it unlocks in proportion to the perfection conditions below (9 of 16 met here), then pays out at the excellence tier's own measured mean. Core 0.3240 + excellence 0.0481.

j_first_use 0.00j_workflow_journey 0.00j_error_state 0.30j_empty_state 1.00console_clean 1.00v_responsive_375 1.00v_dates_readable 1.00t_scene_binding 0.00x_conservation_residual 0.00r_no_row_loss 1.00p_drag_frames 1.00p_idle_flatness 1.00p_stream_apply 1.00p_under_stream 1.00p_api_latency 1.00p_sync_wall 0.00

Earned score 0.134 · Admission ceiling 0.599 · Final score 0.134

Final score is the lower of earned credit and the admission ceiling. Passing admission adds no points.

Visible: not metMatching: not metGood: not metVisual excellence: not metBackend excellence: not met

Visible data-backed 3D: s_visible_surface (maximum 0.599)

Tower structure, currency mapping and scene truth: s_visible_surface, s_tower_geometry, s_currency_collar, t_layout_basis, t_scene_binding, t_height_pixels, t_vs7dbg_truth (maximum 0.699)

3D interaction and overview legibility: t_pick_buffer, t_click_semantics, t_camera_math, t_coast_reality, t_labels_culling, t_brush_link, t_stream_diff, q_overview_legibility, q_inspector_framing (maximum 0.799)

Event animation and backend recovery: m_committed_event_replay, x_l1_no_invented_states, x_l2_per_key_order, x_l3_monotonic_reads, x_l4_convergence, x_l5_group_atomicity, x_m1_amount_immutability, x_m2_pair_conservation, x_m3_terminal_conservation, x_conservation_residual, x_no_lost_write, x_ooo_dup_forged, r_b3_sigkill_resync, r_b4_vendor_down_boot, r_no_dupe_effect, r_cache_truth (maximum 0.899)

Graded browser recording

Loading duration…

Watch the full graded browser recording at its original speed: the payment field, structure inspection, committed updates and replay. The check results and screenshots on this page provide the wider evidence. Pause or scrub to inspect a view.

Each tower represents a payment: height encodes its amount, cap color its status, and the collar its currency. The selected detailed tower is expected to animate its collar after a committed backend update or replay. These are the task requirements; the check results show what this app actually achieved.

Ready to play0:00 / —:—
Open recording file

Drag the timeline to seek. Keyboard: arrow keys move 5 seconds; Home and End jump to the start and end.

Full graded browser recording: payment field, structure inspection, committed updates and replay
Recording integrity

SHA-256: 470fa4ac4e7c3a62e0c3a37ab564a18b7768ffea725481cd98dc8e6db9b56ba6

Screenshots

Captured from the built application during browser grading. Captions identify the recorded views and checks; these images do not represent repair rounds.

1Initial app view
2Error state
3Payment workflow
4First captured render
5Latest captured render

Run details

Model
openai/gpt-6-luna-pro
Engine events
0
Repair rounds
0
Started
Oct 2, 2026, 10:31 PM
Finished
Oct 3, 2026, 09:53 AM