Cloud baseline

openai/gpt-6.1-sol, single model via OpenRouter — 0.7990 on Gauntlet 7.2

openai/gpt-6.1-sol · scorer sb-7.2

openai/gpt-6.1-sol · single agentOct 3, 2026
Overall
0.7990
Excellent
no
Benchmark
Gauntlet 7.2
Model build
26m 35s
Prompt tok
5.4M
Gen tok
56.4k

Tier breakdown

A · structure & runtime1.00
B · behaviour0.98
C · vendor contract0.93
D · finesse0.93

Scoring detail

How this score was built

Per-check scorer output, as posted by the app. Expand a section to inspect its checks.

Earned credit before admission

(0.88 × 0.9149 core + 0.12 × 0.96 gate × 0.90 excellence) × 1.0000 critical = 0.9084

The core (88% of the total) is the weighted mean of the thirteen measured tiers. The last 12% is the excellence slice: it unlocks in proportion to the perfection conditions below (15 of 16 met here), then pays out at the excellence tier's own measured mean. Core 0.8051 + excellence 0.1032.

j_first_use 1.00j_workflow_journey 1.00j_error_state 0.30j_empty_state 1.00console_clean 0.00v_responsive_375 1.00v_dates_readable 1.00t_scene_binding 1.00x_conservation_residual 1.00r_no_row_loss 1.00p_drag_frames 1.00p_idle_flatness 1.00p_stream_apply 1.00p_under_stream 1.00p_api_latency 1.00p_sync_wall 1.00

Earned score 0.908 · Admission ceiling 0.799 · Final score 0.799

Final score is the lower of earned credit and the admission ceiling. Passing admission adds no points.

Visible: passedMatching: passedGood: not metVisual excellence: not metBackend excellence: not met

3D interaction and overview legibility: t_coast_identity, t_coast_reality, q_overview_legibility, q_inspector_framing (maximum 0.799)

Event animation and backend recovery: x_l5_group_atomicity, x_m2_pair_conservation, r_b3_sigkill_resync, r_notification_multiset (maximum 0.899)

Token rates

Measured by the engine itself, one record per completed model call: prefill rate is prompt tokens over time-to-first-token, generation rate is completion tokens over the decode window. Medians per node.

NodeCallsPrompt tokGen tokPrefill tok/sGen tok/s
gemini197251092.183.3
gpt1055,401,61256,3855839.54078.9
fleet1065,402,58456,3905618.73233.9

Graded browser recording

Loading duration…

Watch the full graded browser recording at its original speed: the payment field, structure inspection, committed updates and replay. The check results and screenshots on this page provide the wider evidence. Pause or scrub to inspect a view.

Each tower represents a payment: height encodes its amount, cap color its status, and the collar its currency. The selected detailed tower is expected to animate its collar after a committed backend update or replay. These are the task requirements; the check results show what this app actually achieved.

Ready to play0:00 / —:—
Open recording file

Drag the timeline to seek. Keyboard: arrow keys move 5 seconds; Home and End jump to the start and end.

Full graded browser recording: payment field, structure inspection, committed updates and replay
Recording integrity

SHA-256: 6fb8ea230c049678ec1d976c7932af979dad8356ffedbc873afff6a8026c4753

Screenshots

Captured from the built application during browser grading. Captions identify the recorded views and checks; these images do not represent repair rounds.

1Initial app view
2Error state
3Payment workflow
4First captured render
5Latest captured render

Run details

Model
openai/gpt-6.1-sol
Engine events
0
Repair rounds
0
Started
Oct 3, 2026, 07:25 AM
Finished
Oct 3, 2026, 08:02 AM