Cloud baseline

~openai/gpt-terra-latest, single model via OpenRouter — 0.5781 on Gauntlet 7.2

~openai/gpt-terra-latest · scorer sb-7.2

~openai/gpt-terra-latest · single agentOct 6, 2026
Overall
0.5781
Excellent
no
Benchmark
Gauntlet 7.2
Model build
5m 36s
Prompt tok
2.1M
Gen tok
29.9k

Tier breakdown

A · structure & runtime1.00
B · behaviour0.94
C · vendor contract0.93
D · finesse0.93

Scoring detail

How this score was built

Per-check scorer output, as posted by the app. Expand a section to inspect its checks.

Earned credit before admission

(0.88 × 0.6568 core + 0.12 × 0.71 gate × 0.49 excellence) × 0.9429 critical = 0.5837

Critical defects compound a multiplier on the whole score (pre-severity 0.6191):

j_workflow_journey 0.86 → ×0.94

The core (88% of the total) is the weighted mean of the thirteen measured tiers. The last 12% is the excellence slice: it unlocks in proportion to the perfection conditions below (10 of 16 met here), then pays out at the excellence tier's own measured mean. Core 0.5780 + excellence 0.0411.

j_first_use 1.00j_workflow_journey 0.86j_error_state 0.00j_empty_state 1.00console_clean 7.00v_responsive_375 0.00v_dates_readable 1.00t_scene_binding 0.43x_conservation_residual 1.00r_no_row_loss 1.00p_drag_frames 1.00p_idle_flatness 1.00p_stream_apply 0.00p_under_stream 1.00p_api_latency 1.00p_sync_wall 1.00

Earned score 0.584 · Admission ceiling 0.599 · Final score 0.578

Final score is the lower of earned credit and the admission ceiling. Passing admission adds no points.

Visible: not metMatching: not metGood: not metVisual excellence: not metBackend excellence: not met

Visible data-backed 3D: s_visible_surface (maximum 0.599)

Tower structure, currency mapping and scene truth: s_visible_surface, s_tower_geometry, s_currency_collar, t_scene_binding, t_height_pixels, t_vs7dbg_truth (maximum 0.599)

3D interaction and overview legibility: t_pick_buffer, t_pick_real_pass, t_click_semantics, t_camera_math, t_coast_identity, t_coast_reality, t_labels_culling, t_brush_link, t_stream_diff, q_overview_legibility, q_inspector_framing (maximum 0.699)

Event animation and backend recovery: m_committed_event_replay, r_b6_outbox_atomic (maximum 0.869)

Token rates

Measured by the engine itself, one record per completed model call: prefill rate is prompt tokens over time-to-first-token, generation rate is completion tokens over the decode window. Medians per node.

NodeCallsPrompt tokGen tokPrefill tok/sGen tok/s
gemini197251089.769.4
gpt482,087,94729,91715434.41390.4
fleet492,088,91929,92215147.9545.5

Graded browser recording

Loading duration…

Watch an excerpt of the generated application during browser grading. The check results and screenshots on this page provide the wider evidence. Pause or scrub to inspect a view.

Each tower represents a payment: height encodes its amount, cap color its status, and the collar its currency. The selected detailed tower is expected to animate its collar after a committed backend update or replay. These are the task requirements; the check results show what this app actually achieved.

Ready to play0:00 / —:—
Open recording file

Drag the timeline to seek. Keyboard: arrow keys move 5 seconds; Home and End jump to the start and end.

Graded browser recording re-encoded to 946x592 at 25 fps to fit 4 MiB, same 461.9 s; graded original sha256 cc55f3dad3d15571e838667ad7123e550754bf3a505f117158854e47f3acc65b
Recording integrity

SHA-256: 70910f77d40018f079f9684fb9d5809f094be3f732cdfdd25f957b3a8d4846ad

Screenshots

Captured from the built application during browser grading. Captions identify the recorded views and checks; these images do not represent repair rounds.

1Initial app view
2Error state
3Payment workflow
4First captured render
5Latest captured render

Run details

Model
~openai/gpt-terra-latest
Engine events
0
Repair rounds
0
Started
Oct 6, 2026, 05:39 PM
Finished
Oct 6, 2026, 06:06 PM