Cloud baseline

stepfun/step-5-preview, single model via OpenRouter — 0.5806 on Gauntlet 7.2

stepfun/step-5-preview · scorer sb-7.2

stepfun/step-5-preview · single agentOct 9, 2026
Overall
0.5806
Excellent
no
Benchmark
Gauntlet 7.2
Model build
72m 6s
Prompt tok
32.8M
Gen tok
333.2k

Tier breakdown

A · structure & runtime1.00
B · behaviour0.98
C · vendor contract1.00
D · finesse0.90

Scoring detail

How this score was built

Per-check scorer output, as posted by the app. Expand a section to inspect its checks.

Earned credit before admission

(0.88 × 0.7382 core + 0.12 × 0.90 gate × 0.75 excellence) × 0.8674 critical = 0.6333

Critical defects compound a multiplier on the whole score (pre-severity 0.7301):

b_money_rendered 0.80 → ×0.92j_workflow_journey 0.86 → ×0.94

The core (88% of the total) is the weighted mean of the thirteen measured tiers. The last 12% is the excellence slice: it unlocks in proportion to the perfection conditions below (13 of 16 met here), then pays out at the excellence tier's own measured mean. Core 0.6496 + excellence 0.0804.

j_first_use 1.00j_workflow_journey 0.86j_error_state 0.00j_empty_state 1.00console_clean 0.00v_responsive_375 1.00v_dates_readable 1.00t_scene_binding 1.00x_conservation_residual 1.00r_no_row_loss 1.00p_drag_frames 1.00p_idle_flatness 1.00p_stream_apply 0.50p_under_stream 1.00p_api_latency 1.00p_sync_wall 1.00

Earned score 0.633 · Admission ceiling 0.599 · Final score 0.581

Final score is the lower of earned credit and the admission ceiling. Passing admission adds no points.

Visible: not metMatching: not metGood: not metVisual excellence: not metBackend excellence: passed

Visible data-backed 3D: s_visible_surface (maximum 0.599)

Tower structure, currency mapping and scene truth: s_visible_surface, s_tower_geometry, s_currency_collar, t_height_pixels, t_vs7dbg_truth (maximum 0.609)

3D interaction and overview legibility: t_pick_buffer, t_pick_real_pass, t_camera_math, t_coast_reality, t_brush_link, t_stream_diff, q_overview_legibility, q_inspector_framing (maximum 0.699)

Event animation and backend recovery: m_committed_event_replay (maximum 0.899)

Token rates

Measured by the engine itself, one record per completed model call: prefill rate is prompt tokens over time-to-first-token, generation rate is completion tokens over the decode window. Medians per node.

NodeCallsPrompt tokGen tokPrefill tok/sGen tok/s
gemini197251112.1161.3
step14932,756,945333,19147905.2141.7
fleet15032,757,917333,19647732.2142.0

Graded browser recording

Loading duration…

Watch an excerpt of the generated application during browser grading. The check results and screenshots on this page provide the wider evidence. Pause or scrub to inspect a view.

Each tower represents a payment: height encodes its amount, cap color its status, and the collar its currency. The selected detailed tower is expected to animate its collar after a committed backend update or replay. These are the task requirements; the check results show what this app actually achieved.

Ready to play0:00 / —:—
Open recording file

Drag the timeline to seek. Keyboard: arrow keys move 5 seconds; Home and End jump to the start and end.

Graded browser recording re-encoded to 568x356 at 25 fps to fit 4 MiB, same 241.3 s; graded original sha256 2d5480867e83ccdf314ff5b70306c5da327e1ad176a66c67617ad9755c1891f2
Recording integrity

SHA-256: efdf07fb6c26d2afd87da63b1d4368ddb07d1c9ba3b6f5f3323f2cd63ab49c14

Screenshots

Captured from the built application during browser grading. Captions identify the recorded views and checks; these images do not represent repair rounds.

1Initial app view
2Error state
3Payment workflow
4First captured render
5Latest captured render

Run details

Model
stepfun/step-5-preview
Engine events
0
Repair rounds
0
Started
Oct 9, 2026, 02:46 PM
Finished
Oct 9, 2026, 04:09 PM