Cloud baseline

z-ai/glm-5.3-flash, single model via OpenRouter — 0.4143 on Gauntlet 7.2

z-ai/glm-5.3-flash · scorer sb-7.2

z-ai/glm-5.3-flash · single agentOct 6, 2026
Overall
0.4143
Excellent
no
Benchmark
Gauntlet 7.2
Model build
46m 10s
Prompt tok
8.0M
Gen tok
76.5k

Tier breakdown

A · structure & runtime1.00
B · behaviour1.00
C · vendor contract0.97
D · finesse0.90

Scoring detail

How this score was built

Per-check scorer output, as posted by the app. Expand a section to inspect its checks.

Earned credit before admission

(0.88 × 0.7019 core + 0.12 × 0.88 gate × 0.69 excellence) × 0.6000 critical = 0.4143

Critical defects compound a multiplier on the whole score (pre-severity 0.6905):

j_workflow_journey 0.00 → ×0.60

The core (88% of the total) is the weighted mean of the thirteen measured tiers. The last 12% is the excellence slice: it unlocks in proportion to the perfection conditions below (14 of 16 met here), then pays out at the excellence tier's own measured mean. Core 0.6177 + excellence 0.0728.

j_first_use 1.00j_workflow_journey 0.00j_error_state 1.00j_empty_state 1.00console_clean 2.00v_responsive_375 1.00v_dates_readable 1.00t_scene_binding 1.00x_conservation_residual 1.00r_no_row_loss 1.00p_drag_frames 1.00p_idle_flatness 1.00p_stream_apply 1.00p_under_stream 1.00p_api_latency 1.00p_sync_wall 1.00

Earned score 0.414 · Admission ceiling 0.599 · Final score 0.414

Final score is the lower of earned credit and the admission ceiling. Passing admission adds no points.

Visible: not metMatching: not metGood: not metVisual excellence: not metBackend excellence: passed

Visible data-backed 3D: s_visible_surface (maximum 0.599)

Tower structure, currency mapping and scene truth: s_visible_surface, s_tower_geometry, s_currency_collar, t_height_pixels, t_vs7dbg_truth (maximum 0.609)

3D interaction and overview legibility: t_pick_buffer, t_pick_real_pass, t_click_semantics, t_camera_math, t_coast_reality, t_labels_culling, t_brush_link, t_stream_diff, q_overview_legibility, q_inspector_framing (maximum 0.699)

Event animation and backend recovery: m_committed_event_replay (maximum 0.899)

Token rates

Measured by the engine itself, one record per completed model call: prefill rate is prompt tokens over time-to-first-token, generation rate is completion tokens over the decode window. Medians per node.

NodeCallsPrompt tokGen tokPrefill tok/sGen tok/s
z1457,958,91876,48717492.61051.7
fleet1457,958,91876,48717492.61051.7

Graded browser recording

Loading duration…

Watch the full graded browser recording at its original speed: the payment field, structure inspection, committed updates and replay. The check results and screenshots on this page provide the wider evidence. Pause or scrub to inspect a view.

Each tower represents a payment: height encodes its amount, cap color its status, and the collar its currency. The selected detailed tower is expected to animate its collar after a committed backend update or replay. These are the task requirements; the check results show what this app actually achieved.

Ready to play0:00 / —:—
Open recording file

Drag the timeline to seek. Keyboard: arrow keys move 5 seconds; Home and End jump to the start and end.

Full graded browser recording: payment field, structure inspection, committed updates and replay
Recording integrity

SHA-256: b585e2cfa1c0b7c53623f414c9432dc27dfacae4ed4a378aab3208eb9630d7c2

Screenshots

Captured from the built application during browser grading. Captions identify the recorded views and checks; these images do not represent repair rounds.

1Initial app view
2Error state
3Payment workflow
4First captured render
5Latest captured render

Run details

Model
z-ai/glm-5.3-flash
Engine events
0
Repair rounds
0
Started
Oct 6, 2026, 01:15 AM
Finished
Oct 6, 2026, 02:08 AM