Cloud baseline

z-ai/glm-5.3-flashx, single model via OpenRouter — 0.6990 on Gauntlet 7.2

z-ai/glm-5.3-flashx · scorer sb-7.2

z-ai/glm-5.3-flashx · single agentOct 3, 2026
Overall
0.6990
Excellent
no
Benchmark
Gauntlet 7.2
Model build
49m 53s
Prompt tok
25.1M
Gen tok
226.9k

Tier breakdown

A · structure & runtime1.00
B · behaviour1.00
C · vendor contract0.95
D · finesse0.93

Scoring detail

How this score was built

Per-check scorer output, as posted by the app. Expand a section to inspect its checks.

Earned credit before admission

(0.88 × 0.8422 core + 0.12 × 0.99 gate × 0.90 excellence) × 0.9429 critical = 0.7995

Critical defects compound a multiplier on the whole score (pre-severity 0.8480):

j_workflow_journey 0.86 → ×0.94

The core (88% of the total) is the weighted mean of the thirteen measured tiers. The last 12% is the excellence slice: it unlocks in proportion to the perfection conditions below (15 of 16 met here), then pays out at the excellence tier's own measured mean. Core 0.7411 + excellence 0.1068.

j_first_use 1.00j_workflow_journey 0.86j_error_state 1.00j_empty_state 1.00console_clean 0.00v_responsive_375 1.00v_dates_readable 1.00t_scene_binding 1.00x_conservation_residual 1.00r_no_row_loss 1.00p_drag_frames 1.00p_idle_flatness 1.00p_stream_apply 1.00p_under_stream 1.00p_api_latency 1.00p_sync_wall 1.00

Earned score 0.799 · Admission ceiling 0.699 · Final score 0.699

Final score is the lower of earned credit and the admission ceiling. Passing admission adds no points.

Visible: passedMatching: not metGood: not metVisual excellence: not metBackend excellence: not met

Tower structure, currency mapping and scene truth: s_tower_geometry, t_height_pixels, t_vs7dbg_truth (maximum 0.699)

3D interaction and overview legibility: t_camera_math, t_labels_culling, t_brush_link, t_stream_diff, q_overview_legibility, q_inspector_framing (maximum 0.799)

Event animation and backend recovery: m_committed_event_replay, x_l5_group_atomicity, x_m2_pair_conservation, r_notification_multiset (maximum 0.899)

Token rates

Measured by the engine itself, one record per completed model call: prefill rate is prompt tokens over time-to-first-token, generation rate is completion tokens over the decode window. Medians per node.

NodeCallsPrompt tokGen tokPrefill tok/sGen tok/s
gemini19725909.3138.9
z14125,073,026226,89147422.5140.6
fleet14225,073,998226,89647403.3140.6

Graded browser recording

Loading duration…

Watch an excerpt of the generated application during browser grading. The check results and screenshots on this page provide the wider evidence. Pause or scrub to inspect a view.

Each tower represents a payment: height encodes its amount, cap color its status, and the collar its currency. The selected detailed tower is expected to animate its collar after a committed backend update or replay. These are the task requirements; the check results show what this app actually achieved.

Ready to play0:00 / —:—
Open recording file

Drag the timeline to seek. Keyboard: arrow keys move 5 seconds; Home and End jump to the start and end.

Graded browser recording re-encoded to 932x582 at 25 fps to fit 4 MiB, same 111.7 s; graded original sha256 67cc29cf15b4e0a753bf17ec0106f8d8b4d038726e183e6a394d365c25861bb4
Recording integrity

SHA-256: b3fc4b49af2fca7ad313b669a7fdf06045be419ea14c05f2fdaf677dd61110c0

Screenshots

Captured from the built application during browser grading. Captions identify the recorded views and checks; these images do not represent repair rounds.

1First captured render
2Latest captured render
3After sync
4Error state
5Mobile · 375px

Run details

Model
z-ai/glm-5.3-flashx
Engine events
0
Repair rounds
0
Started
Oct 3, 2026, 02:57 AM
Finished
Oct 3, 2026, 03:55 AM