Cloud baseline

anthropic/claude-haiku-5.5, single model via OpenRouter — 0.6610 on Gauntlet 7.2

anthropic/claude-haiku-5.5 · scorer sb-7.2

anthropic/claude-haiku-5.5 · single agentOct 9, 2026
Overall
0.6610
Excellent
no
Benchmark
Gauntlet 7.2
Model build
36m 36s
Prompt tok
4.7M
Gen tok
459.4k

Tier breakdown

A · structure & runtime1.00
B · behaviour1.00
C · vendor contract1.00
D · finesse0.90

Scoring detail

How this score was built

Per-check scorer output, as posted by the app. Expand a section to inspect its checks.

Earned credit before admission

(0.88 × 0.9349 core + 0.12 × 0.90 gate × 0.65 excellence) × 0.9429 critical = 0.8417

Critical defects compound a multiplier on the whole score (pre-severity 0.8927):

j_workflow_journey 0.86 → ×0.94

The core (88% of the total) is the weighted mean of the thirteen measured tiers. The last 12% is the excellence slice: it unlocks in proportion to the perfection conditions below (13 of 16 met here), then pays out at the excellence tier's own measured mean. Core 0.8227 + excellence 0.0700.

j_first_use 1.00j_workflow_journey 0.86j_error_state 1.00j_empty_state 1.00console_clean 1.00v_responsive_375 1.00v_dates_readable 1.00t_scene_binding 1.00x_conservation_residual 1.00r_no_row_loss 1.00p_drag_frames 1.00p_idle_flatness 1.00p_stream_apply 0.50p_under_stream 1.00p_api_latency 1.00p_sync_wall 1.00

Earned score 0.842 · Admission ceiling 0.669 · Final score 0.661

Final score is the lower of earned credit and the admission ceiling. Passing admission adds no points.

Visible: passedMatching: not metGood: not metVisual excellence: not metBackend excellence: passed

Tower structure, currency mapping and scene truth: s_tower_geometry, t_vs7dbg_truth (maximum 0.669)

3D interaction and overview legibility: t_pick_buffer, t_coast_identity, t_coast_reality, q_overview_legibility (maximum 0.709)

Event animation and backend recovery: an earlier admission band is incomplete (maximum 0.899)

Token rates

Measured by the engine itself, one record per completed model call: prefill rate is prompt tokens over time-to-first-token, generation rate is completion tokens over the decode window. Medians per node.

NodeCallsPrompt tokGen tokPrefill tok/sGen tok/s
claude394,748,880459,36112756.4501.0
gemini197251466.158.8
fleet404,749,852459,36612752.3487.3

Graded browser recording

Loading duration…

Watch the full graded browser recording at its original speed: the payment field, structure inspection, committed updates and replay. The check results and screenshots on this page provide the wider evidence. Pause or scrub to inspect a view.

Each tower represents a payment: height encodes its amount, cap color its status, and the collar its currency. The selected detailed tower is expected to animate its collar after a committed backend update or replay. These are the task requirements; the check results show what this app actually achieved.

Ready to play0:00 / —:—
Open recording file

Drag the timeline to seek. Keyboard: arrow keys move 5 seconds; Home and End jump to the start and end.

Full graded browser recording: payment field, structure inspection, committed updates and replay
Recording integrity

SHA-256: b96a561c97b0eba15cc1b4a89c5603391627f23d5a86dce9dc44ba029b328543

Screenshots

Captured from the built application during browser grading. Captions identify the recorded views and checks; these images do not represent repair rounds.

1Initial app view
2Error state
3Payment workflow
4First captured render
5Latest captured render

Run details

Model
anthropic/claude-haiku-5.5
Engine events
0
Repair rounds
0
Started
Oct 9, 2026, 01:29 PM
Finished
Oct 9, 2026, 02:17 PM