Cloud baseline

DeepSeek V4 Pro — 0.3844 on Gauntlet 7.0

deepseek-v4-pro · scorer sb-7.0

Aug 24, 2026
Overall
0.3844
Excellent
no
Benchmark
Gauntlet 7.0
Recorded duration
80m 53s
Prompt tok
107.1M
Gen tok
363.0k

Tier breakdown

A · structure & runtime1.00
B · behaviour1.00
C · vendor contract0.81
D · finesse0.79

Scoring detail

How this score was built

Per-check scorer output, as posted by the app. Expand a section to inspect its checks.

The number, exactly

(0.88 × 0.7928 core + 0.12 × 0.77 gate × 0.28 excellence) × 0.5314 critical = 0.3844

Critical defects compound a multiplier on the whole score (pre-severity 0.7233):

j_workflow_journey 0.71 → ×0.89r_workflow_durability 0.00 → ×0.60

The core (88% of the total) is the weighted mean of the ten measured tiers. The last 12% is the excellence slice: it unlocks in proportion to the perfection conditions below (11 of 16 met here), then pays out at the excellence tier's own measured mean. Core 0.6977 + excellence 0.0257.

j_first_use 0.60j_workflow_journey 0.71j_error_state 1.00j_empty_state 1.00console_clean 8.00v_responsive_375 1.00v_dates_readable 1.00t_scene_binding 1.00x_conservation_residual 1.00r_no_row_loss 1.00p_drag_frames 0.00p_idle_flatness 1.00p_stream_apply 0.00p_under_stream 1.00p_api_latency 1.00p_sync_wall 1.00

Run notes and corrections

Cloud baseline — a single goose run session on the frozen sb-7 spec (Meridian Payments Console: two services, 12,288-payment collection, signed webhooks, maker/checker approval workflow, instanced WebGL 3D field), scored by the frozen sb-7 scorer. Screenshots are the scorer's own render-gate captures of the built app. Full sb-7 tier means (the schema's tierA–tierD carry only A–D): A 1.0000 · B 1.0000 · C 0.8095 · D 0.7857 · J 0.7592 · V 1.0000 · P 0.6667 · T 0.5157 · X 0.8333 · R 0.8300 · E 0.2141. Smoke qualification provenance: predecessor-carried, not fresh. This result's full build and hermetic score were carried unchanged from campaign cloud-sb7-20260823-live-r2; its contract proof is that predecessor campaign's sealed PASS smoke proof. The successor's fresh smoke attempts made zero provider admissions and were blocked by the unchanged frozen campaign/provider budget envelope.

Token rates

Measured by the engine itself, one record per completed model call: prefill rate is prompt tokens over time-to-first-token, generation rate is completion tokens over the decode window. Medians per node.

NodeCallsPrompt tokGen tokPrefill tok/sGen tok/s
deepseek379107,061,481362,962102943.2114.4
fleet379107,061,481362,962102943.2114.4

Screenshots

Recorded views of the built application. Captions identify the available captures; the number of images does not imply a number of repair rounds.

1First render — before repairs
2Final render
33D field · 12,288 instanced columns · WebGL
4Approval workflow · draft approved
5After sync

Run details

Model
deepseek-v4-pro