Cloud baseline

GPT-5.6 Terra — 0.3252 on Gauntlet 6.0

openai.gpt-5.6-terra · scorer sb-6.0

Aug 19, 2026
Overall
0.3252
Excellent
no
Benchmark
Gauntlet 6.0
Recorded duration
6m 30s

Tier breakdown

A · structure & runtime1.00
B · behaviour1.00
C · vendor contract1.00
D · finesse1.00

Scoring detail

How this score was built

Per-check scorer output, as posted by the app. Expand a section to inspect its checks.

The number, exactly

(0.88 × 0.8247 core + 0.12 × 0.54 gate × 0.75 excellence) × 0.4200 critical = 0.3252

Critical defects compound a multiplier on the whole score (pre-severity 0.7743):

j_loads_data 0.00 → ×0.60j_sync_journey 0.25 → ×0.70

The core (88% of the total) is the weighted mean of the nine measured tiers. The last 12% is the excellence slice: it unlocks in proportion to the perfection conditions below (7 of 14 met here), then pays out at the excellence tier's own measured mean. Core 0.7257 + excellence 0.0485.

j_first_use 0.00j_sync_journey 0.25j_error_state 1.00j_empty_state 0.30console_clean 6.00v_responsive_375 0.00v_dates_readable 0.00t_scene_binding 1.00p_list_latency 1.00p_buckets_latency 1.00p_summary_latency 1.00p_page_interactive 0.00p_first_frame 1.00p_sync_wall 1.00
ComponentWeightEarnedPoints
Core (tiers A–D)60%——
A · Structure25% of core1.000.150
B · Behaviour30% of core1.000.180
C · Vendor contract25% of core1.000.150
D · Finesse20% of core1.000.120
Journey (J)15%0.3880.058
Visual (V)10%0.1710.017
Performance (P)5%0.8330.042
Hard blocks10%——
Total0.1169

Run notes and corrections

Cloud baseline — a single goose run session on the frozen VendorSync Pro spec, scored by the sb-6 scorer with the severity model (critical defects — crash, wrong money, data loss, dead primary flow — compound a multiplier on the composed score). Serial, hermetic, advertised vendor port. NOTE: this session was terminated early by a goose engine defect (compaction shipped reasoning content to Bedrock; fixed in 19b4ed6ef) — the score is a floor for this model, not a ceiling.

Screenshots

Recorded views of the built application. Captions identify the available captures; the number of images does not imply a number of repair rounds.

1Final render
23D buckets · WebGL
3After sync
4Mobile · 375px

Run details

Model
openai.gpt-5.6-terra