Cloud baseline

GPT-5.6 Luna — 0.8671 on Gauntlet 6.0

openai.gpt-5.6-luna · scorer sb-6.0

Aug 19, 2026
Overall
0.8671
Excellent
no
Benchmark
Gauntlet 6.0
Recorded duration
3m 57s

Tier breakdown

A · structure & runtime1.00
B · behaviour1.00
C · vendor contract1.00
D · finesse1.00

Scoring detail

How this score was built

Per-check scorer output, as posted by the app. Expand a section to inspect its checks.

The number, exactly

(0.88 × 0.9038 core + 0.12 × 1.00 gate × 0.60 excellence) × 1.0000 critical = 0.8671

The core (88% of the total) is the weighted mean of the nine measured tiers. The last 12% is the excellence slice: it unlocks in proportion to the perfection conditions below (14 of 14 met here), then pays out at the excellence tier's own measured mean. Core 0.7953 + excellence 0.0718.

j_first_use 1.00j_sync_journey 1.00j_error_state 1.00j_empty_state 1.00console_clean 0.00v_responsive_375 1.00v_dates_readable 1.00t_scene_binding 1.00p_list_latency 1.00p_buckets_latency 1.00p_summary_latency 1.00p_page_interactive 1.00p_first_frame 1.00p_sync_wall 1.00
ComponentWeightEarnedPoints
Core (tiers A–D)60%——
A · Structure25% of core1.000.150
B · Behaviour30% of core1.000.180
C · Vendor contract25% of core1.000.150
D · Finesse20% of core1.000.120
Journey (J)15%1.0000.150
Visual (V)10%1.0000.100
Performance (P)5%1.0000.050
Hard blocks10%——
Total0.3000

Run notes and corrections

Cloud baseline — a single goose run session on the frozen VendorSync Pro spec, scored by the sb-6 scorer with the severity model (critical defects — crash, wrong money, data loss, dead primary flow — compound a multiplier on the composed score). Serial, hermetic, advertised vendor port. NOTE: this session was terminated early by a goose engine defect (compaction shipped reasoning content to Bedrock; fixed in 19b4ed6ef) — the score is a floor for this model, not a ceiling.

Screenshots

Recorded views of the built application. Captions identify the available captures; the number of images does not imply a number of repair rounds.

1Final render
23D buckets · WebGL
3After sync
4Mobile · 375px

Run details

Model
openai.gpt-5.6-luna