Cloud baseline

upstage/solar-mini4, single model via OpenRouter — 0.0083 on Gauntlet 7.2

upstage/solar-mini4 · scorer sb-7.2

upstage/solar-mini4 · single agentOct 3, 2026
Overall
0.0083
Excellent
no
Benchmark
Gauntlet 7.2
Model build
24m 3s
Prompt tok
11.0M
Gen tok
87.6k

Tier breakdown

A · structure & runtime0.42
B · behaviour0.00
C · vendor contract0.00
D · finesse0.08

Scoring detail

How this score was built

Per-check scorer output, as posted by the app. Expand a section to inspect its checks.

Earned credit before admission

(0.88 × 0.0158 core + 0.12 × 0.00 gate × 0.00 excellence) × 0.6000 critical = 0.0083

Critical defects compound a multiplier on the whole score (pre-severity 0.0139):

server_runs 0.15 → ×0.60

sync_completeness: no additional penalty (root:server_runs); data loss — silently missing payments.

b_money_rendered: no additional penalty (vacuous:sync_completeness); wrong money — wrong exponent/digits or a cross-currency sum.

b_buckets_dst: no additional penalty (vacuous:sync_completeness); wrong money — mis-bucketed days.

j_loads_data: no additional penalty (root:server_runs); dead primary flow — no data visible.

j_workflow_journey: no additional penalty (root:server_runs); dead primary flow — approval cannot complete through the UI.

x_conservation_residual: no additional penalty (vacuous:sync_completeness); wrong money — unexplained minor units created/destroyed after dupe/loss attribution.

x_no_lost_write: no additional penalty (vacuous:sync_completeness); wrong money — an acknowledged mutation absent from final state.

r_no_row_loss: no additional penalty (vacuous:sync_completeness); data loss — a committed row missing after any seeded kill.

r_no_dupe_effect: no additional penalty (vacuous:sync_completeness); wrong money — a ledger effect applied twice.

r_cache_truth: no additional penalty (vacuous:sync_completeness); data loss — 304-vs-cache mismatch served as fresh.

r_workflow_durability: no additional penalty (vacuous:server_runs); data loss — submitted/approved state reverting after SIGKILL.

The core (88% of the total) is the weighted mean of the thirteen measured tiers. The last 12% is the excellence slice: it unlocks in proportion to the perfection conditions below (0 of 16 met here), then pays out at the excellence tier's own measured mean. Core 0.0139 + excellence 0.0000.

j_first_use 0.00j_workflow_journey 0.00j_error_state 0.00j_empty_state 0.00console_clean not measuredv_responsive_375 0.00v_dates_readable 0.00t_scene_binding 0.00x_conservation_residual 0.00r_no_row_loss 0.00p_drag_frames 0.00p_idle_flatness 0.00p_stream_apply 0.00p_under_stream 0.00p_api_latency 0.00p_sync_wall 0.00

Earned score 0.008 · Admission ceiling 0.599 · Final score 0.008

Final score is the lower of earned credit and the admission ceiling. Passing admission adds no points.

Visible: not metMatching: not metGood: not metVisual excellence: not metBackend excellence: not met

Visible data-backed 3D: s_visible_surface (maximum 0.599)

Tower structure, currency mapping and scene truth: s_visible_surface, s_tower_geometry, s_currency_collar, t_layout_basis, t_scene_binding, t_height_pixels, t_vs7dbg_truth (maximum 0.699)

3D interaction and overview legibility: t_draw_budget, t_pick_buffer, t_pick_real_pass, t_click_semantics, t_camera_math, t_coast_identity, t_coast_reality, t_labels_culling, t_brush_link, t_stream_diff, q_overview_legibility, q_inspector_framing (maximum 0.799)

Event animation and backend recovery: m_committed_event_replay, x_l1_no_invented_states, x_l2_per_key_order, x_l3_monotonic_reads, x_l4_convergence, x_l5_group_atomicity, x_m1_amount_immutability, x_m2_pair_conservation, x_m3_terminal_conservation, x_m4_no_cross_currency, x_conservation_residual, x_no_lost_write, x_ooo_dup_forged, r_b3_sigkill_resync, r_b4_vendor_down_boot, r_b6_outbox_atomic, r_b7_partition, r_notifier_exactly_once, r_notification_multiset, r_no_row_loss, r_no_dupe_effect, r_cache_truth, r_workflow_durability (maximum 0.899)

Token rates

Measured by the engine itself, one record per completed model call: prefill rate is prompt tokens over time-to-first-token, generation rate is completion tokens over the decode window. Medians per node.

NodeCallsPrompt tokGen tokPrefill tok/sGen tok/s
gemini19725798.058.8
solar15010,993,49287,62023780.46297.1
fleet15110,994,46487,62523457.76260.9

Graded browser recording

No recording: the app never served a page.

server_runs: process survives 5s without binding; serves_page: GET / -> None

Run details

Model
upstage/solar-mini4
Engine events
0
Repair rounds
0
Started
Oct 3, 2026, 01:28 AM
Finished
Oct 3, 2026, 01:54 AM