Cloud baseline

z-ai/glm-5.3, single model via OpenRouter — 0.0292 on Forge 1.0

z-ai/glm-5.3 · scorer forge-1.0

z-ai/glm-5.3 · single agentOct 4, 2026
Overall
0.0292
Excellent
no
Benchmark
Forge 1.0
Model build
17m 18s

Tier breakdown

L · Lint and bundles0.28
K · Platform currency0.50
T · Event pipeline0.00
R · Reconcile0.00
S · Storage0.00
B · Resolvers and permissions0.00
U · UI function0.00
V · Visual0.00
A · Rovo0.25
E · Excellence0.00

Scoring detail

How this score was built

Per-check scorer output, as posted by the app. Expand a section to inspect its checks.

Earned credit before admission

(0.88 × 0.0922 core + 0.12 × 0.28 gate × 0.00 excellence) × 0.3600 critical = 0.0292

Critical defects compound a multiplier on the whole score (pre-severity 0.0812):

l_deployable 0.00 → ×0.60l_bundles_load 0.00 → ×0.60

t_no_double_count: no additional penalty (vacuous:precondition: >= 1 ledger row); wrong numbers — a change counted twice.

r_backfill_complete: no additional penalty (root:l_bundles_load); data loss — changes silently missing.

b_no_permission_leak: no additional penalty (vacuous:precondition: >= 1 person-facing change list returned or rendered); data leak — hidden issues shown to a person.

b_comment_exactly_once: no additional penalty (vacuous:precondition: >= 1 comment POST); duplicate side effect on a customer's Jira.

u_widget_loads: no additional penalty (root:l_bundles_load); dead primary flow — no data on the dashboard.

The core (88% of the total) is the weighted mean of the nine measured tiers. The last 12% is the excellence slice: it unlocks in proportion to the perfection conditions below (5 of 18 met here), then pays out at the excellence tier's own measured mean. Core 0.0811 + excellence 0.0000.

t_event_rows 0.00r_heal_dropped 0.00u_widget_numbers 0.00u_ledger_table 0.00b_comment_adf_as_user 0.00k_consumer_shape 1.00k_current_apis 0.00k_dashboard_widget 1.00k_entity_declared 1.00k_llm_model_current 0.00k_manifest_semantics 0.00k_rovo_mcp 1.00k_rovo_skill 1.00k_runtime_risks 0.00k_widget_edit_bridge 0.00v_console_clean 0.00v_csp_clean 0.00l_lint_warnings 0.00

Earned score 0.029 · Cap 0.499 · Final score 0.029

The final score is the lower of the earned score and the cap. Meeting a cap's conditions adds no points.

  • deployable: maximum 0.499l_deployablel_bundles_load
  • working ledger: maximum 0.699u_widget_loadst_event_rowsr_backfill_completes_storage_scopes_entity_index_used
  • current platform, complete surfaces: maximum 0.799k_widget_edit_bridgeu_widget_edit_configu_ledger_tablea_action_resultv_theme_tokensv_dark_mode
  • production robustness: maximum 0.799t_no_double_countt_out_of_ordert_retry_after_honouredr_heal_droppedr_rate_limitr_paginationb_no_permission_leakb_comment_exactly_onceb_realtime_payload_cleanu_llm_explainu_widget_livev_csp_cleanv_console_cleanu_ledger_sorts_index_orderb_hidden_counta_action_permissionsl_scopesk_llm_model_current

Graded browser recording

Loading duration…

Watch the full graded browser recording at its original speed: the payment field, structure inspection, committed updates and replay. The check results and screenshots on this page provide the wider evidence. Pause or scrub to inspect a view.

Each tower represents a payment: height encodes its amount, cap color its status, and the collar its currency. The selected detailed tower is expected to animate its collar after a committed backend update or replay. These are the task requirements; the check results show what this app actually achieved.

Ready to play0:00 / —:—
Open recording file

Drag the timeline to seek. Keyboard: arrow keys move 5 seconds; Home and End jump to the start and end.

Full graded browser recording: dashboard widget (no config, edit + Save, light and dark, 380 and 1180 px, a second instance), sprint action modal (sort, router, select, comment post through the 429 retry, double click, forbidden post, close) and the not-started sprint
Recording integrity

SHA-256: 9fc32c776b0202c15abc0a69c51e71b1a7a3fe7f513e067684a78dd0421d1a59

Screenshots

Captured from the built application during browser grading. Captions identify the recorded views and checks; these images do not represent repair rounds.

1Every captured surface
2Widget edit view
3Widget before configuration
4Sprint not started
5Sprint action · dark

Run details

Model
z-ai/glm-5.3
Engine events
0
Repair rounds
0
Started
Oct 4, 2026, 01:31 PM
Finished
Oct 4, 2026, 02:06 PM