Cloud baseline

anthropic/claude-sonnet-5.5, single model via OpenRouter — 0.7698 on Forge 2.0

anthropic/claude-sonnet-5.5 · scorer forge-2.0

anthropic/claude-sonnet-5.5 · single agentOct 10, 2026
Overall
0.7698
Benchmark
Forge 2.0
Model build
54m 50s

Tier breakdown

Each bar is the tier's mean, from 0 to 1 (for E, its mean × the excellence gate). The percentage is the tier's share of the score. Mean × share, added over all tiers, is the tests earned figure below.

L · Lint and bundles1.001.76%
K · Platform currency1.002.2%
T · Event pipeline0.97773.52%
R · Reconcile0.95673.08%
S · Storage1.001.76%
B · Resolvers and permissions0.97222.64%
U · UI function0.97083.52%
V · Visual1.001.76%
A · Rovo0.85831.76%
E · Excellence0.72943%
R1 · Migration1.0012%
R2 · Dosing1.0013%
R3 · Time limits1.008%
R4 · Changing world0.785710%
R5 · Admin panel1.0010%
R6 · CI web trigger1.008%
R7 · Custom field0.98785%
R8 · Forge LLM1.004%
R9 · Boot budget1.005%

Scoring detail

How this score was built

Per-test scorer output, as posted by the app. Expand a section to inspect its tests.

The rules behind these numbers: Forge 2.0 requirements, weights and score caps.

The score, step by step

0.9635 tests earned × 1.0000 critical defects × 0.7990 reliability = 0.7698 final score

tests earned = 0.25 × 0.9420 v1 + 0.7280 v2 (out of 0.75)

v1 = 0.88 × 0.9710 tiers L–A + 0.12 × 0.9681 gate × 0.7534 excellence

Reliability: every tier in the breakdown above, except Excellence, is a group of tests. Each group of tests multiplies the score by 1 − 0.10 × the shortfall of its worst test: by 0.90 when that test fails completely, by 0.95 when it scores 0.5.

  • T Event pipeline: worst test t_reestimate_followed, score 0.8333 × 0.9833

    Already counted in this group: t_event_rows, t_multi_sprint_parse

  • R Reconcile: worst test r_heal_dropped, score 0.75 × 0.9750

    Already counted in this group: r_pagination

  • B Resolvers and permissions: worst test b_comment_adf_as_user, score 0.8333 × 0.9833
  • U UI function: worst test u_ledger_sort, score 0.8333 × 0.9833

    Already counted in this group: u_ledger_table

  • A Rovo: worst test a_action_result, score 0.6 × 0.9600

    Already counted in this group: a_action_permissions

  • R4 Changing world: worst test r4_sprint_close_final, score 0 × 0.9000

    Already counted in this group: r4_move_new_board

  • R7 Custom field: worst test r7_values_fresh, score 0.9756 × 0.9976

v1 is this run's score on the tests carried over from Forge 1.0: 25% of the total. v2 is the requirements R1–R9, each at its share of the score: the other 75%. Critical defects multiply the whole score, and so does each group of tests, by 1 − 0.10 × the shortfall of its worst test.

Excellence gate 0.9681: the average of the 18 conditions below, each counted at the value shown. Green is met in full; red is not, and lowers the gate. The 3 excellence tests are paid in proportion to the gate; together they are 3% of the total.

t_event_rows 0.9969r_heal_dropped 0.75u_widget_numbers 1.00u_ledger_table 0.8459b_comment_adf_as_user 0.8333k_consumer_shape 1.00k_current_apis 1.00k_dashboard_widget 1.00k_entity_declared 1.00k_llm_model_current 1.00k_manifest_semantics 1.00k_rovo_mcp 1.00k_rovo_skill 1.00k_runtime_risks 1.00k_widget_edit_bridge 1.00v_console_clean 1.00v_csp_clean 1.00l_lint_warnings 1.00

Earned score 0.7698 · No cap applies · Final score 0.7698

Under a cap the final score is min(earned, cap − 0.05 × (1 − earned)): always below the cap, and closer to it the more the run earned. With no cap it is the earned score. Avoiding a cap adds no points.

Few Jira requests in the first scheduled run and per relevant change in triggers and consumers, and surfaces that paint after one round trip — paid only in proportion to the excellence gate.

e reconcile economy

1.00

mean over 3 scoring sites: backfill 35 background Jira calls (by lineage, every phase) / optimum 20 (incl. 2 forced 429 repeat(s)) = 1.75x; 1.0 at <= 2x; 2 2.0 call(s) not counted (field-value writes, person-facing, web

e event economy

0.2603

mean over 3 scoring sites: 674 Jira calls (by lineage, every phase) / baseline 204 (203 delivered relevant change(s) x 1 read + 1 forced 429 repeat; a normalizer, not a reachable optimum) = 3.30x; 1.0 at <= 2.58x; 182 fi

e ui round trips

1.00

mean over 3 scoring sites: round trips before first paint: {'widget-view': 0, 'widget-edit': 1, 'sprint-action': 1}

Graded browser recording

Loading duration…

Watch the full graded browser recording at its original speed: the dashboard widget, the sprint action, the not-started sprint and the Forge LLM cases, as the scorer drove them on the run's own seed, the first of the three scoring sites. The test results and screenshots on this page provide the wider evidence. Pause or scrub to inspect a view.

The UI Kit admin panel is not in the recording: it is graded from its component tree.

Ready to play0:00 / —:—
Open recording file

Drag the timeline to seek. Keyboard: arrow keys move 5 seconds; Home and End jump to the start and end.

Full graded browser recording: dashboard widget (before configuration, edit and Save, light and dark at 380 px, a second instance at 1180 px), sprint action (light and dark, sort, issue link, select, comment post, double click, forbidden post, explain, close), the not-started sprint and the Forge LLM cases, where the app has each. The UI Kit admin panel is not in the recording.

Screenshots

Pictures from one of the three scoring sites, the one on the run's own seed. The dashboard widget and the sprint action are captured from the built app in a browser during grading. The admin panel picture is drawn by the benchmark from the app's component tree, not by Jira, and nothing is graded from that picture. A test keeps the score of its worst site (the excellence measurements keep their mean), so a test can fail on a site these pictures do not show. A detail that opens "worst of …" names the seed of the site the test scored worst on, which can be the pictured one.

1Dashboard widget · light
2Dashboard widget · dark
3Sprint action · light
4Sprint action · dark
5Admin panel (UI Kit) · light: drawn by the benchmark from the app's component tree, not by Jira.

Run details

Model
anthropic/claude-sonnet-5.5
Engine events
0
Repair rounds
0
Started
Oct 10, 2026, 05:57 PM
Finished
Oct 10, 2026, 07:11 PM

How the benchmark is scored