Cloud baseline
anthropic/claude-sonnet-5.5, single model via OpenRouter — 0.7698 on Forge 2.0
anthropic/claude-sonnet-5.5 · scorer forge-2.0
- Overall
- 0.7698
- Benchmark
- Forge 2.0
- Model build
- 54m 50s
Tier breakdown
Each bar is the tier's mean, from 0 to 1 (for E, its mean × the excellence gate). The percentage is the tier's share of the score. Mean × share, added over all tiers, is the tests earned figure below.
Scoring detail
How this score was built
Per-test scorer output, as posted by the app. Expand a section to inspect its tests.
The rules behind these numbers: Forge 2.0 requirements, weights and score caps.
The score, step by step
0.9635 tests earned × 1.0000 critical defects × 0.7990 reliability = 0.7698 final score
tests earned = 0.25 × 0.9420 v1 + 0.7280 v2 (out of 0.75)
v1 = 0.88 × 0.9710 tiers L–A + 0.12 × 0.9681 gate × 0.7534 excellence
Reliability: every tier in the breakdown above, except Excellence, is a group of tests. Each group of tests multiplies the score by 1 − 0.10 × the shortfall of its worst test: by 0.90 when that test fails completely, by 0.95 when it scores 0.5.
- T Event pipeline: worst test
t_reestimate_followed, score 0.8333 × 0.9833Already counted in this group: t_event_rows, t_multi_sprint_parse
- R Reconcile: worst test
r_heal_dropped, score 0.75 × 0.9750Already counted in this group: r_pagination
- B Resolvers and permissions: worst test
b_comment_adf_as_user, score 0.8333 × 0.9833 - U UI function: worst test
u_ledger_sort, score 0.8333 × 0.9833Already counted in this group: u_ledger_table
- A Rovo: worst test
a_action_result, score 0.6 × 0.9600Already counted in this group: a_action_permissions
- R4 Changing world: worst test
r4_sprint_close_final, score 0 × 0.9000Already counted in this group: r4_move_new_board
- R7 Custom field: worst test
r7_values_fresh, score 0.9756 × 0.9976
v1 is this run's score on the tests carried over from Forge 1.0: 25% of the total. v2 is the requirements R1–R9, each at its share of the score: the other 75%. Critical defects multiply the whole score, and so does each group of tests, by 1 − 0.10 × the shortfall of its worst test.
Excellence gate 0.9681: the average of the 18 conditions below, each counted at the value shown. Green is met in full; red is not, and lowers the gate. The 3 excellence tests are paid in proportion to the gate; together they are 3% of the total.
Earned score 0.7698 · No cap applies · Final score 0.7698
Under a cap the final score is min(earned, cap − 0.05 × (1 − earned)): always below the cap, and closer to it the more the run earned. With no cap it is the earned score. Avoiding a cap adds no points.
Forge's own linter reports no errors (warnings cost points), every manifest function bundles and loads, packages resolve at the kit's pins, the manifest rules the linter misses, and only the scopes the app's calls need.
l deployable
1.00lint: 0 errors, every stage reached
l bundles load
1.007/7 functions bundle, load and export their handler
l lint warnings
1.000 distinct lint warning(s): []
l real packages
1.007/7 bundles resolve @forge/* at the kit pins only
l manifest rules
1.006/6 manifest rules met:
l scopes
1.00scopes: missing [], extra []
Current modules and APIs: dashboards:widget with its edit API, the rovo:skill and rovo:mcp wiring, a Forge LLM model the platform lists as active, no /rest/api/3/search, no @forge/api storage, a nodejs22.x or nodejs24.x runtime, the consumer shape, the declared KVS entity, and deploy readiness: a manifest and permissions real Forge would accept and no code predicted to fail there.
k dashboard widget
1.00dashboards:widget 2/2: {'custom_ui_resource': True, 'edit_resource': True}
k widget edit bridge
1.004/4 board picks reached the dashboard through the widget edit API (updateConfig/onProductSave)
k rovo skill
1.00rovo:skill 7/7:
k current apis
1.00current APIs 5/5:
k consumer shape
1.00consumer shape 3/3
k entity declared
1.00entities with a sprint-partitioned, ranged index: ['scope-change', 'scope-ledger', 'sprint-issue']
k v2 surfacesdiagnostic, no weight
1.00v2 surfaces declared: ['jira:adminPage', 'webtrigger', 'scope-ledger']
k rovo mcp
1.00rovo:mcp 4/4:
k llm model current
1.0011 LLM call(s); model ids list() does not return: []; llm module with claude: True
k manifest semantics
1.00deploy readiness, manifest + permissions: 17 rules checked, none failing here
k runtime risks
1.00deploy readiness, predicted runtime failures: 12 rules checked, none failing here
Issue updates flow trigger to queue to consumer: one ledger row per change under duplicate, reordered, dropped and concurrent deliveries, multi-sprint changelog values, re-estimates followed, Retry-After honoured, no user context in background work.
t trigger handoff
1.001008/1008 deliveries handed off correctly
t event rows
0.9969worst of 3 scoring sites (seed 1d8aa6ad7a780786): event rows: 159/160 exact (precision 1.00, recall 0.99)
t no double count
1.00every change holds at most one ledger row in every phase (duplicated deliveries included); 16 concurrent group(s) raced (same-event 12, same-issue 4)
t out of order
1.00permuted deliveries: 10/10 rows exact, 4/4 sprints with oracle numbers
t multi sprint parse
0.9913worst of 3 scoring sites (seed f0dbf1dfb64b8c3b): 114/115 multi-id Sprint changes recorded exactly
t reestimate followed
0.8333worst of 3 scoring sites (seed 1d8aa6ad7a780786): 5/6 re-estimated sprints show the oracle numbers
t retry after honoured
1.00waited 31.0 virtual s in-invocation
t no user in async
1.00background work is asApp only
The first scheduled run starts the backfill of every change since each active sprint started, removals included; later runs heal what the event stream missed and a run with nothing new writes nothing; they page through results in each endpoint's own style, never retry before Retry-After and finish inside the module timeout, queue continuations included.
r backfill complete
1.00101/101 historical changes v1 never recorded, recorded one virtual hour after the upgrade (the backfill is dosed; v1 rows are R1's, at 2 h)
r removals found
1.0020/20 removed-to-backlog changes found
r heal dropped
0.75worst of 3 scoring sites (seed 1d8aa6ad7a780786): 3/4 dropped changes healed exactly once; 0 other row(s) added
r pagination
0.9038worst of 3 scoring sites (seed 1d8aa6ad7a780786): 47/52 reads the site served in two or more pages walked to their end, each page once, continuing where the site left off (/rest/agile/1.0/board 14/15, /rest/agile/1.0/boa
r rate limit
1.00run 1 (backfill) request 2: waited 3.0 virtual s in-invocation; run 1 (backfill) continuation page: waited 3.0 virtual s in-invocation; run 2 (hour-1) request 2: waited 3.0 virtual s in-invocation
r as app
1.00194/194 Jira calls of the scheduled runs and their queue continuations asApp (scheduled 194)
r completes in timeout
1.0012/12 invocations of the scheduled runs and their queue continuations inside their module timeout (consumer 4, scheduled 8)
r idempotent rerun
1.000 ledger write(s) (scope-ledger) on a no-change run over 368 rows
Forge KVS with the storage scope, ledger reads through the declared entity index in change-time order, no KVS limit errors.
s storage scope
1.00KVS used, no 403 from the KVS proxy
s entity index used
1.00111/111 ledger reads are entity queries on a declared index
s index order
1.006/6 tables list changes in change-time order
s limits
1.00no KVS limit errors
Every invoked resolver exists and answers structured errors, nobody sees changes to issues they cannot browse (not on screen, not in an LLM prompt, not in a Realtime payload), the hidden-change count is right, and each click posts exactly one ADF comment as the viewer.
b invoke contract
1.0057/57 invokes returned a defined, structured result; a Jira 500 on sprintLedger's read: handled
b no permission leak
1.00no hidden key, summary or change id reached a person
b hidden count
1.0018/18 hidden counts equal the oracle
b comment adf as user
0.8333worst of 3 scoring sites (seed 1d8aa6ad7a780786): 10/12 comments are ADF as the viewer naming key, sprint and creep; ['40506 status 201 via user', '39824 status 201 via user']
b comment exactly once
1.0012 click/double-click gesture(s), none posted more than one comment
b realtime payload clean
1.00192/192 realtime payloads hold sprint ids only
The dashboard widget's numbers and chart, its board choice saved through the host, live updates through Forge Realtime, the sprint action's ledger table, sorting, issue links, comment flow, the Forge LLM explanation, close and not-started states.
u widget loadsdiagnostic, no weight
1.006/8 configured widget views render a sprint with its numbers
u widget numbers
1.00widget numbers: 8.00/8 views exact (sprints by startDate, four §1 numbers each)
u widget chart
1.008/8 charts on one linear scale within 1px
u widget edit config
1.00edit/config 10/10:
u ledger table
0.8459worst of 3 scoring sites (seed f0dbf1dfb64b8c3b): ledger tables: 5.08/6 exact (rows, cells, default order)
u ledger sort
0.8333worst of 3 scoring sites (seed 1d8aa6ad7a780786): 10/12 at-toggle states ordered with aria-sort
u issue router
1.006/6 issue keys open /browse/<KEY> through the Forge router
u comment flow
1.0032/32 comment-flow steps right (select, one comment and a success flag per click and double click, forbidden error flag, modal works)
u modal close
1.006/6 close clicks reached bridge close
u not started
1.00future sprint shows not-started only
u widget live
1.00moved sprints showing the new numbers live 1/1; 0 idle invoke(s)
u llm explain
1.00explain 23/23 over 5 scripted answers:
Atlassian design tokens with 4.5:1 contrast in light and dark, a painted surface, no Content Security Policy violations or console errors, and the widget readable at 380 px wide.
v theme tokens
1.0031/31 surfaces themed with --ds-text* at >= 4.5:1
v dark mode
1.0031/31 surfaces paint a --ds-surface* background in their mode
v csp clean
1.000 CSP violation(s), 0 failed asset request(s)
v console clean
1.000 console/page error(s) on nominal scenarios: []
v widget sizes
1.006/6 widget sizes without horizontal overflow, every sprint visible, no number clipped
The get-sprint-scope action returns exact numbers and the visible changes for the invoking person, returns errors instead of throwing, and the skill's SKILL.md tells the agent how to use it.
a action result
0.603/5 active sprints answered exactly
a action errors
1.002/2 bad sprintId inputs return {error} without throwing
a action permissions
0.8333worst of 3 scoring sites (seed 1d8aa6ad7a780786): 10/12 per-person answers list exactly what that person may browse
a skill instructions
1.00SKILL.md body 5/5: result fields named ['sprintName', 'committed', 'added', 'removed', 'creepPercent', 'hiddenChanges', 'changes']
Few Jira requests in the first scheduled run and per relevant change in triggers and consumers, and surfaces that paint after one round trip — paid only in proportion to the excellence gate.
e reconcile economy
1.00mean over 3 scoring sites: backfill 35 background Jira calls (by lineage, every phase) / optimum 20 (incl. 2 forced 429 repeat(s)) = 1.75x; 1.0 at <= 2x; 2 2.0 call(s) not counted (field-value writes, person-facing, web
e event economy
0.2603mean over 3 scoring sites: 674 Jira calls (by lineage, every phase) / baseline 204 (203 delivered relevant change(s) x 1 read + 1 forced 429 repeat; a normalizer, not a reachable optimum) = 3.30x; 1.0 at <= 2.58x; 182 fi
e ui round trips
1.00mean over 3 scoring sites: round trips before first paint: {'widget-view': 0, 'widget-edit': 1, 'sprint-action': 1}
Every v1 row moves to the new scope-ledger entity exactly once, with its change id and time, while events keep flowing and across invocation limits, within 2 virtual hours of the upgrade, with its progress in the admin panel; v1's entity stays intact.
r1 v1 rows migrated
1.00h2: 85 of 85 v1 rows in scope-ledger with their changeId and time
r1 v1 rows intact
1.00all 85 v1 rows intact in scope-change after 85 were migrated
r1 events during migration
1.0057 of 57 changes made while the migration ran are in scope-ledger at the end
r1 progress visible
1.006 of 6 panel reads show the migration right (total 85; complete after the h2 read)
On a site of about 1,000 issues, background work stays inside its share of the 2,400-point hour, no person-facing request meets a quota 429, and Retry-After and RateLimit-Reason (quota, burst, per-issue write) decide every retry; the backfill still completes.
r2 background share
1.009 of 9 hours kept background within 70 % (1680 points; the contract default; the probe recorded no change); busiest hour 350; dosing optimum for the whole run 255
r2 person never quota refused
1.000 of 266 person-facing requests refused for quota
r2 rate limit reaction
1.0016 of 16 429s answered by their reason (quota: all background paused; burst: the endpoint; per-issue: the issue)
r2 backfill complete
1.00101 of 101 pre-upgrade changes v1 never recorded are in scope-ledger, right
Work longer than one invocation continues in the next, no function waits past its limit (a long Retry-After is handed back, not waited out), and a killed invocation leaves no duplicate.
r3 never killed
1.007 of 7 functions never exceeded their limit
r3 long retry after deferred
1.001 of 1 invocations handed a Retry-After longer than their limit stopped and resumed later (results: retry)
r3 no duplicate rows
1.00every (changeId, sprintId) appears once at each of 7 checkpoints; 16 concurrent group(s) raced (same-event 12, same-issue 4)
Sprints close, issues move to another board's sprint, a board's estimation field changes, issues are deleted and a person loses browse permission while the app runs; the ledger and every view come out right.
r4 sprint close final
0.00worst of 3 scoring sites (seed 1d8aa6ad7a780786): 0 of 1 closed sprints kept exactly their final ledger; sprint 923: 2 rows gone (e.g. ('1227844', '923'), ('1229292', '923'))
r4 move new board
0.9286worst of 3 scoring sites (seed f0dbf1dfb64b8c3b): 13 of 14 cross-board moves recorded with the new board's estimate; BILL-410 (1433032): removed row wrong/missing, added row ok (estimate 3, want 3)
r4 field switch
1.0032 of 32 rows after the switch use the new field; 21 of 21 earlier rows keep their estimate
r4 deleted history
1.003 of 3 rows of deleted issues kept as history with deleted: true
r4 permission revoked
1.005 answer(s) after a browse revoke showed ledger data and none of the newly hidden issues
A jira:adminPage built in UI Kit with the contract's controls, and every admin resolver checking Jira's ADMINISTER permission on the server, never trusting the payload.
r5 panel labels
1.008 of 8 §2.6 controls found by label; new secret shown once: yes; then masked ••••<last4>: yes
r5 admin actions apply
1.002 of 2 admin-page invokes of the admin's save and rotate applied
r5 nonadmin refused
1.004 non-admin/forged admin actions changed nothing — judged on the attempted effect landed in storage
Signed deployment events: a missing or wrong signature, a tampered body or a stale timestamp is refused with no effect, a replay applies once, a valid event records the deployment, and the secret is never disclosed.
r6 unsigned no effect
1.005 invalid events (bad-signature, stale, stale-future, tampered, unsigned) had no effect
r6 replay once
1.001 replay(s) made no Jira write and doubled no deployment; bookkeeping writes 0
r6 status codes
1.008 of 8 cases answered with their status
r6 valid applies
1.002 of 2 validly signed events accepted (202) and recorded as `Deployed to <env>` on every ledger row of the issues they name
r6 secret never disclosed
1.00the CI secret (panel) never appeared outside its one-time display
The read-only scope-status field holds committed or added +points for an issue in an active sprint, removed for one that has left an active sprint since that sprint started, and nothing for any other issue, written in bulk through Jira's app field-value API and fresh within the virtual hour.
r7 values fresh
0.9756worst of 3 scoring sites (seed 9f606c78c75c2ce9): mean over 7 checkpoints; worst h4: 304 of 318 issues right; SEC-264 'added +0' (want added +0.5); SEC-294 'added +0' (want added +13)
r7 writes accepted
1.00177 of 177 bulk field writes accepted; 11 rate-limited (429, R2's)
Tool calls validated against the viewer's sprint scope, back-off on a 429 without Retry-After, answers with no finish reason refused, a 10-minute cache, and the admin's kill switch and daily token budget enforced on the server.
r8 tool call scope
1.001 of 1 manipulated tool calls validated against the viewer's sprint scope before acting
r8 failure handling
1.002 of 2 LLM failures handled
r8 cost controls
1.003 of 3 cache/kill-switch/budget cases made no model call
The widget and the sprint action each make at most one invoke and load at most 150 KB before their first data paint, with no external origin; the admin page renders after one invoke.
r9 widget boot
1.00widget-view: 1 invoke(s) before first paint (<= 1), 6525 bytes of JS+CSS (<= 153600), 0 external request(s)
r9 sprint boot
1.00sprint-modal: 1 invoke(s) before first paint (<= 1), 10395 bytes of JS+CSS (<= 153600), 0 external request(s)
r9 admin boot
1.00admin-page: 1 invoke(s) before first paint (<= 1)
Graded browser recording
Loading duration…Watch the full graded browser recording at its original speed: the dashboard widget, the sprint action, the not-started sprint and the Forge LLM cases, as the scorer drove them on the run's own seed, the first of the three scoring sites. The test results and screenshots on this page provide the wider evidence. Pause or scrub to inspect a view.
The UI Kit admin panel is not in the recording: it is graded from its component tree.
Drag the timeline to seek. Keyboard: arrow keys move 5 seconds; Home and End jump to the start and end.
SHA-256: cae2387b626074fc14b5d5947cf266ffc7fb676964156a9adf4876c444bf8287
Screenshots
Pictures from one of the three scoring sites, the one on the run's own seed. The dashboard widget and the sprint action are captured from the built app in a browser during grading. The admin panel picture is drawn by the benchmark from the app's component tree, not by Jira, and nothing is graded from that picture. A test keeps the score of its worst site (the excellence measurements keep their mean), so a test can fail on a site these pictures do not show. A detail that opens "worst of …" names the seed of the site the test scored worst on, which can be the pictured one.
Run details
- Model
- anthropic/claude-sonnet-5.5
- Engine events
- 0
- Repair rounds
- 0
- Started
- Oct 10, 2026, 05:57 PM
- Finished
- Oct 10, 2026, 07:11 PM