Cloud baseline
anthropic/claude-opus-5.5, single model via OpenRouter — 0.8100 on Forge 2.0
anthropic/claude-opus-5.5 · scorer forge-2.0
- Overall
- 0.8100
- Benchmark
- Forge 2.0
- Model build
- 104m 57s
Tier breakdown
Each bar is the tier's mean, from 0 to 1 (for E, its mean × the excellence gate). The percentage is the tier's share of the score. Mean × share, added over all tiers, is the tests earned figure below.
Scoring detail
How this score was built
Per-test scorer output, as posted by the app. Expand a section to inspect its tests.
The rules behind these numbers: Forge 2.0 requirements, weights and score caps.
The score, step by step
0.9685 tests earned × 1.0000 critical defects × 0.8363 reliability = 0.8100 final score
tests earned = 0.25 × 0.9567 v1 + 0.7293 v2 (out of 0.75)
v1 = 0.88 × 0.9823 tiers L–A + 0.12 × 0.9886 gate × 0.7778 excellence
Reliability: every tier in the breakdown above, except Excellence, is a group of tests. Each group of tests multiplies the score by 1 − 0.10 × the shortfall of its worst test: by 0.90 when that test fails completely, by 0.95 when it scores 0.5.
- T Event pipeline: worst test
t_reestimate_followed, score 0.8333 × 0.9833Already counted in this group: t_event_rows, t_multi_sprint_parse
- B Resolvers and permissions: worst test
b_comment_adf_as_user, score 0.8182 × 0.9818 - U UI function: worst test
u_ledger_sort, score 0.8333 × 0.9833Already counted in this group: u_ledger_table, u_comment_flow
- A Rovo: worst test
a_action_result, score 0.8 × 0.9800Already counted in this group: a_action_permissions
- R4 Changing world: worst test
r4_sprint_close_final, score 0 × 0.9000Already counted in this group: r4_field_switch
- R7 Custom field: worst test
r7_writes_accepted, score 0.9877 × 0.9988
v1 is this run's score on the tests carried over from Forge 1.0: 25% of the total. v2 is the requirements R1–R9, each at its share of the score: the other 75%. Critical defects multiply the whole score, and so does each group of tests, by 1 − 0.10 × the shortfall of its worst test.
Excellence gate 0.9886: the average of the 18 conditions below, each counted at the value shown. Green is met in full; red is not, and lowers the gate. The 3 excellence tests are paid in proportion to the gate; together they are 3% of the total.
Earned score 0.8100 · No cap applies · Final score 0.8100
Under a cap the final score is min(earned, cap − 0.05 × (1 − earned)): always below the cap, and closer to it the more the run earned. With no cap it is the earned score. Avoiding a cap adds no points.
Forge's own linter reports no errors (warnings cost points), every manifest function bundles and loads, packages resolve at the kit's pins, the manifest rules the linter misses, and only the scopes the app's calls need.
l deployable
1.00lint: 0 errors, every stage reached
l bundles load
1.007/7 functions bundle, load and export their handler
l lint warnings
1.000 distinct lint warning(s): []
l real packages
1.007/7 bundles resolve @forge/* at the kit pins only
l manifest rules
1.006/6 manifest rules met:
l scopes
1.00scopes: missing [], extra []
Current modules and APIs: dashboards:widget with its edit API, the rovo:skill and rovo:mcp wiring, a Forge LLM model the platform lists as active, no /rest/api/3/search, no @forge/api storage, a nodejs22.x or nodejs24.x runtime, the consumer shape, the declared KVS entity, and deploy readiness: a manifest and permissions real Forge would accept and no code predicted to fail there.
k dashboard widget
1.00dashboards:widget 2/2: {'custom_ui_resource': True, 'edit_resource': True}
k widget edit bridge
1.004/4 board picks reached the dashboard through the widget edit API (updateConfig/onProductSave)
k rovo skill
1.00rovo:skill 7/7:
k current apis
1.00current APIs 5/5:
k consumer shape
1.00consumer shape 3/3
k entity declared
1.00entities with a sprint-partitioned, ranged index: ['scope-change', 'scope-ledger', 'scope-member', 'sprint-issue']
k v2 surfacesdiagnostic, no weight
1.00v2 surfaces declared: ['jira:adminPage', 'webtrigger', 'scope-ledger']
k rovo mcp
1.00rovo:mcp 4/4:
k llm model current
1.0011 LLM call(s); model ids list() does not return: []; llm module with claude: True
k manifest semantics
1.00deploy readiness, manifest + permissions: 17 rules checked, none failing here
k runtime risks
1.00deploy readiness, predicted runtime failures: 12 rules checked, none failing here
Issue updates flow trigger to queue to consumer: one ledger row per change under duplicate, reordered, dropped and concurrent deliveries, multi-sprint changelog values, re-estimates followed, Retry-After honoured, no user context in background work.
t trigger handoff
1.001004/1004 deliveries handed off correctly
t event rows
0.9915worst of 3 scoring sites (seed b2f7d3e899f76986): event rows: 175/178 exact (precision 1.00, recall 0.98)
t no double count
1.00every change holds at most one ledger row in every phase (duplicated deliveries included); 12 concurrent group(s) raced (same-event 8, same-issue 4)
t out of order
1.00permuted deliveries: 11/11 rows exact, 4/4 sprints with oracle numbers
t multi sprint parse
0.9915worst of 3 scoring sites (seed b2f7d3e899f76986): 117/118 multi-id Sprint changes recorded exactly
t reestimate followed
0.8333worst of 3 scoring sites (seed b2f7d3e899f76986): 5/6 re-estimated sprints show the oracle numbers
t retry after honoured
1.00waited 30.5 virtual s in-invocation
t no user in async
1.00background work is asApp only
The first scheduled run starts the backfill of every change since each active sprint started, removals included; later runs heal what the event stream missed and a run with nothing new writes nothing; they page through results in each endpoint's own style, never retry before Retry-After and finish inside the module timeout, queue continuations included.
r backfill complete
1.00134/134 historical changes v1 never recorded, recorded one virtual hour after the upgrade (the backfill is dosed; v1 rows are R1's, at 2 h)
r removals found
1.0013/13 removed-to-backlog changes found
r heal dropped
1.005/5 dropped changes healed exactly once; 0 other row(s) added
r pagination
1.0054/54 reads the site served in two or more pages walked to their end, each page once, continuing where the site left off (/rest/agile/1.0/board 15/15, /rest/agile/1.0/board/{id}/sprint 25/25, /rest/api/3/changelog/bulkfe
r rate limit
1.00run 1 (backfill) request 2: waited 2.5 virtual s in-invocation; run 1 (backfill) continuation page: waited 2.5 virtual s in-invocation; run 2 (hour-1) request 2: waited 2.5 virtual s in-invocation
r as app
1.00171/171 Jira calls of the scheduled runs and their queue continuations asApp (scheduled 171)
r completes in timeout
1.008/8 invocations of the scheduled runs and their queue continuations inside their module timeout (scheduled 8)
r idempotent rerun
1.000 ledger write(s) (scope-ledger) on a no-change run over 397 rows
Forge KVS with the storage scope, ledger reads through the declared entity index in change-time order, no KVS limit errors.
s storage scope
1.00KVS used, no 403 from the KVS proxy
s entity index used
1.00314/314 ledger reads are entity queries on a declared index
s index order
1.006/6 tables list changes in change-time order
s limits
1.00no KVS limit errors
Every invoked resolver exists and answers structured errors, nobody sees changes to issues they cannot browse (not on screen, not in an LLM prompt, not in a Realtime payload), the hidden-change count is right, and each click posts exactly one ADF comment as the viewer.
b invoke contract
1.0057/57 invokes returned a defined, structured result; a Jira 500 on sprintLedger's read: handled
b no permission leak
1.00no hidden key, summary or change id reached a person
b hidden count
1.0018/18 hidden counts equal the oracle
b comment adf as user
0.8182worst of 3 scoring sites (seed a077e1430dd5520c): 9/11 comments are ADF as the viewer naming key, sprint and creep; ['25709 status 201 via user', '25949 status 201 via user']
b comment exactly once
1.0012 click/double-click gesture(s), none posted more than one comment
b realtime payload clean
1.00186/186 realtime payloads hold sprint ids only
The dashboard widget's numbers and chart, its board choice saved through the host, live updates through Forge Realtime, the sprint action's ledger table, sorting, issue links, comment flow, the Forge LLM explanation, close and not-started states.
u widget loadsdiagnostic, no weight
1.006/8 configured widget views render a sprint with its numbers
u widget numbers
1.00widget numbers: 8.00/8 views exact (sprints by startDate, four §1 numbers each)
u widget chart
1.008/8 charts on one linear scale within 1px
u widget edit config
1.00edit/config 10/10:
u ledger table
0.9859worst of 3 scoring sites (seed b2f7d3e899f76986): ledger tables: 5.92/6 exact (rows, cells, default order)
u ledger sort
0.8333worst of 3 scoring sites (seed b2f7d3e899f76986): 15/18 at-toggle states ordered with aria-sort (6 after another sort was active)
u issue router
1.006/6 issue keys open /browse/<KEY> through the Forge router
u comment flow
0.9722worst of 3 scoring sites (seed a077e1430dd5520c): 35/36 comment-flow steps right (select, one comment and a success flag per click and double click, forbidden error flag, modal works); failed: ['post_one_comment@1077']
u modal close
1.006/6 close clicks reached bridge close
u not started
1.00future sprint shows not-started only
u widget live
1.00moved sprints showing the new numbers live 1/1; 0 idle invoke(s)
u llm explain
1.00explain 23/23 over 5 scripted answers:
Atlassian design tokens with 4.5:1 contrast in light and dark, a painted surface, no Content Security Policy violations or console errors, and the widget readable at 380 px wide.
v theme tokens
1.0031/31 surfaces themed with --ds-text* at >= 4.5:1
v dark mode
1.0031/31 surfaces paint a --ds-surface* background in their mode
v csp clean
1.000 CSP violation(s), 0 failed asset request(s)
v console clean
1.000 console/page error(s) on nominal scenarios: []
v widget sizes
1.006/6 widget sizes without horizontal overflow, every sprint visible, no number clipped
The get-sprint-scope action returns exact numbers and the visible changes for the invoking person, returns errors instead of throwing, and the skill's SKILL.md tells the agent how to use it.
a action result
0.804/5 active sprints answered exactly
a action errors
1.002/2 bad sprintId inputs return {error} without throwing
a action permissions
0.8333worst of 3 scoring sites (seed b2f7d3e899f76986): 10/12 per-person answers list exactly what that person may browse
a skill instructions
1.00SKILL.md body 5/5: result fields named ['sprintName', 'committed', 'added', 'removed', 'creepPercent', 'hiddenChanges', 'changes']
Few Jira requests in the first scheduled run and per relevant change in triggers and consumers, and surfaces that paint after one round trip — paid only in proportion to the excellence gate.
e reconcile economy
1.00mean over 3 scoring sites: backfill 30 background Jira calls (by lineage, every phase) / optimum 20 (incl. 2 forced 429 repeat(s)) = 1.50x; 1.0 at <= 2x; 2 2.0 call(s) not counted (field-value writes, person-facing, web
e event economy
0.3333mean over 3 scoring sites: event rows inexact — economy is not credited on unfinished work
e ui round trips
1.00mean over 3 scoring sites: round trips before first paint: {'widget-view': 0, 'widget-edit': 1, 'sprint-action': 1}
Every v1 row moves to the new scope-ledger entity exactly once, with its change id and time, while events keep flowing and across invocation limits, within 2 virtual hours of the upgrade, with its progress in the admin panel; v1's entity stays intact.
r1 v1 rows migrated
1.00h2: 83 of 83 v1 rows in scope-ledger with their changeId and time
r1 v1 rows intact
1.00all 83 v1 rows intact in scope-change after 83 were migrated
r1 events during migration
1.0058 of 58 changes made while the migration ran are in scope-ledger at the end
r1 progress visible
1.006 of 6 panel reads show the migration right (total 83; complete after the h2 read)
On a site of about 1,000 issues, background work stays inside its share of the 2,400-point hour, no person-facing request meets a quota 429, and Retry-After and RateLimit-Reason (quota, burst, per-issue write) decide every retry; the backfill still completes.
r2 background share
1.009 of 9 hours kept background within 70 % (1680 points; the contract default; the probe recorded no change); busiest hour 223; dosing optimum for the whole run 254
r2 person never quota refused
1.000 of 341 person-facing requests refused for quota
r2 rate limit reaction
1.0014 of 14 429s answered by their reason (quota: all background paused; burst: the endpoint; per-issue: the issue)
r2 backfill complete
1.00134 of 134 pre-upgrade changes v1 never recorded are in scope-ledger, right
Work longer than one invocation continues in the next, no function waits past its limit (a long Retry-After is handed back, not waited out), and a killed invocation leaves no duplicate.
r3 never killed
1.007 of 7 functions never exceeded their limit
r3 long retry after deferred
1.001 of 1 invocations handed a Retry-After longer than their limit stopped and resumed later (results: retry)
r3 no duplicate rows
1.00every (changeId, sprintId) appears once at each of 7 checkpoints; 12 concurrent group(s) raced (same-event 8, same-issue 4)
Sprints close, issues move to another board's sprint, a board's estimation field changes, issues are deleted and a person loses browse permission while the app runs; the ledger and every view come out right.
r4 sprint close final
0.00worst of 3 scoring sites (seed b2f7d3e899f76986): 0 of 1 closed sprints kept exactly their final ledger; sprint 987: 3 rows gone (e.g. ('1432222', '987'), ('1432439', '987') (+1))
r4 move new board
1.0018 of 18 cross-board moves recorded with the new board's estimate
r4 field switch
0.9808worst of 3 scoring sites (seed b2f7d3e899f76986): 29 of 30 rows after the switch use the new field; 22 of 22 earlier rows keep their estimate; wrong after: ('1430325', '1245')
r4 deleted history
1.004 of 4 rows of deleted issues kept as history with deleted: true
r4 permission revoked
1.004 answer(s) after a browse revoke showed ledger data and none of the newly hidden issues
A jira:adminPage built in UI Kit with the contract's controls, and every admin resolver checking Jira's ADMINISTER permission on the server, never trusting the payload.
r5 panel labels
1.008 of 8 §2.6 controls found by label; new secret shown once: yes; then masked ••••<last4>: yes
r5 admin actions apply
1.002 of 2 admin-page invokes of the admin's save and rotate applied
r5 nonadmin refused
1.004 non-admin/forged admin actions changed nothing — judged on the attempted effect landed in storage
Signed deployment events: a missing or wrong signature, a tampered body or a stale timestamp is refused with no effect, a replay applies once, a valid event records the deployment, and the secret is never disclosed.
r6 unsigned no effect
1.005 invalid events (bad-signature, stale, stale-future, tampered, unsigned) had no effect
r6 replay once
1.001 replay(s) made no Jira write and doubled no deployment; bookkeeping writes 0
r6 status codes
1.008 of 8 cases answered with their status
r6 valid applies
1.002 of 2 validly signed events accepted (202) and recorded as `Deployed to <env>` on every ledger row of the issues they name
r6 secret never disclosed
1.00the CI secret (panel) never appeared outside its one-time display
The read-only scope-status field holds committed or added +points for an issue in an active sprint, removed for one that has left an active sprint since that sprint started, and nothing for any other issue, written in bulk through Jira's app field-value API and fresh within the virtual hour.
r7 values fresh
1.00mean over 7 checkpoints; worst h1: 361 of 361 issues right
r7 writes accepted
0.9877worst of 3 scoring sites (seed b2f7d3e899f76986): 160 of 162 bulk field writes accepted; refused: 404; 9 rate-limited (429, R2's)
Tool calls validated against the viewer's sprint scope, back-off on a 429 without Retry-After, answers with no finish reason refused, a 10-minute cache, and the admin's kill switch and daily token budget enforced on the server.
r8 tool call scope
1.001 of 1 manipulated tool calls validated against the viewer's sprint scope before acting
r8 failure handling
1.002 of 2 LLM failures handled
r8 cost controls
1.003 of 3 cache/kill-switch/budget cases made no model call
The widget and the sprint action each make at most one invoke and load at most 150 KB before their first data paint, with no external origin; the admin page renders after one invoke.
r9 widget boot
1.00widget-view: 1 invoke(s) before first paint (<= 1), 7365 bytes of JS+CSS (<= 153600), 0 external request(s)
r9 sprint boot
1.00sprint-modal: 1 invoke(s) before first paint (<= 1), 11668 bytes of JS+CSS (<= 153600), 0 external request(s)
r9 admin boot
1.00admin-page: 1 invoke(s) before first paint (<= 1)
Graded browser recording
Loading duration…Watch the full graded browser recording at its original speed: the dashboard widget, the sprint action, the not-started sprint and the Forge LLM cases, as the scorer drove them on the run's own seed, the first of the three scoring sites. The test results and screenshots on this page provide the wider evidence. Pause or scrub to inspect a view.
The UI Kit admin panel is not in the recording: it is graded from its component tree.
Drag the timeline to seek. Keyboard: arrow keys move 5 seconds; Home and End jump to the start and end.
SHA-256: 042f6cf2d4a9dd3df1d32488a14740a4913ac9b8f398558fbee65a3888af54df
Screenshots
Pictures from one of the three scoring sites, the one on the run's own seed. The dashboard widget and the sprint action are captured from the built app in a browser during grading. The admin panel picture is drawn by the benchmark from the app's component tree, not by Jira, and nothing is graded from that picture. A test keeps the score of its worst site (the excellence measurements keep their mean), so a test can fail on a site these pictures do not show. A detail that opens "worst of …" names the seed of the site the test scored worst on, which can be the pictured one.
Run details
- Model
- anthropic/claude-opus-5.5
- Engine events
- 0
- Repair rounds
- 0
- Started
- Oct 11, 2026, 01:41 AM
- Finished
- Oct 11, 2026, 03:44 AM