Cloud baseline
openai/gpt-6.1-sol, single model via OpenRouter — 0.8111 on Forge 2.0
openai/gpt-6.1-sol · scorer forge-2.0
- Overall
- 0.8111
- Benchmark
- Forge 2.0
- Model build
- 32m 2s
Tier breakdown
Each bar is the tier's mean, from 0 to 1 (for E, its mean × the excellence gate). The percentage is the tier's share of the score. Mean × share, added over all tiers, is the tests earned figure below.
Scoring detail
How this score was built
Per-test scorer output, as posted by the app. Expand a section to inspect its tests.
The rules behind these numbers: Forge 2.0 requirements, weights and score caps.
The score, step by step
0.9845 tests earned × 1.0000 critical defects × 0.8238 reliability = 0.8111 final score
tests earned = 0.25 × 0.9556 v1 + 0.7456 v2 (out of 0.75)
v1 = 0.88 × 0.9676 tiers L–A + 0.12 × 0.9971 gate × 0.8702 excellence
Reliability: every tier in the breakdown above, except Excellence, is a group of tests. Each group of tests multiplies the score by 1 − 0.10 × the shortfall of its worst test: by 0.90 when that test fails completely, by 0.95 when it scores 0.5.
- R Reconcile: worst test
r_rate_limit, score 0.3333 × 0.9333 - U UI function: worst test
u_ledger_table, score 0.9475 × 0.9948 - A Rovo: worst test
a_action_result, score 0 × 0.9000 - R5 Admin panel: worst test
r5_panel_labels, score 0.9 × 0.9900 - R7 Custom field: worst test
r7_values_fresh, score 0.959 × 0.9959
v1 is this run's score on the tests carried over from Forge 1.0: 25% of the total. v2 is the requirements R1–R9, each at its share of the score: the other 75%. Critical defects multiply the whole score, and so does each group of tests, by 1 − 0.10 × the shortfall of its worst test.
Excellence gate 0.9971: the average of the 18 conditions below, each counted at the value shown. Green is met in full; red is not, and lowers the gate. The 3 excellence tests are paid in proportion to the gate; together they are 3% of the total.
Earned score 0.8111 · No cap applies · Final score 0.8111
Under a cap the final score is min(earned, cap − 0.05 × (1 − earned)): always below the cap, and closer to it the more the run earned. With no cap it is the earned score. Avoiding a cap adds no points.
Forge's own linter reports no errors (warnings cost points), every manifest function bundles and loads, packages resolve at the kit's pins, the manifest rules the linter misses, and only the scopes the app's calls need.
l deployable
1.00lint: 0 errors, every stage reached
l bundles load
1.007/7 functions bundle, load and export their handler
l lint warnings
1.000 distinct lint warning(s): []
l real packages
1.007/7 bundles resolve @forge/* at the kit pins only
l manifest rules
1.006/6 manifest rules met:
l scopes
1.00scopes: missing [], extra []
Current modules and APIs: dashboards:widget with its edit API, the rovo:skill and rovo:mcp wiring, a Forge LLM model the platform lists as active, no /rest/api/3/search, no @forge/api storage, a nodejs22.x or nodejs24.x runtime, the consumer shape, the declared KVS entity, and deploy readiness: a manifest and permissions real Forge would accept and no code predicted to fail there.
k dashboard widget
1.00dashboards:widget 2/2: {'custom_ui_resource': True, 'edit_resource': True}
k widget edit bridge
1.004/4 board picks reached the dashboard through the widget edit API (updateConfig/onProductSave)
k rovo skill
1.00rovo:skill 7/7:
k current apis
1.00current APIs 5/5:
k consumer shape
1.00consumer shape 3/3
k entity declared
1.00entities with a sprint-partitioned, ranged index: ['scope-change', 'scope-ledger', 'sprint-issue']
k v2 surfacesdiagnostic, no weight
1.00v2 surfaces declared: ['jira:adminPage', 'webtrigger', 'scope-ledger']
k rovo mcp
1.00rovo:mcp 4/4:
k llm model current
1.0011 LLM call(s); model ids list() does not return: []; llm module with claude: True
k manifest semantics
1.00deploy readiness, manifest + permissions: 17 rules checked, none failing here
k runtime risks
1.00deploy readiness, predicted runtime failures: 12 rules checked, none failing here
Issue updates flow trigger to queue to consumer: one ledger row per change under duplicate, reordered, dropped and concurrent deliveries, multi-sprint changelog values, re-estimates followed, Retry-After honoured, no user context in background work.
t trigger handoff
1.001046/1046 deliveries handed off correctly
t event rows
1.00event rows: 193/193 exact (precision 1.00, recall 1.00)
t no double count
1.00every change holds at most one ledger row in every phase (duplicated deliveries included); 11 concurrent group(s) raced (same-event 8, same-issue 3)
t out of order
1.00permuted deliveries: 10/10 rows exact, 4/4 sprints with oracle numbers
t multi sprint parse
1.00123/123 multi-id Sprint changes recorded exactly
t reestimate followed
1.006/6 re-estimated sprints show the oracle numbers
t retry after honoured
1.00InvocationError retryAfter 30s >= 30s
t no user in async
1.00background work is asApp only
The first scheduled run starts the backfill of every change since each active sprint started, removals included; later runs heal what the event stream missed and a run with nothing new writes nothing; they page through results in each endpoint's own style, never retry before Retry-After and finish inside the module timeout, queue continuations included.
r backfill complete
1.00114/114 historical changes v1 never recorded, recorded one virtual hour after the upgrade (the backfill is dosed; v1 rows are R1's, at 2 h)
r removals found
1.0023/23 removed-to-backlog changes found
r heal dropped
1.007/7 dropped changes healed exactly once; 0 other row(s) added
r pagination
1.0055/55 reads the site served in two or more pages walked to their end, each page once, continuing where the site left off (/rest/agile/1.0/board 16/16, /rest/agile/1.0/board/{id}/sprint 23/23, /rest/api/3/changelog/bulkfe
r rate limit
0.3333run 1 (backfill) request 2: the 429 was never retried; run 1 (backfill) continuation page: InvocationError retryAfter 2s >= 2s; run 2 (hour-1) request 2: the 429 was never retried
r as app
1.00332/332 Jira calls of the scheduled runs and their queue continuations asApp (consumer 176, scheduled 156)
r completes in timeout
1.00788/788 invocations of the scheduled runs and their queue continuations inside their module timeout (consumer 780, scheduled 8)
r idempotent rerun
1.000 ledger write(s) (scope-ledger) on a no-change run over 399 rows
Forge KVS with the storage scope, ledger reads through the declared entity index in change-time order, no KVS limit errors.
s storage scope
1.00KVS used, no 403 from the KVS proxy
s entity index used
1.0090/90 ledger reads are entity queries on a declared index
s index order
1.006/6 tables list changes in change-time order
s limits
1.00no KVS limit errors
Every invoked resolver exists and answers structured errors, nobody sees changes to issues they cannot browse (not on screen, not in an LLM prompt, not in a Realtime payload), the hidden-change count is right, and each click posts exactly one ADF comment as the viewer.
b invoke contract
1.0057/57 invokes returned a defined, structured result; a Jira 500 on sprintLedger's read: handled
b no permission leak
1.00no hidden key, summary or change id reached a person
b hidden count
1.0018/18 hidden counts equal the oracle
b comment adf as user
1.0012/12 comments are ADF as the viewer naming key, sprint and creep
b comment exactly once
1.0012 click/double-click gesture(s), none posted more than one comment
b realtime payload clean
1.00331/331 realtime payloads hold sprint ids only
The dashboard widget's numbers and chart, its board choice saved through the host, live updates through Forge Realtime, the sprint action's ledger table, sorting, issue links, comment flow, the Forge LLM explanation, close and not-started states.
u widget loadsdiagnostic, no weight
1.006/8 configured widget views render a sprint with its numbers
u widget numbers
1.00widget numbers: 8.00/8 views exact (sprints by startDate, four §1 numbers each)
u widget chart
1.008/8 charts on one linear scale within 1px
u widget edit config
1.00edit/config 10/10:
u ledger table
0.9475worst of 3 scoring sites (seed d65c448ce6b52432): ledger tables: 5.69/6 exact (rows, cells, default order)
u ledger sort
1.0018/18 at-toggle states ordered with aria-sort (6 after another sort was active)
u issue router
1.006/6 issue keys open /browse/<KEY> through the Forge router
u comment flow
1.0032/32 comment-flow steps right (select, one comment and a success flag per click and double click, forbidden error flag, modal works)
u modal close
1.006/6 close clicks reached bridge close
u not started
1.00future sprint shows not-started only
u widget live
1.00moved sprints showing the new numbers live 1/1; 0 idle invoke(s)
u llm explain
1.00explain 23/23 over 5 scripted answers:
Atlassian design tokens with 4.5:1 contrast in light and dark, a painted surface, no Content Security Policy violations or console errors, and the widget readable at 380 px wide.
v theme tokens
1.0031/31 surfaces themed with --ds-text* at >= 4.5:1
v dark mode
1.0031/31 surfaces paint a --ds-surface* background in their mode
v csp clean
1.000 CSP violation(s), 0 failed asset request(s)
v console clean
1.000 console/page error(s) on nominal scenarios: []
v widget sizes
1.006/6 widget sizes without horizontal overflow, every sprint visible, no number clipped
The get-sprint-scope action returns exact numbers and the visible changes for the invoking person, returns errors instead of throwing, and the skill's SKILL.md tells the agent how to use it.
a action result
0.00worst of 3 scoring sites (seed d65c448ce6b52432): 0/5 active sprints answered exactly
a action errors
1.002/2 bad sprintId inputs return {error} without throwing
a action permissions
1.0012/12 per-person answers list exactly what that person may browse
a skill instructions
1.00SKILL.md body 5/5: result fields named ['sprintName', 'committed', 'added', 'removed', 'creepPercent', 'hiddenChanges', 'changes']
Few Jira requests in the first scheduled run and per relevant change in triggers and consumers, and surfaces that paint after one round trip — paid only in proportion to the excellence gate.
e reconcile economy
1.00mean over 3 scoring sites: backfill 31 background Jira calls (by lineage, every phase) / optimum 20 (incl. 2 forced 429 repeat(s)) = 1.55x; 1.0 at <= 2x; 97 2.0 call(s) not counted (field-value writes, person-facing, web
e event economy
0.6105mean over 3 scoring sites: 898 Jira calls (by lineage, every phase) / baseline 216 (215 delivered relevant change(s) x 1 read + 1 forced 429 repeat; a normalizer, not a reachable optimum) = 4.16x; 1.0 at <= 2.58x; 159 fi
e ui round trips
1.00mean over 3 scoring sites: round trips before first paint: {'widget-view': 0, 'widget-edit': 1, 'sprint-action': 1}
Every v1 row moves to the new scope-ledger entity exactly once, with its change id and time, while events keep flowing and across invocation limits, within 2 virtual hours of the upgrade, with its progress in the admin panel; v1's entity stays intact.
r1 v1 rows migrated
1.00h2: 85 of 85 v1 rows in scope-ledger with their changeId and time
r1 v1 rows intact
1.00all 85 v1 rows intact in scope-change after 85 were migrated
r1 events during migration
1.0069 of 69 changes made while the migration ran are in scope-ledger at the end
r1 progress visible
1.006 of 6 panel reads show the migration right (total 85; complete after the h2 read)
On a site of about 1,000 issues, background work stays inside its share of the 2,400-point hour, no person-facing request meets a quota 429, and Retry-After and RateLimit-Reason (quota, burst, per-issue write) decide every retry; the backfill still completes.
r2 background share
1.009 of 9 hours kept background within 70 % (1680 points; the contract default; the probe recorded no change); busiest hour 398; dosing optimum for the whole run 269
r2 person never quota refused
1.000 of 267 person-facing requests refused for quota
r2 rate limit reaction
1.0014 of 14 429s answered by their reason (quota: all background paused; burst: the endpoint; per-issue: the issue)
r2 backfill complete
1.00114 of 114 pre-upgrade changes v1 never recorded are in scope-ledger, right
Work longer than one invocation continues in the next, no function waits past its limit (a long Retry-After is handed back, not waited out), and a killed invocation leaves no duplicate.
r3 never killed
1.007 of 7 functions never exceeded their limit
r3 long retry after deferred
1.001 of 1 invocations handed a Retry-After longer than their limit stopped and resumed later (results: retry)
r3 no duplicate rows
1.00every (changeId, sprintId) appears once at each of 7 checkpoints; 11 concurrent group(s) raced (same-event 8, same-issue 3)
Sprints close, issues move to another board's sprint, a board's estimation field changes, issues are deleted and a person loses browse permission while the app runs; the ledger and every view come out right.
r4 sprint close final
1.001 of 1 closed sprints kept exactly their final ledger
r4 move new board
1.0014 of 14 cross-board moves recorded with the new board's estimate
r4 field switch
1.0037 of 37 rows after the switch use the new field; 34 of 34 earlier rows keep their estimate
r4 deleted history
1.002 of 2 rows of deleted issues kept as history with deleted: true
r4 permission revoked
1.003 answer(s) after a browse revoke showed ledger data and none of the newly hidden issues
A jira:adminPage built in UI Kit with the contract's controls, and every admin resolver checking Jira's ADMINISTER permission on the server, never trusting the payload.
r5 panel labels
0.907 of 8 §2.6 controls found by label; new secret shown once: yes; then masked ••••<last4>: yes; missing migration
r5 admin actions apply
1.002 of 2 admin-page invokes of the admin's save and rotate applied
r5 nonadmin refused
1.004 non-admin/forged admin actions changed nothing — judged on the attempted effect landed in storage
Signed deployment events: a missing or wrong signature, a tampered body or a stale timestamp is refused with no effect, a replay applies once, a valid event records the deployment, and the secret is never disclosed.
r6 unsigned no effect
1.005 invalid events (bad-signature, stale, stale-future, tampered, unsigned) had no effect
r6 replay once
1.001 replay(s) made no Jira write and doubled no deployment; bookkeeping writes 0
r6 status codes
1.008 of 8 cases answered with their status
r6 valid applies
1.002 of 2 validly signed events accepted (202) and recorded as `Deployed to <env>` on every ledger row of the issues they name
r6 secret never disclosed
1.00the CI secret (panel) never appeared outside its one-time display
The read-only scope-status field holds committed or added +points for an issue in an active sprint, removed for one that has left an active sprint since that sprint started, and nothing for any other issue, written in bulk through Jira's app field-value API and fresh within the virtual hour.
r7 values fresh
0.959worst of 3 scoring sites (seed d65c448ce6b52432): mean over 7 checkpoints; worst h3: 274 of 361 issues right; DATA-247 'added +3' (want added +8, removed); OPS-273 'added +0' (want added +2)
r7 writes accepted
1.00272 of 272 bulk field writes accepted; 9 rate-limited (429, R2's)
Tool calls validated against the viewer's sprint scope, back-off on a 429 without Retry-After, answers with no finish reason refused, a 10-minute cache, and the admin's kill switch and daily token budget enforced on the server.
r8 tool call scope
1.001 of 1 manipulated tool calls validated against the viewer's sprint scope before acting
r8 failure handling
1.002 of 2 LLM failures handled
r8 cost controls
1.003 of 3 cache/kill-switch/budget cases made no model call
The widget and the sprint action each make at most one invoke and load at most 150 KB before their first data paint, with no external origin; the admin page renders after one invoke.
r9 widget boot
1.00widget-view: 1 invoke(s) before first paint (<= 1), 39246 bytes of JS+CSS (<= 153600), 0 external request(s)
r9 sprint boot
1.00sprint-modal: 1 invoke(s) before first paint (<= 1), 43360 bytes of JS+CSS (<= 153600), 0 external request(s)
r9 admin boot
1.00admin-page: 1 invoke(s) before first paint (<= 1)
Graded browser recording
Loading duration…Watch the full graded browser recording at its original speed: the dashboard widget, the sprint action, the not-started sprint and the Forge LLM cases, as the scorer drove them on the run's own seed, the first of the three scoring sites. The test results and screenshots on this page provide the wider evidence. Pause or scrub to inspect a view.
The UI Kit admin panel is not in the recording: it is graded from its component tree.
Drag the timeline to seek. Keyboard: arrow keys move 5 seconds; Home and End jump to the start and end.
SHA-256: 66fdea9ece979d805241f8b9a1efe6925f0816c7ecf31e12e5e4434771e911c5
Screenshots
Pictures from one of the three scoring sites, the one on the run's own seed. The dashboard widget and the sprint action are captured from the built app in a browser during grading. The admin panel picture is drawn by the benchmark from the app's component tree, not by Jira, and nothing is graded from that picture. A test keeps the score of its worst site (the excellence measurements keep their mean), so a test can fail on a site these pictures do not show. A detail that opens "worst of …" names the seed of the site the test scored worst on, which can be the pictured one.
Run details
- Model
- openai/gpt-6.1-sol
- Engine events
- 0
- Repair rounds
- 0
- Started
- Oct 10, 2026, 07:17 PM
- Finished
- Oct 10, 2026, 08:14 PM