Cloud baseline
openai/gpt-6-luna, single model via OpenRouter — 0.6250 on Forge 2.0
openai/gpt-6-luna · scorer forge-2.0
- Overall
- 0.6250
- Benchmark
- Forge 2.0
- Model build
- 44m 15s
Tier breakdown
Each bar is the tier's mean, from 0 to 1 (for E, its mean × the excellence gate). The percentage is the tier's share of the score. Mean × share, added over all tiers, is the tests earned figure below.
Scoring detail
How this score was built
Per-test scorer output, as posted by the app. Expand a section to inspect its tests.
The rules behind these numbers: Forge 2.0 requirements, weights and score caps.
The score, step by step
0.9567 tests earned × 1.0000 critical defects × 0.6532 reliability = 0.6250 final score
tests earned = 0.25 × 0.9482 v1 + 0.7197 v2 (out of 0.75)
v1 = 0.88 × 0.9442 tiers L–A + 0.12 × 0.9771 gate × 1.0000 excellence
Reliability: every tier in the breakdown above, except Excellence, is a group of tests. Each group of tests multiplies the score by 1 − 0.10 × the shortfall of its worst test: by 0.90 when that test fails completely, by 0.95 when it scores 0.5.
- K Platform currency: worst test
k_manifest_semantics, score 0.9286 × 0.9929 - T Event pipeline: worst test
t_reestimate_followed, score 0.8333 × 0.9833 - R Reconcile: worst test
r_pagination, score 0.4937 × 0.9494 - B Resolvers and permissions: worst test
b_comment_adf_as_user, score 0.8333 × 0.9833 - U UI function: worst test
u_widget_live, score 0 × 0.9000Already counted in this group: u_widget_numbers, u_widget_chart, u_ledger_table, u_llm_explain
- A Rovo: worst test
a_action_result, score 0 × 0.9000 - R2 Dosing: worst test
r2_rate_limit_reaction, score 0.8333 × 0.9833 - R4 Changing world: worst test
r4_field_switch, score 0.9825 × 0.9983 - R5 Admin panel: worst test
r5_panel_labels, score 0.9 × 0.9900 - R6 CI web trigger: worst test
r6_valid_applies, score 0.5 × 0.9500 - R7 Custom field: worst test
r7_values_fresh, score 0.9149 × 0.9915 - R9 Boot budget: worst test
r9_widget_boot, score 0.6667 × 0.9667Already counted in this group: r9_sprint_boot
v1 is this run's score on the tests carried over from Forge 1.0: 25% of the total. v2 is the requirements R1–R9, each at its share of the score: the other 75%. Critical defects multiply the whole score, and so does each group of tests, by 1 − 0.10 × the shortfall of its worst test.
Excellence gate 0.9771: the average of the 18 conditions below, each counted at the value shown. Green is met in full; red is not, and lowers the gate. The 3 excellence tests are paid in proportion to the gate; together they are 3% of the total.
Earned score 0.6250 · No cap applies · Final score 0.6250
Under a cap the final score is min(earned, cap − 0.05 × (1 − earned)): always below the cap, and closer to it the more the run earned. With no cap it is the earned score. Avoiding a cap adds no points.
Forge's own linter reports no errors (warnings cost points), every manifest function bundles and loads, packages resolve at the kit's pins, the manifest rules the linter misses, and only the scopes the app's calls need.
l deployable
1.00lint: 0 errors, every stage reached
l bundles load
1.0010/10 functions bundle, load and export their handler
l lint warnings
1.000 distinct lint warning(s): []
l real packages
1.0010/10 bundles resolve @forge/* at the kit pins only
l manifest rules
1.006/6 manifest rules met:
l scopes
1.00scopes: missing [], extra []
Current modules and APIs: dashboards:widget with its edit API, the rovo:skill and rovo:mcp wiring, a Forge LLM model the platform lists as active, no /rest/api/3/search, no @forge/api storage, a nodejs22.x or nodejs24.x runtime, the consumer shape, the declared KVS entity, and deploy readiness: a manifest and permissions real Forge would accept and no code predicted to fail there.
k dashboard widget
1.00dashboards:widget 2/2: {'custom_ui_resource': True, 'edit_resource': True}
k widget edit bridge
1.004/4 board picks reached the dashboard through the widget edit API (updateConfig/onProductSave)
k rovo skill
1.00rovo:skill 7/7:
k current apis
1.00current APIs 5/5:
k consumer shape
1.00consumer shape 3/3
k entity declared
1.00entities with a sprint-partitioned, ranged index: ['scope-change', 'scope-ledger', 'sprint-issue']
k v2 surfacesdiagnostic, no weight
1.00v2 surfaces declared: ['jira:adminPage', 'webtrigger', 'scope-ledger']
k rovo mcp
1.00rovo:mcp 4/4:
k llm model current
1.0011 LLM call(s); model ids list() does not return: []; llm module with claude: True
k manifest semantics
0.9286deploy readiness, manifest + permissions: 17 rules checked; failing (priced here): P2 declared scopes cover every route in the code (OAuth2 rule) [src/index.js: /rest/api/3/group/member needs one of [['manage:jira-config
k runtime risks
1.00deploy readiness, predicted runtime failures: 12 rules checked, none failing here
Issue updates flow trigger to queue to consumer: one ledger row per change under duplicate, reordered, dropped and concurrent deliveries, multi-sprint changelog values, re-estimates followed, Retry-After honoured, no user context in background work.
t trigger handoff
1.001044/1044 deliveries handed off correctly
t event rows
1.00event rows: 173/173 exact (precision 1.00, recall 1.00)
t no double count
1.00every change holds at most one ledger row in every phase (duplicated deliveries included); 11 concurrent group(s) raced (same-event 8, same-issue 3)
t out of order
1.00permuted deliveries: 9/9 rows exact, 4/4 sprints with oracle numbers
t multi sprint parse
1.0094/94 multi-id Sprint changes recorded exactly
t reestimate followed
0.8333worst of 3 scoring sites (seed ae58692bcd067755): 5/6 re-estimated sprints show the oracle numbers
t retry after honoured
1.00InvocationError retryAfter 30s >= 30s
t no user in async
1.00background work is asApp only
The first scheduled run starts the backfill of every change since each active sprint started, removals included; later runs heal what the event stream missed and a run with nothing new writes nothing; they page through results in each endpoint's own style, never retry before Retry-After and finish inside the module timeout, queue continuations included.
r backfill complete
1.0079/79 historical changes v1 never recorded, recorded one virtual hour after the upgrade (the backfill is dosed; v1 rows are R1's, at 2 h)
r removals found
1.0014/14 removed-to-backlog changes found
r heal dropped
1.004/4 dropped changes healed exactly once; 0 other row(s) added
r pagination
0.493739/79 reads the site served in two or more pages walked to their end, each page once, continuing where the site left off (/rest/agile/1.0/board 8/16, /rest/agile/1.0/board/{id}/sprint 7/39, /rest/api/3/changelog/bulkfetc
r rate limit
1.00run 1 (backfill) request 2: waited 3.0 virtual s in-invocation; run 1 (backfill) continuation page: waited 3.0 virtual s in-invocation; run 2 (hour-1) request 2: waited 3.0 virtual s in-invocation
r as app
1.00353/353 Jira calls of the scheduled runs and their queue continuations asApp (scheduled 353)
r completes in timeout
1.008/8 invocations of the scheduled runs and their queue continuations inside their module timeout (scheduled 8)
r idempotent rerun
1.000 ledger write(s) (scope-ledger) on a no-change run over 362 rows
Forge KVS with the storage scope, ledger reads through the declared entity index in change-time order, no KVS limit errors.
s storage scope
1.00KVS used, no 403 from the KVS proxy
s entity index used
1.0022961/22961 ledger reads are entity queries on a declared index
s index order
1.006/6 tables list changes in change-time order
s limits
1.00no KVS limit errors
Every invoked resolver exists and answers structured errors, nobody sees changes to issues they cannot browse (not on screen, not in an LLM prompt, not in a Realtime payload), the hidden-change count is right, and each click posts exactly one ADF comment as the viewer.
b invoke contract
1.0057/57 invokes returned a defined, structured result; a Jira 500 on sprintLedger's read: handled
b no permission leak
1.00no hidden key, summary or change id reached a person
b hidden count
1.0018/18 hidden counts equal the oracle
b comment adf as user
0.8333worst of 3 scoring sites (seed ae58692bcd067755): 10/12 comments are ADF as the viewer naming key, sprint and creep; ['27574 status 201 via user', '28116 status 201 via user']
b comment exactly once
1.0012 click/double-click gesture(s), none posted more than one comment
b realtime payload clean
1.00193/193 realtime payloads hold sprint ids only
The dashboard widget's numbers and chart, its board choice saved through the host, live updates through Forge Realtime, the sprint action's ledger table, sorting, issue links, comment flow, the Forge LLM explanation, close and not-started states.
u widget loadsdiagnostic, no weight
1.006/8 configured widget views render a sprint with its numbers
u widget numbers
0.875worst of 3 scoring sites (seed 657d0d781d10ed46): widget numbers: 7.00/8 views exact (sprints by startDate, four §1 numbers each)
u widget chart
0.8757/8 charts on one linear scale within 1px; ['worst bar off its linear scale by 6.44px']
u widget edit config
1.00edit/config 10/10:
u ledger table
0.9504worst of 3 scoring sites (seed 657d0d781d10ed46): ledger tables: 5.70/6 exact (rows, cells, default order)
u ledger sort
1.0018/18 at-toggle states ordered with aria-sort (6 after another sort was active)
u issue router
1.006/6 issue keys open /browse/<KEY> through the Forge router
u comment flow
1.0032/32 comment-flow steps right (select, one comment and a success flag per click and double click, forbidden error flag, modal works)
u modal close
1.006/6 close clicks reached bridge close
u not started
1.00future sprint shows not-started only
u widget live
0.00moved sprints showing the new numbers live 0/1; 0 idle invoke(s)
u llm explain
0.9565worst of 3 scoring sites (seed ae58692bcd067755): explain 22/23 over 5 scripted answers: 1:digits:every_number_is_a_ledger_number
Atlassian design tokens with 4.5:1 contrast in light and dark, a painted surface, no Content Security Policy violations or console errors, and the widget readable at 380 px wide.
v theme tokens
1.0031/31 surfaces themed with --ds-text* at >= 4.5:1
v dark mode
1.0031/31 surfaces paint a --ds-surface* background in their mode
v csp clean
1.000 CSP violation(s), 0 failed asset request(s)
v console clean
1.000 console/page error(s) on nominal scenarios: []
v widget sizes
1.006/6 widget sizes without horizontal overflow, every sprint visible, no number clipped
The get-sprint-scope action returns exact numbers and the visible changes for the invoking person, returns errors instead of throwing, and the skill's SKILL.md tells the agent how to use it.
a action result
0.00worst of 3 scoring sites (seed ae58692bcd067755): 0/5 active sprints answered exactly
a action errors
1.002/2 bad sprintId inputs return {error} without throwing
a action permissions
1.0012/12 per-person answers list exactly what that person may browse
a skill instructions
1.00SKILL.md body 5/5: result fields named ['sprintName', 'committed', 'added', 'removed', 'creepPercent', 'hiddenChanges', 'changes']
Few Jira requests in the first scheduled run and per relevant change in triggers and consumers, and surfaces that paint after one round trip — paid only in proportion to the excellence gate.
e reconcile economy
1.00mean over 3 scoring sites: backfill 40 background Jira calls (by lineage, every phase) / optimum 20 (incl. 2 forced 429 repeat(s)) = 2.00x; 1.0 at <= 2x; 3 2.0 call(s) not counted (field-value writes, person-facing, web
e event economy
1.00mean over 3 scoring sites: 221 Jira calls (by lineage, every phase) / baseline 199 (198 delivered relevant change(s) x 1 read + 1 forced 429 repeat; a normalizer, not a reachable optimum) = 1.11x; 1.0 at <= 2.58x; 0 fiel
e ui round trips
1.00mean over 3 scoring sites: round trips before first paint: {'widget-view': 0, 'widget-edit': 1, 'sprint-action': 1}
Every v1 row moves to the new scope-ledger entity exactly once, with its change id and time, while events keep flowing and across invocation limits, within 2 virtual hours of the upgrade, with its progress in the admin panel; v1's entity stays intact.
r1 v1 rows migrated
1.00h2: 106 of 106 v1 rows in scope-ledger with their changeId and time
r1 v1 rows intact
1.00all 106 v1 rows intact in scope-change after 106 were migrated
r1 events during migration
1.0054 of 54 changes made while the migration ran are in scope-ledger at the end
r1 progress visible
1.006 of 6 panel reads show the migration right (total 106; complete after the h2 read)
On a site of about 1,000 issues, background work stays inside its share of the 2,400-point hour, no person-facing request meets a quota 429, and Retry-After and RateLimit-Reason (quota, burst, per-issue write) decide every retry; the backfill still completes.
r2 background share
1.009 of 9 hours kept background within 70 % (1680 points; the contract default; the probe recorded no change); busiest hour 196; dosing optimum for the whole run 252
r2 person never quota refused
1.000 of 275 person-facing requests refused for quota
r2 rate limit reaction
0.8333worst of 3 scoring sites (seed ae58692bcd067755): 5 of 6 429s answered by their reason (quota: all background paused; burst: the endpoint; per-issue: the issue); broken: jira-burst-based at 1791486819629 ms: 1 request(s)
r2 backfill complete
1.0079 of 79 pre-upgrade changes v1 never recorded are in scope-ledger, right
Work longer than one invocation continues in the next, no function waits past its limit (a long Retry-After is handed back, not waited out), and a killed invocation leaves no duplicate.
r3 never killed
1.008 of 8 functions never exceeded their limit
r3 long retry after deferred
1.001 of 1 invocations handed a Retry-After longer than their limit stopped and resumed later (results: retry)
r3 no duplicate rows
1.00every (changeId, sprintId) appears once at each of 7 checkpoints; 11 concurrent group(s) raced (same-event 8, same-issue 3)
Sprints close, issues move to another board's sprint, a board's estimation field changes, issues are deleted and a person loses browse permission while the app runs; the ledger and every view come out right.
r4 sprint close final
1.001 of 1 closed sprints kept exactly their final ledger
r4 move new board
1.0012 of 12 cross-board moves recorded with the new board's estimate
r4 field switch
0.9825worst of 3 scoring sites (seed 657d0d781d10ed46): 47 of 48 rows after the switch use the new field; 9 of 9 earlier rows keep their estimate; wrong after: ('1187295', '989')
r4 deleted history
1.002 of 2 rows of deleted issues kept as history with deleted: true
r4 permission revoked
1.006 answer(s) after a browse revoke showed ledger data and none of the newly hidden issues
A jira:adminPage built in UI Kit with the contract's controls, and every admin resolver checking Jira's ADMINISTER permission on the server, never trusting the payload.
r5 panel labels
0.907 of 8 §2.6 controls found by label; new secret shown once: yes; then masked ••••<last4>: yes; missing migration
r5 admin actions apply
1.004 of 4 admin-page invokes of the admin's save and rotate applied
r5 nonadmin refused
1.008 non-admin/forged admin actions changed nothing — judged on the attempted effect landed in storage
Signed deployment events: a missing or wrong signature, a tampered body or a stale timestamp is refused with no effect, a replay applies once, a valid event records the deployment, and the secret is never disclosed.
r6 unsigned no effect
1.005 invalid events (bad-signature, stale, stale-future, tampered, unsigned) had no effect
r6 replay once
1.001 replay(s) made no Jira write and doubled no deployment; bookkeeping writes 0
r6 status codes
1.008 of 8 cases answered with their status
r6 valid applies
0.501 of 2 validly signed events accepted (202) and recorded as `Deployed to <env>` on every ledger row of the issues they name; valid: no `Deployed to staging` on every ledger row of DATA-253
r6 secret never disclosed
1.00the CI secret (panel) never appeared outside its one-time display
The read-only scope-status field holds committed or added +points for an issue in an active sprint, removed for one that has left an active sprint since that sprint started, and nothing for any other issue, written in bulk through Jira's app field-value API and fresh within the virtual hour.
r7 values fresh
0.9149worst of 3 scoring sites (seed ae58692bcd067755): mean over 7 checkpoints; worst h5: 274 of 316 issues right; CHK-112 'added +0.5' (want added +2, removed); CHK-157 'added +5' (want removed)
r7 writes accepted
1.0029 of 29 bulk field writes accepted
Tool calls validated against the viewer's sprint scope, back-off on a 429 without Retry-After, answers with no finish reason refused, a 10-minute cache, and the admin's kill switch and daily token budget enforced on the server.
r8 tool call scope
1.001 of 1 manipulated tool calls validated against the viewer's sprint scope before acting
r8 failure handling
1.002 of 2 LLM failures handled
r8 cost controls
1.003 of 3 cache/kill-switch/budget cases made no model call
The widget and the sprint action each make at most one invoke and load at most 150 KB before their first data paint, with no external origin; the admin page renders after one invoke.
r9 widget boot
0.6667widget-view: 1 invoke(s) before first paint (<= 1), 244622 bytes of JS+CSS (<= 153600), 0 external request(s)
r9 sprint boot
0.6667sprint-modal: 1 invoke(s) before first paint (<= 1), 248482 bytes of JS+CSS (<= 153600), 0 external request(s)
r9 admin boot
1.00admin-page: 1 invoke(s) before first paint (<= 1)
Graded browser recording
Loading duration…Watch the full graded browser recording at its original speed: the dashboard widget, the sprint action, the not-started sprint and the Forge LLM cases, as the scorer drove them on the run's own seed, the first of the three scoring sites. The test results and screenshots on this page provide the wider evidence. Pause or scrub to inspect a view.
The UI Kit admin panel is not in the recording: it is graded from its component tree.
Drag the timeline to seek. Keyboard: arrow keys move 5 seconds; Home and End jump to the start and end.
SHA-256: f7fef5995f559d11832c8955d9f3645eadde2597eff5693fbc61002bd5408fb3
Screenshots
Pictures from one of the three scoring sites, the one on the run's own seed. The dashboard widget and the sprint action are captured from the built app in a browser during grading. The admin panel picture is drawn by the benchmark from the app's component tree, not by Jira, and nothing is graded from that picture. A test keeps the score of its worst site (the excellence measurements keep their mean), so a test can fail on a site these pictures do not show. A detail that opens "worst of …" names the seed of the site the test scored worst on, which can be the pictured one.
Run details
- Model
- openai/gpt-6-luna
- Engine events
- 0
- Repair rounds
- 0
- Started
- Oct 10, 2026, 08:15 PM
- Finished
- Oct 10, 2026, 09:18 PM