Cloud baseline
fireworks/ember-1, single model via OpenRouter — 0.4580 on Forge 1.0
fireworks/ember-1 · scorer forge-1.0
- Overall
- 0.4580
- Excellent
- no
- Benchmark
- Forge 1.0
- Model build
- 25m 14s
Tier breakdown
Scoring detail
How this score was built
Per-check scorer output, as posted by the app. Expand a section to inspect its checks.
Earned credit before admission
(0.88 × 0.7075 core + 0.12 × 0.68 gate × 0.33 excellence) × 0.7049 critical = 0.4580
Critical defects compound a multiplier on the whole score (pre-severity 0.6498):
The core (88% of the total) is the weighted mean of the nine measured tiers. The last 12% is the excellence slice: it unlocks in proportion to the perfection conditions below (10 of 18 met here), then pays out at the excellence tier's own measured mean. Core 0.6226 + excellence 0.0271.
Earned score 0.458 · Cap 0.699 · Final score 0.458
The final score is the lower of the earned score and the cap. Meeting a cap's conditions adds no points.
- working ledger: maximum 0.699
t_event_rowsr_backfill_completes_entity_index_used - current platform, complete surfaces: maximum 0.799
k_widget_edit_bridgek_rovo_skillu_widget_edit_configu_ledger_tablea_action_result - production robustness: maximum 0.799
t_out_of_ordert_retry_after_honouredr_heal_droppedr_paginationu_llm_explainu_widget_liveu_ledger_sortb_hidden_counta_action_permissions
Forge's own linter reports no errors (warnings cost points), every manifest function bundles and loads, packages resolve at the kit's pins, the manifest rules the linter misses, and only the scopes the app's calls need.
l deployable
1.00lint: 0 errors, every stage reached
l bundles load
1.006/6 functions bundle, load and export their handler
l lint warnings
1.000 distinct lint warning(s): []
l real packages
1.006/6 bundles resolve @forge/* at the kit pins only
l manifest rules
1.006/6 manifest rules met:
l scopes
1.00scopes: missing [], extra []
Current modules and APIs: dashboards:widget with its edit API, the rovo:skill and rovo:mcp wiring, a Forge LLM model the platform lists as active, no /rest/api/3/search, no @forge/api storage, a nodejs22.x or nodejs24.x runtime, the consumer shape, the declared KVS entity, and deploy readiness: a manifest and permissions real Forge would accept and no code predicted to fail there.
k dashboard widget
1.00dashboards:widget 2/2: {'custom_ui_resource': True, 'edit_resource': True}
k widget edit bridge
0.501/2 board picks reached the dashboard through the widget edit API (updateConfig/onProductSave)
k rovo skill
0.86rovo:skill 6/7: agent_lists_skill
k current apis
1.00current APIs 5/5:
k consumer shape
1.00consumer shape 3/3
k entity declared
0.00entities with a sprint-partitioned, ranged index: []
k rovo mcp
1.00rovo:mcp 4/4:
k llm model current
1.005 LLM call(s); model ids list() does not return: []; llm module with claude: True
k manifest semantics
1.00deploy readiness, manifest + permissions: 17 rules checked, none failing here
k runtime risks
1.00deploy readiness, predicted runtime failures: 12 rules checked, none failing here
Issue updates flow trigger to queue to consumer: one ledger row per change under duplicate, reordered and dropped deliveries, multi-sprint changelog values, re-estimates followed, Retry-After honoured, no user context in background work.
t trigger handoff
0.87worst of 3 scoring sites (seed 2f4e096558a50b25): 34/39 deliveries handed off correctly; ['86062: jira 0, pushes 0', '86092: jira 0, pushes 0', '86092: jira 0, pushes 0']
t event rows
0.16worst of 3 scoring sites (seed f4ab5f2d33c80122): event rows: 2/23 exact (precision 1.00, recall 0.09)
t no double count
1.00every change holds at most one ledger row in every phase (duplicated deliveries included)
t out of order
0.08worst of 3 scoring sites (seed d7ae43c78dfbd9a9): permuted deliveries: 0/11 rows exact, 1/2 sprints with oracle numbers
t multi sprint parse
0.34worst of 3 scoring sites (seed f4ab5f2d33c80122): 12/35 multi-id Sprint changes recorded exactly
t reestimate followed
0.33worst of 3 scoring sites (seed 2f4e096558a50b25): 1/3 re-estimated sprints show the oracle numbers
t retry after honoured
0.00worst of 3 scoring sites (seed 2f4e096558a50b25): waited 30.0 virtual s in-invocation; the faulted change did not land exactly once
t no user in async
1.00background work is asApp only
The first scheduled run backfills every change since each active sprint started, removals included; later runs heal what the event stream missed, page through results in each endpoint's own style, wait out rate limits and finish inside the module timeout.
r backfill complete
0.26worst of 3 scoring sites (seed d7ae43c78dfbd9a9): 16/61 historical changes recorded exactly after the first scheduled run
r removals found
0.00worst of 3 scoring sites (seed d7ae43c78dfbd9a9): 0/11 removed-to-backlog changes found
r heal dropped
0.33worst of 3 scoring sites (seed f4ab5f2d33c80122): 2/6 dropped changes healed exactly once; 0 other row(s) added
r pagination
0.58worst of 3 scoring sites (seed f4ab5f2d33c80122): 70/120 reads the site served in two or more pages walked to their end, each page once, continuing where the site left off (/rest/agile/1.0/board 0/7, /rest/agile/1.0/boar
r rate limit
1.00backfill request 2: waited 2.0 virtual s in-invocation; backfill continuation page: waited 2.0 virtual s in-invocation; heal request 2: waited 2.0 virtual s in-invocation
r as app
1.0038/38 scheduled-run Jira calls asApp
r completes in timeout
1.003/3 scheduled invocations inside the module timeout
r idempotent rerun
1.000 entity write(s) on a no-change run over 30 rows
Forge KVS with the storage scope, ledger reads through the declared entity index in change-time order, no KVS limit errors.
s storage scope
1.00KVS used, no 403 from the KVS proxy
s entity index used
0.000/29 ledger reads are entity queries on the sprint index
s index order
1.001/1 tables list changes in change-time order
s limits
1.00no KVS limit errors
Every invoked resolver exists and answers structured errors, nobody sees changes to issues they cannot browse (not on screen, not in an LLM prompt, not in a Realtime payload), the hidden-change count is right, and each click posts exactly one ADF comment as the viewer.
b invoke contract
1.0028/28 invokes returned a defined, structured result; a Jira 500 on getSprintView's read: handled
b no permission leak
1.00no hidden key, summary or change id reached a person
b hidden count
0.56worst of 3 scoring sites (seed d7ae43c78dfbd9a9): 5/9 hidden counts equal the oracle
b comment adf as user
0.00worst of 3 scoring sites (seed 2f4e096558a50b25): 0/2 comments are ADF as the viewer naming key, sprint and creep; ['CHK-442 status 201 via user', 'CHK-442 status 201 via user']
b comment exactly once
1.002 click/double-click posts each ended with exactly one comment and one success flag (the 429-once fault included)
b realtime payload clean
1.009/9 realtime payloads hold sprint ids only
The dashboard widget's numbers and chart, its board choice saved through the host, live updates through Forge Realtime, the sprint action's ledger table, sorting, issue links, comment flow, the Forge LLM explanation, close and not-started states.
u widget loads
1.004/4 configured widget views render a sprint with its numbers
u widget numbers
0.13worst of 3 scoring sites (seed 2f4e096558a50b25): widget numbers: 0.50/4 views exact (sprints by startDate, four §1 numbers each)
u widget chart
0.000/4 charts on one linear scale within 1px; ['3 rects for 6 sprint x series', '3 rects for 6 sprint x series']
u widget edit config
0.33edit/config 2/6: pick0_view_shows_board, pick1_view_shows_board, pick1_reopen_pressed, second_instance_own_board
u ledger table
0.24worst of 3 scoring sites (seed d7ae43c78dfbd9a9): ledger tables: 0.72/3 exact (rows, cells, default order)
u ledger sort
0.00worst of 3 scoring sites (seed 2f4e096558a50b25): 0/2 at-toggle states ordered with aria-sort
u issue router
1.001/1 issue keys open /browse/<KEY> through the Forge router
u comment flow
1.004/4 comment-flow steps right (select, success flag, forbidden error flag, modal works)
u modal close
1.003/3 close clicks reached bridge close
u not started
1.00future sprint shows not-started only
u widget live
0.75live widget 3/4: shows_new_numbers; 0 idle invoke(s), sprints right after the live changes 1/2
u llm explain
0.96worst of 3 scoring sites (seed 2f4e096558a50b25): explain 22/23 over 5 scripted answers: 1:digits:every_number_is_a_ledger_number
Atlassian design tokens with 4.5:1 contrast in light and dark, a painted surface, no Content Security Policy violations or console errors, and the widget readable at 380 px wide.
v theme tokens
1.0017/17 surfaces themed with --ds-text* at >= 4.5:1
v dark mode
1.0017/17 surfaces paint a --ds-surface* background in their mode
v csp clean
1.000 CSP violation(s), 0 failed asset request(s)
v console clean
1.000 console/page error(s) on nominal scenarios: []
v widget sizes
1.004/4 widget sizes without horizontal overflow, every sprint visible, no number clipped
The get-sprint-scope action returns exact numbers and the visible changes for the invoking person, returns errors instead of throwing, and the skill's SKILL.md tells the agent how to use it.
a action result
0.331/3 active sprints answered exactly
a action errors
1.002/2 bad sprintId inputs return {error} without throwing
a action permissions
0.332/6 per-person answers list exactly what that person may browse
a skill instructions
1.00SKILL.md body 5/5: result fields named ['sprintName', 'committed', 'added', 'removed', 'creepPercent', 'hiddenChanges', 'changes']
Few Jira requests in the backfill and per event, few round trips before each surface paints, and a scheduled run with nothing new writes nothing — paid only in proportion to the excellence gate.
e reconcile economy
0.00mean over 3 scoring sites: backfill incomplete — economy is not credited on unfinished work
e event economy
0.00mean over 3 scoring sites: event rows inexact — economy is not credited on unfinished work
e ui round trips
1.00mean over 3 scoring sites: round trips before first paint: {'widget-view': 0, 'widget-edit': 1, 'sprint-action': 1}
Graded browser recording
Loading duration…Watch the full graded browser recording at its original speed: the payment field, structure inspection, committed updates and replay. The check results and screenshots on this page provide the wider evidence. Pause or scrub to inspect a view.
Each tower represents a payment: height encodes its amount, cap color its status, and the collar its currency. The selected detailed tower is expected to animate its collar after a committed backend update or replay. These are the task requirements; the check results show what this app actually achieved.
Drag the timeline to seek. Keyboard: arrow keys move 5 seconds; Home and End jump to the start and end.
Recording integrity
SHA-256: 45c1470c6d43801e118e4fa8b25987fb294ebc8977662fc6743e8cd5236d4f6c
Screenshots
Captured from the built application during browser grading. Captions identify the recorded views and checks; these images do not represent repair rounds.
Run details
- Model
- fireworks/ember-1
- Engine events
- 0
- Repair rounds
- 0
- Started
- Oct 6, 2026, 08:30 PM
- Finished
- Oct 6, 2026, 09:03 PM