Cloud baseline

openai/gpt-6.1-sol, single model via OpenRouter — 0.8111 on Forge 2.0

openai/gpt-6.1-sol · scorer forge-2.0

openai/gpt-6.1-sol · single agentOct 10, 2026
Overall
0.8111
Benchmark
Forge 2.0
Model build
32m 2s

Tier breakdown

Each bar is the tier's mean, from 0 to 1 (for E, its mean × the excellence gate). The percentage is the tier's share of the score. Mean × share, added over all tiers, is the tests earned figure below.

L · Lint and bundles1.001.76%
K · Platform currency1.002.2%
T · Event pipeline1.003.52%
R · Reconcile0.91673.08%
S · Storage1.001.76%
B · Resolvers and permissions1.002.64%
U · UI function0.99523.52%
V · Visual1.001.76%
A · Rovo0.751.76%
E · Excellence0.86763%
R1 · Migration1.0012%
R2 · Dosing1.0013%
R3 · Time limits1.008%
R4 · Changing world1.0010%
R5 · Admin panel0.966710%
R6 · CI web trigger1.008%
R7 · Custom field0.97955%
R8 · Forge LLM1.004%
R9 · Boot budget1.005%

Scoring detail

How this score was built

Per-test scorer output, as posted by the app. Expand a section to inspect its tests.

The rules behind these numbers: Forge 2.0 requirements, weights and score caps.

The score, step by step

0.9845 tests earned × 1.0000 critical defects × 0.8238 reliability = 0.8111 final score

tests earned = 0.25 × 0.9556 v1 + 0.7456 v2 (out of 0.75)

v1 = 0.88 × 0.9676 tiers L–A + 0.12 × 0.9971 gate × 0.8702 excellence

Reliability: every tier in the breakdown above, except Excellence, is a group of tests. Each group of tests multiplies the score by 1 − 0.10 × the shortfall of its worst test: by 0.90 when that test fails completely, by 0.95 when it scores 0.5.

  • R Reconcile: worst test r_rate_limit, score 0.3333 × 0.9333
  • U UI function: worst test u_ledger_table, score 0.9475 × 0.9948
  • A Rovo: worst test a_action_result, score 0 × 0.9000
  • R5 Admin panel: worst test r5_panel_labels, score 0.9 × 0.9900
  • R7 Custom field: worst test r7_values_fresh, score 0.959 × 0.9959

v1 is this run's score on the tests carried over from Forge 1.0: 25% of the total. v2 is the requirements R1–R9, each at its share of the score: the other 75%. Critical defects multiply the whole score, and so does each group of tests, by 1 − 0.10 × the shortfall of its worst test.

Excellence gate 0.9971: the average of the 18 conditions below, each counted at the value shown. Green is met in full; red is not, and lowers the gate. The 3 excellence tests are paid in proportion to the gate; together they are 3% of the total.

t_event_rows 1.00r_heal_dropped 1.00u_widget_numbers 1.00u_ledger_table 0.9475b_comment_adf_as_user 1.00k_consumer_shape 1.00k_current_apis 1.00k_dashboard_widget 1.00k_entity_declared 1.00k_llm_model_current 1.00k_manifest_semantics 1.00k_rovo_mcp 1.00k_rovo_skill 1.00k_runtime_risks 1.00k_widget_edit_bridge 1.00v_console_clean 1.00v_csp_clean 1.00l_lint_warnings 1.00

Earned score 0.8111 · No cap applies · Final score 0.8111

Under a cap the final score is min(earned, cap − 0.05 × (1 − earned)): always below the cap, and closer to it the more the run earned. With no cap it is the earned score. Avoiding a cap adds no points.

The get-sprint-scope action returns exact numbers and the visible changes for the invoking person, returns errors instead of throwing, and the skill's SKILL.md tells the agent how to use it.

a action result

0.00

worst of 3 scoring sites (seed d65c448ce6b52432): 0/5 active sprints answered exactly

a action errors

1.00

2/2 bad sprintId inputs return {error} without throwing

a action permissions

1.00

12/12 per-person answers list exactly what that person may browse

a skill instructions

1.00

SKILL.md body 5/5: result fields named ['sprintName', 'committed', 'added', 'removed', 'creepPercent', 'hiddenChanges', 'changes']

Graded browser recording

Loading duration…

Watch the full graded browser recording at its original speed: the dashboard widget, the sprint action, the not-started sprint and the Forge LLM cases, as the scorer drove them on the run's own seed, the first of the three scoring sites. The test results and screenshots on this page provide the wider evidence. Pause or scrub to inspect a view.

The UI Kit admin panel is not in the recording: it is graded from its component tree.

Ready to play0:00 / —:—
Open recording file

Drag the timeline to seek. Keyboard: arrow keys move 5 seconds; Home and End jump to the start and end.

Full graded browser recording: dashboard widget (before configuration, edit and Save, light and dark at 380 px, a second instance at 1180 px), sprint action (light and dark, sort, issue link, select, comment post, double click, forbidden post, explain, close), the not-started sprint and the Forge LLM cases, where the app has each. The UI Kit admin panel is not in the recording.

Screenshots

Pictures from one of the three scoring sites, the one on the run's own seed. The dashboard widget and the sprint action are captured from the built app in a browser during grading. The admin panel picture is drawn by the benchmark from the app's component tree, not by Jira, and nothing is graded from that picture. A test keeps the score of its worst site (the excellence measurements keep their mean), so a test can fail on a site these pictures do not show. A detail that opens "worst of …" names the seed of the site the test scored worst on, which can be the pictured one.

1Dashboard widget · light
2Dashboard widget · dark
3Sprint action · light
4Sprint action · dark
5Admin panel (UI Kit) · light: drawn by the benchmark from the app's component tree, not by Jira.

Run details

Model
openai/gpt-6.1-sol
Engine events
0
Repair rounds
0
Started
Oct 10, 2026, 07:17 PM
Finished
Oct 10, 2026, 08:14 PM

How the benchmark is scored