Cloud baseline

openai/gpt-6-luna, single model via OpenRouter — 0.6250 on Forge 2.0

openai/gpt-6-luna · scorer forge-2.0

openai/gpt-6-luna · single agentOct 10, 2026
Overall
0.6250
Benchmark
Forge 2.0
Model build
44m 15s

Tier breakdown

Each bar is the tier's mean, from 0 to 1 (for E, its mean × the excellence gate). The percentage is the tier's share of the score. Mean × share, added over all tiers, is the tests earned figure below.

L · Lint and bundles1.001.76%
K · Platform currency0.99292.2%
T · Event pipeline0.97923.52%
R · Reconcile0.93673.08%
S · Storage1.001.76%
B · Resolvers and permissions0.97222.64%
U · UI function0.87793.52%
V · Visual1.001.76%
A · Rovo0.751.76%
E · Excellence0.97713%
R1 · Migration1.0012%
R2 · Dosing0.958313%
R3 · Time limits1.008%
R4 · Changing world0.996510%
R5 · Admin panel0.966710%
R6 · CI web trigger0.908%
R7 · Custom field0.95755%
R8 · Forge LLM1.004%
R9 · Boot budget0.77785%

Scoring detail

How this score was built

Per-test scorer output, as posted by the app. Expand a section to inspect its tests.

The rules behind these numbers: Forge 2.0 requirements, weights and score caps.

The score, step by step

0.9567 tests earned × 1.0000 critical defects × 0.6532 reliability = 0.6250 final score

tests earned = 0.25 × 0.9482 v1 + 0.7197 v2 (out of 0.75)

v1 = 0.88 × 0.9442 tiers L–A + 0.12 × 0.9771 gate × 1.0000 excellence

Reliability: every tier in the breakdown above, except Excellence, is a group of tests. Each group of tests multiplies the score by 1 − 0.10 × the shortfall of its worst test: by 0.90 when that test fails completely, by 0.95 when it scores 0.5.

  • K Platform currency: worst test k_manifest_semantics, score 0.9286 × 0.9929
  • T Event pipeline: worst test t_reestimate_followed, score 0.8333 × 0.9833
  • R Reconcile: worst test r_pagination, score 0.4937 × 0.9494
  • B Resolvers and permissions: worst test b_comment_adf_as_user, score 0.8333 × 0.9833
  • U UI function: worst test u_widget_live, score 0 × 0.9000

    Already counted in this group: u_widget_numbers, u_widget_chart, u_ledger_table, u_llm_explain

  • A Rovo: worst test a_action_result, score 0 × 0.9000
  • R2 Dosing: worst test r2_rate_limit_reaction, score 0.8333 × 0.9833
  • R4 Changing world: worst test r4_field_switch, score 0.9825 × 0.9983
  • R5 Admin panel: worst test r5_panel_labels, score 0.9 × 0.9900
  • R6 CI web trigger: worst test r6_valid_applies, score 0.5 × 0.9500
  • R7 Custom field: worst test r7_values_fresh, score 0.9149 × 0.9915
  • R9 Boot budget: worst test r9_widget_boot, score 0.6667 × 0.9667

    Already counted in this group: r9_sprint_boot

v1 is this run's score on the tests carried over from Forge 1.0: 25% of the total. v2 is the requirements R1–R9, each at its share of the score: the other 75%. Critical defects multiply the whole score, and so does each group of tests, by 1 − 0.10 × the shortfall of its worst test.

Excellence gate 0.9771: the average of the 18 conditions below, each counted at the value shown. Green is met in full; red is not, and lowers the gate. The 3 excellence tests are paid in proportion to the gate; together they are 3% of the total.

t_event_rows 1.00r_heal_dropped 1.00u_widget_numbers 0.875u_ledger_table 0.9504b_comment_adf_as_user 0.8333k_consumer_shape 1.00k_current_apis 1.00k_dashboard_widget 1.00k_entity_declared 1.00k_llm_model_current 1.00k_manifest_semantics 0.9286k_rovo_mcp 1.00k_rovo_skill 1.00k_runtime_risks 1.00k_widget_edit_bridge 1.00v_console_clean 1.00v_csp_clean 1.00l_lint_warnings 1.00

Earned score 0.6250 · No cap applies · Final score 0.6250

Under a cap the final score is min(earned, cap − 0.05 × (1 − earned)): always below the cap, and closer to it the more the run earned. With no cap it is the earned score. Avoiding a cap adds no points.

The get-sprint-scope action returns exact numbers and the visible changes for the invoking person, returns errors instead of throwing, and the skill's SKILL.md tells the agent how to use it.

a action result

0.00

worst of 3 scoring sites (seed ae58692bcd067755): 0/5 active sprints answered exactly

a action errors

1.00

2/2 bad sprintId inputs return {error} without throwing

a action permissions

1.00

12/12 per-person answers list exactly what that person may browse

a skill instructions

1.00

SKILL.md body 5/5: result fields named ['sprintName', 'committed', 'added', 'removed', 'creepPercent', 'hiddenChanges', 'changes']

Graded browser recording

Loading duration…

Watch the full graded browser recording at its original speed: the dashboard widget, the sprint action, the not-started sprint and the Forge LLM cases, as the scorer drove them on the run's own seed, the first of the three scoring sites. The test results and screenshots on this page provide the wider evidence. Pause or scrub to inspect a view.

The UI Kit admin panel is not in the recording: it is graded from its component tree.

Ready to play0:00 / —:—
Open recording file

Drag the timeline to seek. Keyboard: arrow keys move 5 seconds; Home and End jump to the start and end.

Full graded browser recording: dashboard widget (before configuration, edit and Save, light and dark at 380 px, a second instance at 1180 px), sprint action (light and dark, sort, issue link, select, comment post, double click, forbidden post, explain, close), the not-started sprint and the Forge LLM cases, where the app has each. The UI Kit admin panel is not in the recording.

Screenshots

Pictures from one of the three scoring sites, the one on the run's own seed. The dashboard widget and the sprint action are captured from the built app in a browser during grading. The admin panel picture is drawn by the benchmark from the app's component tree, not by Jira, and nothing is graded from that picture. A test keeps the score of its worst site (the excellence measurements keep their mean), so a test can fail on a site these pictures do not show. A detail that opens "worst of …" names the seed of the site the test scored worst on, which can be the pictured one.

1Dashboard widget · light
2Dashboard widget · dark
3Sprint action · light
4Sprint action · dark
5Admin panel (UI Kit) · light: drawn by the benchmark from the app's component tree, not by Jira.

Run details

Model
openai/gpt-6-luna
Engine events
0
Repair rounds
0
Started
Oct 10, 2026, 08:15 PM
Finished
Oct 10, 2026, 09:18 PM

How the benchmark is scored