Cloud baseline

deepseek/deepseek-v4.1-flash, single model via OpenRouter — 0.7979 on Forge 2.0

deepseek/deepseek-v4.1-flash · scorer forge-2.0

deepseek/deepseek-v4.1-flash · single agentOct 11, 2026
Overall
0.7979
Benchmark
Forge 2.0
Model build
42m 14s

Tier breakdown

Each bar is the tier's mean, from 0 to 1 (for E, its mean × the excellence gate). The percentage is the tier's share of the score. Mean × share, added over all tiers, is the tests earned figure below.

L · Lint and bundles1.001.76%
K · Platform currency1.002.2%
T · Event pipeline0.97923.52%
R · Reconcile1.003.08%
S · Storage1.001.76%
B · Resolvers and permissions0.97222.64%
U · UI function0.96813.52%
V · Visual1.001.76%
A · Rovo1.001.76%
E · Excellence0.98973%
R1 · Migration1.0012%
R2 · Dosing0.825513%
R3 · Time limits1.008%
R4 · Changing world1.0010%
R5 · Admin panel1.0010%
R6 · CI web trigger0.908%
R7 · Custom field0.99815%
R8 · Forge LLM1.004%
R9 · Boot budget1.005%

Scoring detail

How this score was built

Per-test scorer output, as posted by the app. Expand a section to inspect its tests.

The rules behind these numbers: Forge 2.0 requirements, weights and score caps.

The score, step by step

0.9663 tests earned × 1.0000 critical defects × 0.8257 reliability = 0.7979 final score

tests earned = 0.25 × 0.9884 v1 + 0.7192 v2 (out of 0.75)

v1 = 0.88 × 0.9882 tiers L–A + 0.12 × 0.9897 gate × 1.0000 excellence

Reliability: every tier in the breakdown above, except Excellence, is a group of tests. Each group of tests multiplies the score by 1 − 0.10 × the shortfall of its worst test: by 0.90 when that test fails completely, by 0.95 when it scores 0.5.

  • T Event pipeline: worst test t_reestimate_followed, score 0.8333 × 0.9833
  • B Resolvers and permissions: worst test b_comment_adf_as_user, score 0.8333 × 0.9833
  • U UI function: worst test u_issue_router, score 0.6667 × 0.9667

    Already counted in this group: u_ledger_table

  • R2 Dosing: worst test r2_rate_limit_reaction, score 0.3019 × 0.9302
  • R6 CI web trigger: worst test r6_valid_applies, score 0.5 × 0.9500
  • R7 Custom field: worst test r7_values_fresh, score 0.9962 × 0.9996

v1 is this run's score on the tests carried over from Forge 1.0: 25% of the total. v2 is the requirements R1–R9, each at its share of the score: the other 75%. Critical defects multiply the whole score, and so does each group of tests, by 1 − 0.10 × the shortfall of its worst test.

Excellence gate 0.9897: the average of the 18 conditions below, each counted at the value shown. Green is met in full; red is not, and lowers the gate. The 3 excellence tests are paid in proportion to the gate; together they are 3% of the total.

t_event_rows 1.00r_heal_dropped 1.00u_widget_numbers 1.00u_ledger_table 0.9821b_comment_adf_as_user 0.8333k_consumer_shape 1.00k_current_apis 1.00k_dashboard_widget 1.00k_entity_declared 1.00k_llm_model_current 1.00k_manifest_semantics 1.00k_rovo_mcp 1.00k_rovo_skill 1.00k_runtime_risks 1.00k_widget_edit_bridge 1.00v_console_clean 1.00v_csp_clean 1.00l_lint_warnings 1.00

Earned score 0.7979 · No cap applies · Final score 0.7979

Under a cap the final score is min(earned, cap − 0.05 × (1 − earned)): always below the cap, and closer to it the more the run earned. With no cap it is the earned score. Avoiding a cap adds no points.

On a site of about 1,000 issues, background work stays inside its share of the 2,400-point hour, no person-facing request meets a quota 429, and Retry-After and RateLimit-Reason (quota, burst, per-issue write) decide every retry; the backfill still completes.

r2 background share

1.00

9 of 9 hours kept background within 70 % (1680 points; the contract default; the probe recorded no change); busiest hour 200; dosing optimum for the whole run 256

r2 person never quota refused

1.00

0 of 291 person-facing requests refused for quota

r2 rate limit reaction

0.3019

worst of 3 scoring sites (seed 0b27cc0bf1493673): 16 of 53 429s answered by their reason (quota: all background paused; burst: the endpoint; per-issue: the issue); broken: jira-burst-based at 1791742313089 ms: 1 request(

r2 backfill complete

1.00

108 of 108 pre-upgrade changes v1 never recorded are in scope-ledger, right

Graded browser recording

Loading duration…

Watch the full graded browser recording at its original speed: the dashboard widget, the sprint action, the not-started sprint and the Forge LLM cases, as the scorer drove them on the run's own seed, the first of the three scoring sites. The test results and screenshots on this page provide the wider evidence. Pause or scrub to inspect a view.

The UI Kit admin panel is not in the recording: it is graded from its component tree.

Ready to play0:00 / —:—
Open recording file

Drag the timeline to seek. Keyboard: arrow keys move 5 seconds; Home and End jump to the start and end.

Full graded browser recording: dashboard widget (before configuration, edit and Save, light and dark at 380 px, a second instance at 1180 px), sprint action (light and dark, sort, issue link, select, comment post, double click, forbidden post, explain, close), the not-started sprint and the Forge LLM cases, where the app has each. The UI Kit admin panel is not in the recording.

Screenshots

Pictures from one of the three scoring sites, the one on the run's own seed. The dashboard widget and the sprint action are captured from the built app in a browser during grading. The admin panel picture is drawn by the benchmark from the app's component tree, not by Jira, and nothing is graded from that picture. A test keeps the score of its worst site (the excellence measurements keep their mean), so a test can fail on a site these pictures do not show. A detail that opens "worst of …" names the seed of the site the test scored worst on, which can be the pictured one.

1Dashboard widget · light
2Dashboard widget · dark
3Sprint action · light
4Sprint action · dark
5Admin panel (UI Kit) · light: drawn by the benchmark from the app's component tree, not by Jira.

Run details

Model
deepseek/deepseek-v4.1-flash
Engine events
0
Repair rounds
0
Started
Oct 11, 2026, 12:39 AM
Finished
Oct 11, 2026, 01:39 AM

How the benchmark is scored