Cloud baseline

xiaomi/mimo-v2.6-flash, single model via OpenRouter — 0.8461 on Forge 2.0

xiaomi/mimo-v2.6-flash · scorer forge-2.0

xiaomi/mimo-v2.6-flash · single agentOct 10, 2026
Overall
0.8461
Benchmark
Forge 2.0
Model build
115m 59s

Tier breakdown

Each bar is the tier's mean, from 0 to 1 (for E, its mean × the excellence gate). The percentage is the tier's share of the score. Mean × share, added over all tiers, is the tests earned figure below.

L · Lint and bundles1.001.76%
K · Platform currency1.002.2%
T · Event pipeline0.97923.52%
R · Reconcile1.003.08%
S · Storage1.001.76%
B · Resolvers and permissions0.97222.64%
U · UI function0.99773.52%
V · Visual1.001.76%
A · Rovo0.901.76%
E · Excellence0.97943%
R1 · Migration1.0012%
R2 · Dosing0.853513%
R3 · Time limits1.008%
R4 · Changing world1.0010%
R5 · Admin panel1.0010%
R6 · CI web trigger1.008%
R7 · Custom field0.97425%
R8 · Forge LLM1.004%
R9 · Boot budget1.005%

Scoring detail

How this score was built

Per-test scorer output, as posted by the app. Expand a section to inspect its tests.

The rules behind these numbers: Forge 2.0 requirements, weights and score caps.

The score, step by step

0.9757 tests earned × 1.0000 critical defects × 0.8671 reliability = 0.8461 final score

tests earned = 0.25 × 0.9843 v1 + 0.7297 v2 (out of 0.75)

v1 = 0.88 × 0.9850 tiers L–A + 0.12 × 0.9893 gate × 0.9899 excellence

Reliability: every tier in the breakdown above, except Excellence, is a group of tests. Each group of tests multiplies the score by 1 − 0.10 × the shortfall of its worst test: by 0.90 when that test fails completely, by 0.95 when it scores 0.5.

  • T Event pipeline: worst test t_reestimate_followed, score 0.8333 × 0.9833
  • B Resolvers and permissions: worst test b_comment_adf_as_user, score 0.8333 × 0.9833
  • U UI function: worst test u_ledger_table, score 0.9748 × 0.9975
  • A Rovo: worst test a_action_result, score 0.6 × 0.9600
  • R2 Dosing: worst test r2_rate_limit_reaction, score 0.4138 × 0.9414
  • R7 Custom field: worst test r7_values_fresh, score 0.9484 × 0.9948

v1 is this run's score on the tests carried over from Forge 1.0: 25% of the total. v2 is the requirements R1–R9, each at its share of the score: the other 75%. Critical defects multiply the whole score, and so does each group of tests, by 1 − 0.10 × the shortfall of its worst test.

Excellence gate 0.9893: the average of the 18 conditions below, each counted at the value shown. Green is met in full; red is not, and lowers the gate. The 3 excellence tests are paid in proportion to the gate; together they are 3% of the total.

t_event_rows 1.00r_heal_dropped 1.00u_widget_numbers 1.00u_ledger_table 0.9748b_comment_adf_as_user 0.8333k_consumer_shape 1.00k_current_apis 1.00k_dashboard_widget 1.00k_entity_declared 1.00k_llm_model_current 1.00k_manifest_semantics 1.00k_rovo_mcp 1.00k_rovo_skill 1.00k_runtime_risks 1.00k_widget_edit_bridge 1.00v_console_clean 1.00v_csp_clean 1.00l_lint_warnings 1.00

Earned score 0.8461 · No cap applies · Final score 0.8461

Under a cap the final score is min(earned, cap − 0.05 × (1 − earned)): always below the cap, and closer to it the more the run earned. With no cap it is the earned score. Avoiding a cap adds no points.

On a site of about 1,000 issues, background work stays inside its share of the 2,400-point hour, no person-facing request meets a quota 429, and Retry-After and RateLimit-Reason (quota, burst, per-issue write) decide every retry; the backfill still completes.

r2 background share

1.00

10 of 10 hours kept background within 70 % (1680 points; the contract default; the probe recorded no change); busiest hour 255; dosing optimum for the whole run 249

r2 person never quota refused

1.00

0 of 271 person-facing requests refused for quota

r2 rate limit reaction

0.4138

worst of 3 scoring sites (seed 1969058b8b32be42): 12 of 29 429s answered by their reason (quota: all background paused; burst: the endpoint; per-issue: the issue); broken: jira-quota-tenant-based at 1795005153234 ms: 17

r2 backfill complete

1.00

135 of 135 pre-upgrade changes v1 never recorded are in scope-ledger, right

Graded browser recording

Loading duration…

Watch the full graded browser recording at its original speed: the dashboard widget, the sprint action, the not-started sprint and the Forge LLM cases, as the scorer drove them on the run's own seed, the first of the three scoring sites. The test results and screenshots on this page provide the wider evidence. Pause or scrub to inspect a view.

The UI Kit admin panel is not in the recording: it is graded from its component tree.

Ready to play0:00 / —:—
Open recording file

Drag the timeline to seek. Keyboard: arrow keys move 5 seconds; Home and End jump to the start and end.

Full graded browser recording: dashboard widget (before configuration, edit and Save, light and dark at 380 px, a second instance at 1180 px), sprint action (light and dark, sort, issue link, select, comment post, double click, forbidden post, explain, close), the not-started sprint and the Forge LLM cases, where the app has each. The UI Kit admin panel is not in the recording.

Screenshots

Pictures from one of the three scoring sites, the one on the run's own seed. The dashboard widget and the sprint action are captured from the built app in a browser during grading. The admin panel picture is drawn by the benchmark from the app's component tree, not by Jira, and nothing is graded from that picture. A test keeps the score of its worst site (the excellence measurements keep their mean), so a test can fail on a site these pictures do not show. A detail that opens "worst of …" names the seed of the site the test scored worst on, which can be the pictured one.

1Dashboard widget · light
2Dashboard widget · dark
3Sprint action · light
4Sprint action · dark
5Admin panel (UI Kit) · light: drawn by the benchmark from the app's component tree, not by Jira.

Run details

Model
xiaomi/mimo-v2.6-flash
Engine events
0
Repair rounds
0
Started
Oct 10, 2026, 09:20 PM
Finished
Oct 10, 2026, 11:34 PM

How the benchmark is scored