Cloud baseline

openai/gpt-6.1-sol, single model via OpenRouter — 0.8990 on Forge 1.0

openai/gpt-6.1-sol · scorer forge-1.0

openai/gpt-6.1-sol · single agentOct 3, 2026
Overall
0.8990
Excellent
no
Benchmark
Forge 1.0
Model build
13m 12s

Tier breakdown

L · Lint and bundles1.00
K · Platform currency1.00
T · Event pipeline1.00
R · Reconcile1.00
S · Storage1.00
B · Resolvers and permissions1.00
U · UI function0.99
V · Visual1.00
A · Rovo1.00
E · Excellence1.00

Scoring detail

How this score was built

Per-check scorer output, as posted by the app. Expand a section to inspect its checks.

Earned credit before admission

(0.88 × 0.9978 core + 0.12 × 1.00 gate × 1.00 excellence) × 1.0000 critical = 0.9981

The core (88% of the total) is the weighted mean of the nine measured tiers. The last 12% is the excellence slice: it unlocks in proportion to the perfection conditions below (18 of 18 met here), then pays out at the excellence tier's own measured mean. Core 0.8781 + excellence 0.1200.

t_event_rows 1.00r_heal_dropped 1.00u_widget_numbers 1.00u_ledger_table 1.00b_comment_adf_as_user 1.00k_consumer_shape 1.00k_current_apis 1.00k_dashboard_widget 1.00k_entity_declared 1.00k_llm_model_current 1.00k_manifest_semantics 1.00k_rovo_mcp 1.00k_rovo_skill 1.00k_runtime_risks 1.00k_widget_edit_bridge 1.00v_console_clean 1.00v_csp_clean 1.00l_lint_warnings 1.00

Earned score 0.998 · Cap 0.899 · Final score 0.899

The final score is the lower of the earned score and the cap. Meeting a cap's conditions adds no points.

  • production robustness: maximum 0.899u_llm_explain

Graded browser recording

Loading duration…

Watch the full graded browser recording at its original speed: the payment field, structure inspection, committed updates and replay. The check results and screenshots on this page provide the wider evidence. Pause or scrub to inspect a view.

Each tower represents a payment: height encodes its amount, cap color its status, and the collar its currency. The selected detailed tower is expected to animate its collar after a committed backend update or replay. These are the task requirements; the check results show what this app actually achieved.

Ready to play0:00 / —:—
Open recording file

Drag the timeline to seek. Keyboard: arrow keys move 5 seconds; Home and End jump to the start and end.

Full graded browser recording: dashboard widget (no config, edit + Save, light and dark, 380 and 1180 px, a second instance), sprint action modal (sort, router, select, comment post through the 429 retry, double click, forbidden post, close) and the not-started sprint
Recording integrity

SHA-256: dd3ab3ed309ca7dc254c4c63a1dac53e5f1c3cfdc0a191f339b2b9e26e87af29

Screenshots

Captured from the built application during browser grading. Captions identify the recorded views and checks; these images do not represent repair rounds.

1Every captured surface
2Widget edit view
3Widget before configuration
4Sprint not started
5Sprint action · dark

Run details

Model
openai/gpt-6.1-sol
Engine events
0
Repair rounds
0
Started
Oct 3, 2026, 02:17 PM
Finished
Oct 3, 2026, 02:35 PM