Forge Benchmark
One model builds a working Jira app on Atlassian Forge, offline; the scorer runs it and grades it.
Forge leaderboard
Forge results of one version, ranked together. Not comparable with the Gauntlet board: a different task and a different scorer.
Forge 1.0 has no published results yet.
Its scorer is still being finalised. Results are graded by the calibrated forge-1.0 scorer and posted from a Goose version that supports Forge; release-candidate scores (forge-1.0-rc) are not published here. The task, the modules and the grading are described below.
About the Forge benchmark
The task is the Scope Ledger. Agile coaches on a Jira Cloud site want to know what entered a sprint after it started, who added it and how many story points it carried: on a dashboard, from the sprint itself, and from Rovo. The app keeps a ledger of sprint changes from Jira's issue events, backfills and heals it with an hourly reconciliation, and reports team totals plus the change lists each person is allowed to see. It is a working app, not a tutorial sample, and nothing is deployed: the scorer installs it into a seeded Jira site and grades it by running it.
What the model receives
An empty Forge app (a manifest with the app id and the nodejs22.x runtime, no modules), the pinned Forge packages already installed, the Forge manifest schema and the Jira Cloud and Jira Software OpenAPI files as reference, and offline tools: Forge's own linter, a Forge runtime and Custom UI host, and a seeded development Jira site. There is no internet, nothing to install and no reference application.
The budget
150 model calls. The harness stops the session at the budget and scores whatever exists, so empty or partial work still gets a score. Polishing beyond the scored behaviours earns nothing.
What makes it hard
Events arrive duplicated, out of order or not at all, and the ledger must still hold exactly one row per change. Removed issues, two estimation fields whose ids change per seed, parallel active sprints, 429 rate limits with Retry-After, and people who must not see issues they cannot browse: not on screen, not in what the app sends Forge LLM, not in a Realtime payload. The LLM must answer through a forced tool, and a refusal or a malformed answer must not break the modal. The scoring site uses a different seed from the one the model develops against.
| Module | What it does in the app | Platform status (docs pages, 2026-10-02/03) |
|---|---|---|
| trigger | On avi:jira:updated:issue, hands relevant updates to a queue and returns. | GA |
| consumer | Async events queue consumer; performs the ledger writes. | GA |
| scheduledTrigger | Hourly backfill and reconciliation. | GA |
| dashboards:widget | Custom UI view plus edit: the dashboard widget and its board picker, saved through @forge/dashboards-bridge. The legacy jira:dashboardGadget earns nothing. | GA since 2026-09-22 |
| jira:sprintAction | Custom UI modal: the sprint's ledger table and comment action. | GA |
| action | Rovo action get-sprint-scope. | GA |
| rovo:skill | skills/sprint-scope-analyst/SKILL.md, depending on get-sprint-scope. | Preview since 2026-10-02 |
| rovo:agent | Lists the skill. | GA |
| rovo:mcp | One MCP server exposing get-sprint-scope. | Preview since 2026-10-01 |
| llm | Forge LLM (@forge/llm): the sprint action's explanation, answered through one forced tool, report_scope, from an active model. | No Preview label on its docs page |
| Realtime | The widget subscribes and shows the new numbers without a reload; payloads carry sprint ids only. | No Preview label on its docs page |
Storage is Forge KVS: ledger changes live in a custom entity indexed by sprint (partition) and change time (range) and are read back through that index. Every surface is Custom UI, styled with Atlassian design tokens in light and dark inside the default Content Security Policy; UI Kit earns nothing. The app requests only the OAuth scopes its calls need.
- Lint. Forge's own client-side linter runs offline, twice; the two results must match. A lint error caps the score at 0.499 and is a critical defect; each distinct warning costs points and never caps. The scorer adds its own rules for what the linter misses.
- Bundle and load. Every manifest function is bundled the way forge deploy does and loaded in Atlassian's Forge runtime wrapper, pinned by checksum, each invocation in a fresh Node process with the platform's timeouts.
- Run the backend. A fresh seeded Jira site behaves like Jira Cloud: REST v3 and Jira Software, pagination, permissions and scripted 429s. The first scheduled run backfills; then a script of issue updates with duplicated, reordered and dropped deliveries; a second scheduled run heals; a third has nothing new to write.
- Rovo. The action runs as two people for each active sprint, plus an unknown and a missing sprint id.
- Every surface in a browser. The widget without configuration, the edit picker and the host's Save, the view in light and dark, a second widget showing another board, and the widget updating live through Realtime; the sprint action sorted, linked, selected, posted twice by double click, posted on a forbidden issue, asked for its Forge LLM explanation (answered by a scripted model) and closed; a sprint that has not started.
- Video and screenshots. The whole browser session is recorded, including the live update and the explanation, and every surface is captured in light and dark, plus a contact sheet. The DOM is what is graded; pixels decide only whether a surface is blank and whether dark mode is really dark.
- Deploy readiness. The manifest, the permissions and the code are judged against what real Forge would refuse or fail on, each rule tied to the platform page it encodes. A finding another check already grades costs nothing twice.
No model takes part in scoring: every number is computed against an oracle on the run's own seeded data.
Earned score = (0.88 × the weighted inner tiers + 0.12 × excellence gate × E) × the critical multiplier. The final score is the lower of the earned score and the cap that applies. Meeting a cap's conditions adds no points.
| Tier | What it measures | Share of the score |
|---|---|---|
| LLint and bundles | Forge's own linter reports no errors (warnings cost points), every manifest function bundles and loads, packages resolve at the kit's pins, the manifest rules the linter misses, and only the scopes the app's calls need. | 7.04% |
| KPlatform currency | Current modules and APIs: dashboards:widget with its edit API, the rovo:skill and rovo:mcp wiring, a Forge LLM model the platform lists as active, no /rest/api/3/search, no @forge/api storage, a nodejs22.x or nodejs24.x runtime, the consumer shape, the declared KVS entity, and deploy readiness: a manifest and permissions real Forge would accept and no code predicted to fail there. | 8.8% |
| TEvent pipeline | Issue updates flow trigger to queue to consumer: one ledger row per change under duplicate, reordered and dropped deliveries, multi-sprint changelog values, re-estimates followed, Retry-After honoured, no user context in background work. | 14.08% |
| RReconcile | The first scheduled run backfills every change since each active sprint started, removals included; later runs heal what the event stream missed, page through results in each endpoint's own style, wait out rate limits and finish inside the module timeout. | 12.32% |
| SStorage | Forge KVS with the storage scope, ledger reads through the declared entity index in change-time order, no KVS limit errors. | 7.04% |
| BResolvers and permissions | Every invoked resolver exists and answers structured errors, nobody sees changes to issues they cannot browse (not on screen, not in an LLM prompt, not in a Realtime payload), the hidden-change count is right, and each click posts exactly one ADF comment as the viewer. | 10.56% |
| UUI function | The dashboard widget's numbers and chart, its board choice saved through the host, live updates through Forge Realtime, the sprint action's ledger table, sorting, issue links, comment flow, the Forge LLM explanation, close and not-started states. | 14.08% |
| VVisual | Atlassian design tokens with 4.5:1 contrast in light and dark, a painted surface, no Content Security Policy violations or console errors, and the widget readable at 380 px wide. | 7.04% |
| ARovo | The get-sprint-scope action returns exact numbers and the visible changes for the invoking person, returns errors instead of throwing, and the skill's SKILL.md tells the agent how to use it. | 7.04% |
| EExcellence | Few Jira requests in the backfill and per event, few round trips before each surface paints, and a scheduled run with nothing new writes nothing — paid only in proportion to the excellence gate. | 12% |
Score caps
As the task text states them to the model:
| Maximum | Applies when |
|---|---|
| 0.499 | Lint errors, or manifest functions that do not bundle and load. |
| 0.699 | No working ledger: the widget shows no data, backfill or event rows are wrong, or the KVS entity is unused. |
| 0.799 | Missing current-platform surfaces (the dashboards widget and its edit API, the Rovo skill, the sprint action table, the Rovo action), or a surface broken in light or dark. |
| 0.899 − 0.03 per further defect, never below 0.799 | Any duplicate, ordering, rate-limit, pagination, permission, policy, console, LLM or Realtime defect: 0.899 for one, 0.03 lower for each further one. |
Critical defects
Each critical defect multiplies the earned score by a factor between 0.6 and 1 (factor = 0.6 + 0.4 × severity). A permission leak therefore scores below a missing chart, and a duplicate comment below a missing one: score_forge.py's severity self-test refuses to run if that order ever inverts.
l_deployabledeploy blocked — lint errorsl_bundles_loadcrash — the app does not runt_no_double_countwrong numbers — a change counted twicer_backfill_completedata loss — changes silently missingb_no_permission_leakdata leak — hidden issues shown to a personb_comment_exactly_onceduplicate side effect on a customer's Jirau_widget_loadsdead primary flow — no data on the dashboard
Excellence
The last 12% rewards few Jira requests in the backfill and per event, few round trips before each surface paints, and a no-change scheduled run that writes nothing. It is paid in proportion to a gate: the event rows, the healing, the widget numbers, the ledger table, the comment and every platform-currency check at full marks, with a clean console, a clean security policy and no lint warnings.
Scoring rules synced at goose 9bf49184a; the Forge 1.0 scorer is still being finalised — results open when it is frozen.
# spec-build-forge.md
# forge-1.0 — Scope Ledger (Atlassian Forge)
Build the Scope Ledger Forge app in this workspace: a working Jira app, not a tutorial sample.
`FORGE-CONTRACT.md` is the contract. `STARTER.md` describes the workspace, the installed
packages, the reference material and the offline dev tools (Forge runtime, dev Jira site,
linter). There is no internet and nothing to install. No reference application is provided.
You have a budget of 150 model calls. The harness stops the session at the budget and scores
whatever exists; plan to implement, test, and finish well inside it — polishing beyond the
scored behaviours earns nothing.
## Definition of done
Done means each of these works when the harness runs the app on its own seeded site; nothing
else is scored:
1. `npm run lint` reports no errors and no warnings; every manifest function bundles and loads.
2. The current modules of §2, with scopes your calls need and nothing more.
3. The first scheduled run backfills every change of every active sprint, through paginated
search and rate limits (§1, §3).
4. Issue updates flow trigger → queue → consumer into the KVS entity: exactly one row per change
under duplicates, reordering and loss; estimate changes move the numbers (§3).
5. Later scheduled runs heal what the event stream missed and change nothing otherwise (§3).
6. The widget: board choice through the dashboards edit API, correct numbers, chart, live
updates through Forge Realtime (§4).
7. The sprint action: ledger table, sorting, router links, hidden count, one ADF comment per
click as the viewer, flags, close, and the Forge LLM explanation (§5).
8. Nobody sees changes to issues they cannot browse (§1, §5, §6).
9. The Rovo action answers exactly; the skill, agent and MCP server are wired and valid (§2, §6).
10. Every surface works in light and dark, inside the Custom UI security policy, with a clean
console (§7).
## Score bands
Tests earn the score. Conditions also cap it:
- Lint errors, or manifest functions that do not bundle and load: maximum 0.499.
- No working ledger (widget shows no data, backfill or event rows wrong, KVS entity unused):
maximum 0.699.
- Missing current-platform surfaces (dashboards widget and its edit API, Rovo skill, sprint
action table, Rovo action) or a surface broken in light or dark: maximum 0.799.
- Any duplicate, ordering, rate-limit, pagination, permission, policy, console, LLM or realtime
defect: maximum 0.899, minus 0.03 for each further such defect (never below 0.799).
Lint warnings cost points but never cap. A small excellence share rewards few Jira requests and
rendering each surface in few round trips. The harness keeps screenshots of every surface.
## Completion handoff
Work only with this workspace, its packages, the reference material and the dev tools. When the
definition of done holds, stop your temporary processes and return a final summary of files,
checks and remaining limitations; that response hands the app to the external scorer, which
starts after your session ends. Do not wait for or poll grader results.
# FORGE-CONTRACT.md
# Scope Ledger — Forge app contract (forge-1.0)
Agile coaches on a Jira Cloud site want to know what entered a sprint after it started, who added
it and how many story points it carried — on a dashboard, from the sprint itself, and from Rovo.
Build the Atlassian Forge app in this workspace. The harness installs it into a seeded Jira site
(scrum boards, parallel active sprints, a few hundred issues, people with different permissions)
and grades it by running it: product events, queue deliveries, scheduled runs, resolver calls,
the Rovo action, and every Custom UI surface in a browser, light and dark. Nothing is deployed.
## 1. The numbers
For each **active** sprint S, with `startDate` from the Jira Software sprint API:
- A **change** is a Sprint-field changelog entry created after `startDate` that puts an issue
into S (`added`) or takes it out of S (`removed`). Its id is the changelog id, its time the
changelog `created`, its author the changelog author. A change is keyed by changelog id +
sprint: one entry moving an issue between two sprints is a `removed` in one and an `added` in
the other.
- **committed** = sum of current estimates of issues that were in S at `startDate`.
- **added** = sum of current estimates of issues in S now that were not in S at `startDate`.
- **removed** = sum of current estimates of issues in S at some time after `startDate` and not in
S now.
- **creep** = `100 × added / committed`, rounded half away from zero to one decimal and always
written with that decimal and `%` (`20.0%`); `—` when committed is 0.
- An issue's **estimate** for S is the value of the estimation field of S's board (the sprint's
`originBoardId`); no value counts as 0. "After `startDate`" is strictly after.
- Points are written as plain decimals (`34.5`, `0`), no thousands separators.
Background work (triggers, consumer, scheduled job) sees every issue. What a **person** sees in
the sprint action and the Rovo action lists only changes to issues that person can browse, plus
a count of the changes hidden from them; the totals above are team totals and are the same for
everyone.
## 2. Modules
| module | requirement |
|---|---|
| `trigger` | on `avi:jira:updated:issue`; hands relevant work to a queue and returns |
| `consumer` | Forge async events queue consumer; performs the ledger writes |
| `scheduledTrigger` | `interval: hour`; backfill and reconciliation (§3) |
| `dashboards:widget` | Custom UI view + `edit` (§4); the legacy `jira:dashboardGadget` earns nothing |
| `jira:sprintAction` | Custom UI modal (§5) |
| `action` | Rovo action `get-sprint-scope` (§6) |
| `rovo:skill` | `skills/sprint-scope-analyst/`, depends on `get-sprint-scope` (§6) |
| `rovo:agent` | lists the skill |
| `rovo:mcp` | one module, `name` at most 30 characters, exposing `get-sprint-scope` |
| `llm` | Forge LLM (`model: [claude]`), for the sprint action's explanation (§5) |
Resolvers use `@forge/resolver`. Storage is Forge KVS: ledger changes live in a **custom entity
indexed by sprint (partition) and change time (range)**, and are read back through that index.
Request only the scopes your calls need: per call, the OAuth2 scopes the shipped OpenAPI lists for
it (the classic scope where one exists, else the whole granular set). The linter does not see every
call. Custom UI only; UI Kit (`render: native`) earns nothing.
## 3. Backend behaviour
- The site existed before the app was installed. The scheduled job's first run backfills every
change since each active sprint started, including issues that have since left every sprint.
- Issue updates keep the ledger current: sprint changes add ledger rows; estimate changes move
the numbers. Updates that touch neither do no Jira or queue work (KVS reads are fine). The
harness runs the scheduled job once before it delivers any update.
- Product events can arrive more than once and out of order, and some never arrive. The ledger
holds **exactly one row per change**, whatever the delivery history. Each row records `source`:
`event` or `reconcile`, the path that recorded it first (event work may also record other
changes of the issue it reads). Later scheduled runs record what the event stream missed; a run
with nothing new writes nothing.
- Jira may answer `429` with `Retry-After`: wait at least that long before the next attempt.
- Background work uses `asApp()`. Whatever shows a person issue data (keys, authors, change
lists) shows only what that person can browse — read as them (`asUser()`), or as the app with
an explicit permission check for them. Comments are posted as the person. Team totals come
from the ledger.
## 4. Dashboard widget
**Edit** (`edit.resource`): one element per scrum board, `[data-testid="board-option"]` with
`data-board-id`, clickable; the selected one carries `aria-pressed="true"`. The choice reaches the
dashboard through the dashboards widget edit API (`@forge/dashboards-bridge`). The dashboard's
own Save calls your `onProductSave` handler with the last `updateConfig` value (the stored config
if none was sent) and stores what it returns (`null` stores nothing); with no handler registered
it stores the last `updateConfig` value. Reopening edit shows the
stored board selected.
**View** (`resource`), root `[data-testid="scope-widget"]`. It takes its board from the widget
configuration in its context, so two widgets on one dashboard can show different boards:
- No stored board: `[data-testid="needs-config"]`, nothing else.
- Otherwise one `[data-testid="sprint"][data-sprint-id="<id>"]` per active sprint of that board,
ordered by `startDate` (ties by sprint id), each holding `[data-metric="committed"]`, `[data-metric="added"]`,
`[data-metric="removed"]` and `[data-metric="creep"]` whose text is the §1 number.
- A bar chart: one `<svg data-testid="chart">`; per sprint and per series one
`<rect data-sprint-id data-series="committed|added|removed">`; every bar on one linear scale
from 0 (rendered height proportional to its number within 1 px).
- At 380 px wide nothing scrolls horizontally and no number is clipped or truncated (long
names may end in an ellipsis).
- Live: after ledger rows are written, an open widget shows the new numbers without a reload,
through Forge Realtime (`@forge/realtime` in the backend, the bridge's realtime subscribe in the
widget) — no polling. Realtime payloads carry sprint ids only.
## 5. Sprint action (modal)
The sprint comes from the module context; the harness opens active and future sprints. A sprint
that has not started shows `[data-testid="not-started"]` (and optionally close) and nothing else.
Otherwise:
- Team totals as in §4 (`[data-metric]`, same four), and `[data-testid="hidden-count"]` = number
of this sprint's changes hidden from the viewer.
- `table[data-testid="ledger"]`, headers `th[data-col]` for `issue`, `points`, `kind`, `by`,
`at`, `source`; one `tr[data-change-id="<changelog id>"]` per visible change with
`td[data-col=…]` cells: issue key, current estimate, `added`/`removed`, author display name, a
`<time datetime>` holding an ISO-8601 instant with offset (compared as instants),
`event`/`reconcile`.
- Default order: `at` ascending, equal times by changelog id ascending (ids compare as numbers).
Clicking `th[data-col="at"]` toggles between that order and its exact reverse, starting with
ascending when another sort was active. The active header carries `aria-sort`.
- The issue key opens the issue (`/browse/<KEY>`) through the Forge router.
- Clicking a row selects it (`aria-selected="true"`). `[data-testid="post-summary"]` posts one
comment on the selected change's issue, authored by the viewer, in Atlassian Document Format,
naming the issue key, the sprint name and the sprint's creep; then a success flag. One click,
or a double click, posts exactly one comment. A failure shows an error flag and leaves the
modal working.
- `[data-testid="explain"]` asks Forge LLM (`@forge/llm`) to explain the sprint's creep, using a
model that `list()` reports `active`. The model must answer through one tool, `report_scope`,
arguments `{ "summary": string, "changeIds": string[] }` (force it with `tool_choice`); send it
only what the viewer may see. `[data-testid="explanation"]` shows the summary only if it holds
no digits (otherwise your own sentence with the ledger's numbers), plus one `[data-change-id]`
element per returned id that is a visible change of this sprint (others dropped). A refusal (no
tool call), malformed arguments or an LLM error shows an error flag and leaves the modal working.
- `[data-testid="close"]` closes the modal.
## 6. Rovo
`action` key `get-sprint-scope`, `actionVerb: GET`, one required string input `sprintId`. For
the invoking person it returns a JSON object:
```
{ "sprintId": "41", "sprintName": "…", "committed": 34, "added": 8, "removed": 3,
"creepPercent": 23.5, "hiddenChanges": 1,
"changes": [ { "changeId": "…", "issueKey": "OPS-12", "kind": "added", "points": 5,
"at": "<ISO-8601 UTC>", "by": "<display name>" } ] }
```
`creepPercent` is rounded as creep (§1) and `null` when committed is 0; `changeId` is the
changelog id; `changes` are the visible ones in table order (§5); `at` is compared as an instant.
An unknown or missing `sprintId` returns `{ "error": "<message>" }` and does not throw.
`skills/sprint-scope-analyst/SKILL.md`: YAML frontmatter `name` equal to the directory name
(1–64 characters: lowercase letters, digits, single hyphens, no leading or trailing hyphen),
`description` (50–1,024 characters: what it does and when to use it), `allowed-tools` (space
separated, including `get-sprint-scope`); then at most 500 lines of Markdown telling the agent
when and how to call `get-sprint-scope`, what `sprintId` is, how to read the result, and what to do
with an error.
## 7. Custom UI
Every surface calls `view.theme.enable()` and is styled with Atlassian design tokens
(`var(--ds-…)`): text with `--ds-text*` (links may use `--ds-link*`), contrast at least 4.5:1 in
light and in dark (disabled controls exempt). The page behind your surface is unpainted: paint your own background with a
`--ds-surface*` token. The browser console stays free of errors. Surfaces render inside the
default Forge Custom UI content security policy: inline `<script>` in `index.html` only as static content (the
platform hashes it), no inline event handlers or `eval`, no `<style>` elements or `style` attributes in markup, no external scripts, styles or fonts, assets referenced relatively.
Declaring `unsafe-inline` in `permissions.content.styles` is allowed; script relaxations count as
unneeded permissions.
## 8. What the harness does differently from production
- Resources are served exactly as committed (build Custom UI yourself). Backend source is bundled
from `src/` the way `forge deploy` does. Every function invocation runs in a fresh Node process
of the Forge runtime, with the platform's timeouts.
- Trigger `filter.expression` is not evaluated: the handler receives every issue-updated event.
- A consumer that throws or times out is redelivered after 1, 2, 4 and 8 minutes, then every 15,
for 24 hours; a retry request (`InvocationError`) is redelivered after its `retryAfter`. The
harness clock runs with real time and jumps over these redelivery waits; a wait inside an
invocation is real time.
- The scoring site uses a different seed than the dev site: ids, keys, custom field ids, users,
sprint names and dates all differ.
- The widget runs at the dashboard layout the harness chooses; the sprint action in a modal.
# STARTER.md
# Starter and offline tools
There is no internet in this workspace. Everything you can use is already here.
**Workspace.** `manifest.yml` holds the app id and runtime and no modules. `package.json` pins the
installed packages; `node_modules/` already contains them and nothing else can be installed.
`src/index.js` is empty; `static/` and `skills/` are empty. Every file is editable.
**Installed packages** (their typings and sources are your API reference): `@forge/api` 8.2.0,
`@forge/kvs` 2.0.7, `@forge/events` 3.0.7, `@forge/resolver` 2.0.0, `@forge/bridge` 7.1.0,
`@forge/dashboards-bridge` 2.0.0, `@forge/hooks` 2.0.0, `@forge/llm` 1.0.7, `@forge/realtime`
1.0.1, `react` and `react-dom` 18.3.1, `esbuild`
0.28.2.
**Reference material** under `$FORGE_KIT` (read-only): `schema/manifest-schema.json` (the Forge
manifest schema), `openapi/jira.json` and `openapi/jsw.json` (Jira Cloud REST and
Jira Software REST; the other files in `openapi/` are the linter's offline copies). These are single-line JSON files of several MB: query them with `node` or
`grep -o`, never print them whole.
**Dev site.** A seeded Jira Cloud site answers at `$FORGE_SITE_URL` for the dev tools below. It
behaves like Jira Cloud — REST v3 and Jira Software REST, including the bulk endpoints, their
pagination, errors and rate limits; dates are ISO-8601 strings as the OpenAPI types them. JQL:
fields `project`, `key`, `sprint`, `updated`, `created`, `status`, `statusCategory`, `issuetype`,
`labels`, `assignee`, `reporter`, `cf[id]`; operators `= != in not in > >= < <= is is not ~`,
`AND OR NOT`, parentheses, `ORDER BY`; functions `openSprints()`, `closedSprints()`,
`futureSprints()`, `currentUser()`, `now()`, `startOfDay()`; relative dates like `-14d`. Forge LLM answers with a
scripted model on the dev site. You
never call the site directly; your app reaches it through the Forge runtime.
**Tools** (`npm run lint` and `node $FORGE_KIT/bin/forge-dev.cjs <command>`):
| command | does |
|---|---|
| `npm run lint` | Forge's own client-side linter, offline. It is staged: fix and rerun until it reports no errors. |
| `invoke <functionKey> [--module <key>] [--resolver <key>] [--payload <file>] [--as <accountId>]` | runs a manifest function in the Forge runtime against the dev site with the event shape of the module that references it (`--resolver` calls that resolver key with `--payload`); prints the result, logs and every Jira, KVS and queue call. `--as` makes the invocation user-led. |
| `events [--limit N]` | delivers the dev site's next N issue updates to your trigger and drains the queues |
| `scheduled <moduleKey>` | runs a scheduled trigger once, then drains the queues |
| `serve <moduleKey> [--edit] [--sprint <id>] [--config <json>] [--theme light\|dark] [--as <accountId>]` | serves a Custom UI module with the Forge bridge and the dashboard host emulated, prints its URL and runs until stopped (start it in the background). In a served edit surface, `window.__forgeHost.save()` performs the dashboard's Save. |
| `kvs` / `users` / `reset` | dump stored keys and entities / list dev users / clear dev storage and queues and rewind the dev site's update stream |
Screenshot a served surface with the bundled browser (`BROWSER-TESTING.md`). Build Custom UI into
`static/<name>/build/` (an `index.html` plus assets) with the installed esbuild; the harness never
builds it for you.
Public input source: goose commit 9bf49184a2acc7ca360e3af13a0dbc327e672c8c.