Qwen3.8 Flash vs Max Prime: the cheap one beat the flagship on our Forge benchmark
Mihai Perdum
Author
38 min readOctober 8, 2026
Key takeaways
Qwen3.8 Flash scored 0.7961 on Forge 1.0, an Atlassian Forge app built with no internet, and Qwen3.8 Max Prime scored 0.6953. On Gauntlet 7.2, a payments web app, the order flips: Max Prime 0.8364, Flash 0.6846.
On Forge the two earned almost the same work before the ceilings, 0.9426 and 0.9269. One manifest line decided it: Max Prime gave a storage index two range attributes where Forge allows exactly one. Atlassian's server-side check, reached through the current forge lint and forge deploy, rejects that manifest; forge lint passes Flash's.
Asked afterwards how many attributes a range may hold, both models mostly said one (8 of 8 and 6 of 8). Asked the same about the partition, which Atlassian says can hold several, they said one again (7 of 8 and 7 of 7). Given the benchmark contract's sentence, both wrote two ordering attributes (7 of 8 and 7 of 7). Given Atlassian's docs sentence as well, both kept to one range attribute, 8 of 8, though only 7 of those 16 answers were valid Forge manifest YAML.
Max Prime still built the more polished Forge app and won 7 of the 12 checks where the two differ, including the Forge LLM explain feature, 23 of 23 assertions. Flash's explain step made no LLM call on any of the three scoring sites.
On Gauntlet Max Prime won 13 checks to Flash's 2. Flash's live stream never applied, its inspector framed a tower in 0 of 8 poses, and an approved payment never reached its table. Flash won only the two crash-recovery checks.
Flash cost $0.47 for its Forge run and $0.59 for Gauntlet; Max Prime cost $16.46 and $16.06. When we fine-tuned the open Qwen3.8-27B, Atlassian multiple choice on our home probe fell from 72.86 at step 506 to 64.32 at step 3,541, with 79% of the tokens agent work, while a general-text drift check stayed near 0.006.
Qwen3.8 Flash vs Max Prime is the oddest result inside one model family on our boards. The far cheaper Flash beat Qwen3.8 Max Prime on Forge 1.0, an Atlassian Forge app built with no internet, 0.7961 to 0.6953. Then it lost to Max Prime on Gauntlet 7.2, a payments web app, 0.6846 to 0.8364. Flash's Forge run cost 47 cents. Max Prime's cost $16.46.
Max Prime is sold as a faster version of Qwen3.8 Max, the 2.4-trillion-parameter flagship. Flash has a 125-billion-parameter main model and uses 6 billion per token. Same agent, goose, though on two builds, 3.0.100 for Flash and 3.0.107 for Max Prime. Same 150-call budget, same Alibaba host behind OpenRouter, runs a day apart. On Forge the two earned almost the same work before a ceiling cut Max Prime down. The ceiling came from one line in its manifest: a storage index with two range attributes, where Forge allows exactly one. Afterwards I asked both models about the rule, 100 calls for about 60 cents. Both mostly said a range takes one attribute, then said "one" for the partition too, where Atlassian's answer is several. Given only the contract, they wrote two ordering attributes in 14 of 15 answers. Once Atlassian's sentence about range was in the prompt, both kept to one range attribute, 8 times out of 8.
Neither model showed it knew the rule. Handed the sentence, both kept to it.
On a payments app the internet's habit is usually the right answer, so a missing platform rule never shows up in the score. On Forge one missing sentence cost a band. Below is every check behind both results, the probe, what Qwen says about the two models and what it doesn't, the other 28 models that ran both boards, and what fine-tuning our own Qwen3.8-27B taught me about a model forgetting what it knew. This is the sequel to Claude Sonnet 5.5 vs Opus 5.5, which explains how both benchmarks work and lists every Forge trap. I don't repeat that here.
Qwen3.8 Flash vs Max Prime benchmark scores on both boards
These are the boards as I read them on Thursday 8 October, 08:08 EEST. Max Prime's Forge run was re-scored at 07:38 the morning before and moved from 0.6954 to 0.6953. The Gauntlet board and the Forge board are the current word, not this table.
model
Gauntlet 7.2
rank of 31
Forge 1.0
rank of 30
Qwen3.8 Max Prime
0.8364
2
0.6953
8
Qwen3.8 Flash
0.6846
12
0.7961
6
GLM 5.3 Prime
0.6797
13
0.4736
13
Qwen3.8 Omni Flash
0.4269
21
0.5529
10
Qwen3.8 27B
0.0915
25
0.1582
20
6 rows × 5 columnsHeader row enabled
Every run is a single model through OpenRouter, driven by goose, our fork of the open-source agent Block started, with a 150-call budget and no human help. The names are OpenRouter's: qwen/qwen3.8-flash, qwen/qwen3.8-max-prime, qwen/qwen3.8-omni-flash and qwen/qwen3.8-27b. The last is the untouched open model our own fine-tune starts from. GLM 5.3 Prime is there as a neighbour: it sits a hair under Flash on Gauntlet.
One caveat, said once and meant for the whole piece. Each model has one run per board. One run can't separate a model from its luck, and a single different decision early in a Forge build can move a run a whole band. Read everything below as "this is what happened in these runs, and here is why the scorer says so". It isn't a law about the models.
What we found in the two Qwen runs, a probe and our own training
On Forge the work was nearly a tie. Before the ceilings, Flash earned 0.9426 and Max Prime 0.9269. The 0.1008 gap on the board is the ceiling: 0.799 for Flash, 0.699 for Max Prime.
Max Prime's ceiling came from one line. It gave its ledger index two range attributes, at and changeId. Atlassian's docs say the range "can only have one attribute", and the contract the model read named one.
Flash wanted the same tie-breaker and packed it into the one attribute Forge allows: a zero-padded timestamp, a #, and the change id, in a single string.
The rule isn't in the linter the model could run. Our offline kit pins the same lint and manifest packages the current Forge CLI ships. The check that rejects the index lives on Atlassian's servers, which the sandbox can't reach.
Asked directly, both models mostly said a range takes one attribute, and said the same, wrongly, about the partition. Asked to build the index from the contract's sentence, they wrote two ordering attributes in 14 of 15 answers. Handed Atlassian's sentence, all 16 answers kept to one range attribute.
Max Prime still built the better-looking app, and it passed the Forge LLM explain feature (23 of 23 assertions) and the live Realtime widget, which Flash didn't.
On normal programming Max Prime was the better programmer in this run. It won 13 of the 15 Gauntlet checks where the two differ, and Flash won only the two crash-recovery checks.
Max Prime cost $16.06 and $16.46 a run. Flash cost $0.59 and $0.47.
Across the 30 Forge runs, failures cluster on how the platform behaves, not on stale names. Of the apps that got far enough to show them, 10 of 17 failed the LLM explain check and 8 of 17 the live widget. Not one run used the deprecated dashboard gadget.
In our own 27B fine-tune, the Atlassian multiple-choice score slid as general agent training went on while a general drift check read flat, and our attempt to pull the knowledge back pulled an old habit back with it.
Same two models, two benchmarks, tier by tier. On Gauntlet, Flash trails in nearly every 3D and journey tier. On Forge, Flash wins storage and reconcile, the data layer, and Max Prime wins what you can see. Plain mean of each tier's checks, before ceilings and criticals, from the published runs' rows. The Gauntlet tier names are my plain-English descriptions of the checks in each tier.
Qwen3.8 Flash vs Max Prime on Forge: one manifest line, 0.7961 to 0.6953
Forge 1.0 asks for a Jira app on Atlassian Forge: a sprint scope-creep ledger. It listens to Jira events and reconciles them against the real sprint history. It stores every change in Forge's key-value store. It shows a dashboard widget and a sprint action, explains the creep with the Forge LLM API, pushes live updates with Realtime and exposes a Rovo skill. It's built inside a sandbox with no internet and graded by running the app against three seeded Jira sites, where the worst site counts. The prequel walks through the emulator, the fence and the traps.
They came out of it level. Almost. Flash's weighted core score was 0.9537 and Max Prime's 0.9519. Their earned scores, with the excellence slice added, were 0.9426 and 0.9269. Neither took a critical defect. On lint, platform currency, the event pipeline, the resolvers and Rovo, they were level.
Then the ceilings. Forge has four bands, each a set of checks an app must fully clear. Fail any check in a band and the score can't go above that band's cap. Flash's worst failure sat in the band for "current platform, complete surfaces", which caps at 0.799: dark mode. Max Prime failed one check in the band below, "working ledger", which caps at 0.699. Then the scorer pulls each capped score a little lower, by how much the app missed, so two models can't tie on a cap. The arithmetic: 0.799 minus 0.05 times (1 minus 0.9426) is 0.7961, and 0.699 minus 0.05 times (1 minus 0.9269) is 0.6953.
That one check was Max Prime's only band failure. Take it away and nothing else capped it, so its earned 0.9269 would have stood, fifth on the Forge board behind GPT-6.1 Sol, Claude Opus 5.5, GPT-6.1 Sol Pro and Pareto 26.10 Preview. That's arithmetic on one run. I didn't re-score it. But it tells you what one line was worth.
It cuts the other way too. Against Atlassian's servers, forge deploy stopped on that line. Our scorer prices it as one failed readiness rule, not as an app that can't deploy, because I couldn't test what Atlassian does with the linter skipped. If the platform refuses it there as well, Max Prime's real-world Forge result is worse than 0.6953, not better.
Two range attributes, where Forge allows one
Forge's custom entities are the structured side of its key-value store. You declare an entity in manifest.yml with its attributes and its indexes, and you query through an index. An index has a partition, the key you look up by, and a range, the key results come back sorted by. This is Max Prime's index, from the manifest it shipped:
yaml
1indexes:2-name: by-sprint
3partition:4- sprintId
5range:6- at
7- changeId
And this is what Atlassian's custom entities reference says about range: "Optimizes your index for the use of query conditions. This parameter can only have one attribute."
I can see why a strong programmer writes Max Prime's version. Plenty would. A change time isn't unique, so you add the change's id as a tie-breaker. Then the sort is stable. In a SQL database a composite index on sprint, time and id is ordinary good practice. On Forge it's against the rules. And the contract the model was given said, in its storage paragraph, that ledger changes live "in a custom entity indexed by sprint (partition) and change time (range)". It named one range attribute, though it never said a second wasn't allowed.
Because the index isn't a valid one-range index, the scorer doesn't accept it as the sprint index the contract asked for. So "declared a sprint-partitioned, ranged index" scored 0. "Ledger reads go through that index" scored 0 as well, with 0 of 90 reads counted. That second check is the one in the 0.699 band.
Max Prime had two more data-layer defects. They stand on their own. A no-change rerun recorded 104 entity writes over 104 rows, where Flash's recorded 0 over 91, so the idempotency check scored 0. Max Prime's code writes with Forge's FAIL_IF_EXISTS key policy and swallows the "already exists" error, so some of those 104 may be refused attempts rather than overwrites; I haven't confirmed which. And its event path made 61 Jira calls where 26 would do. Flash made 36 against an optimum of 28.
Flash packed the tie-breaker into one attribute
Flash wanted the same stable sort. Its manifest declares one range attribute:
Its code builds timeSort as the timestamp padded to 14 digits, a #, and the change id, padded to 12 digits when it's a number:
js
1constpad=(ms)=>String(Math.max(0,Math.floor(ms))).padStart(14,'0');2constidPart=(id)=>(/^\d+$/.test(String(id))?String(id).padStart(12,'0'):String(id).replace(/[^a-zA-Z0-9._-]/g,'_'));3exportconstrowKeyOf=(sprintId, ms, changeId)=>`${sprintId}:${pad(ms)}#${idPart(changeId)}`;4// each stored row: timeSort: key.split(':')[1]
So it sorts by time, then by id, in one string. The id padding matters: the contract compares change ids as numbers, and unpadded, "10" sorts before "9". A string is also the simple way to carry both in one sortable value: a Forge integer is 32-bit, too small for a millisecond timestamp, and a JavaScript number can't hold a 13-digit time and a 12-digit id together without losing digits. It's the composite sort key DynamoDB people build all the time, so the right Forge answer is an internet pattern too; Max Prime's has the shape of the SQL one. If you need a tie-breaker in a Forge custom entity, I'd copy this, padding included. One caution: Flash's app also re-sorted in memory, and the benchmark runs on an emulator, so I haven't seen this order come back from Atlassian's hosted store.
So both models wanted a composite key. One of them fitted it into the platform.
The app that scored lower looks better. Max Prime puts the totals in cards and the actions first; Flash's modal is in the browser's default serif, because its stylesheet sets `font-family: inherit` and nothing in its iframe sets a font to inherit; there are no buttons in view. The points Max Prime lost are in storage, where no screenshot can see them. Scorer captures in dark mode, two different seeded sprints, from the two run pages, Forge 1.0, captured 5 and 7 October 2026.
The rule lives on Atlassian's servers, not in the linter the model had
Max Prime's closing note says npm run lint gave "0 errors / 0 warnings, all 26 manifest stages complete". That's true. The benchmark's offline kit pins the Forge linter and the manifest schema, so every model sees the same platform. In the pinned schema (@forge/manifest 13.6.0) a range is an array of strings with at least one item. No maximum.
So on 7 October I ran Atlassian's real CLI on both manifests, on a throwaway app registered to our test account. The current @forge/cli, 14.1.0, printed this for Max Prime's:
text
1manifest.yml
20:0 error Storage entity named index must include exactly one range attribute. MANIFEST_INVALID_RULE
34X 1 issue (1 error, 0 warnings, 0 approvals)
On Flash's manifest the same command printed "No issues found." Then I tried forge deploy on Max Prime's. It stopped at that line: "The deploy failed due to errors in the app code." I didn't run deploy on Flash's. I didn't get Max Prime's to bundle with the linter skipped either, so I can't tell you what the platform does with that index if you force it through.
I didn't expect where the rule lives. CLI 14.1.0 ships the same @forge/lint 6.3.0 and @forge/manifest 13.6.0 that our kit pins. MANIFEST_INVALID_RULE is one of the CLI's pre-deployment check rules, answered by Atlassian's servers. That's why forge lint refused to run until I'd registered a real app, and why the error has no line number. The older 12.21.0 CLI installed on this Mac reported two errors about Rovo modules its schema predates, and nothing about the index: it bundles @forge/lint 5.19.1, which has no server-side check at all. If you already have a two-attribute range index deployed, I can't tell you what Atlassian does with it at runtime. A current forge deploy will stop on it the next time you ship.
So a model in a fenced sandbox can lint clean. Atlassian will still refuse what it wrote. Our scorer's storage checks have always required one range attribute, which is why Max Prime's cap was there from the first scoring. The explicit deploy rule, "range has 2 attributes (one allowed)", went into the scorer at 06:51 on 7 October, and the re-score at 07:38 knocked 0.0001 off. Before that, our own deploy-readiness check said "would deploy". We caught up that morning.
Illustration, generated locally with FLUX. Max Prime gave its index two range attributes where Forge allows one, and Atlassian's pre-deployment check refused the manifest. Flash's passed the same check.
Max Prime built the better-looking app
Count the checks where the two differ and Max Prime wins more of them, 7 to 5. Flash won the data layer: the index declaration, index use, the deploy rule, the no-change rerun and event economy. Max Prime won what a user sees and touches. Its explain button produced an explanation that passed 23 of 23 assertions over 5 scripted answers. It passed the live widget check, sorted the ledger right on every seeded site, passed the comment flow and passed dark mode. It made 5 Forge LLM calls with a model id the platform's list() returns. And it didn't hard-code a custom field id. Flash did.
That cuts against the easy story. The Forge LLM API, which reached Preview in June, and Realtime are among the newer surfaces in the task, and Max Prime handled both better than Flash. "The cheap model knows newer Forge" is not what these runs show.
Flash never asked the Forge LLM anything
Flash's app has the LLM code. Its resolver imports list and chat from @forge/llm, picks an active model and asks it to explain the sprint's creep, and its modal has an explain button. On all three scoring sites the explain step recorded no LLM call at all: 0 of 5 explanations, no success flag, no error flag. With no call, the check that the model id is current had nothing to check, so it scored 0 as vacuous.
I can't give you the root cause. The code path looks plausible, and Max Prime's nearly identical model selection worked. Flash's own closing note says it "ran out of call budget before screenshotting the three served surfaces (light/dark) and before exercising the LLM phases". It used 148 of its 150 calls. That explains why Flash never tested the feature. It doesn't explain why the feature made no call when the scorer pressed the button.
It isn't a Qwen thing, either. Of the 17 Forge apps that showed the explain button, 4 made no LLM call the scorer could see, Claude Sonnet 5.5 among them.
Flash's other losses were small and specific. The ledger sorted correctly in 3 of 9 toggle states on the worst site. The comment flow passed 6 of 8 steps. The live widget subscribed, didn't poll and didn't reload, but after the live changes only 1 of its 2 sprints showed the right numbers. In dark mode 4 of 17 surfaces didn't paint the theme's surface token, and three of them came out mostly [18, 18, 18]. Our visual checks look at theme tokens, dark mode, console and policy errors and widget sizes, not at fonts, which is how Flash's serif modal still scored 0.95 on visual. And it hard-coded a custom field id in src/core.js, which the deploy-readiness check flags as a runtime failure on any site where that id differs.
Did either model know the one-range rule? A probe for under a dollar
The runs left me with a question. Did Max Prime not know the rule, or did it know it and write the habit anyway? On 7 and 8 October I asked both models directly, through OpenRouter, on the same Alibaba host, at the same medium reasoning effort Forge uses. Every prompt was a single question. No tools, no other context, 8 tries per model per prompt. A few of Flash's calls hit the host's rate limit and never came back, so some of its counts are out of 7.
The first prompt asked how many attributes a Forge custom entity index's range list may contain, and to start the answer with a number. Max Prime said 1 in all 8 answers. Flash said 1 in 6 and 2 in the other 2. Flash also thought much harder. It used a median of 6,961 output tokens to Max Prime's 354. At a 6,000-token cap most of Flash's first answers came back empty, so every Flash number here is from runs with a 32,000-token cap.
That looked like Max Prime knowing the rule. Then I asked the control question: the same words, about the partition list instead. Atlassian's page is plain on that one: the partition "can have multiple attributes". Max Prime said 1 in 7 of 8, and one of those answers added that the range "may contain multiple attributes", the exact reverse of the docs. Flash said 1 in 7 of 7. The "1" doesn't show knowledge of Forge's rule. It's what both models say about either list. Two of Max Prime's range answers explained it by analogy to a key-value store, one as "a single sort/range key in a key-value store". That one also said, correctly, that the partition may hold several; asked about the partition directly, Max Prime said one in 7 of 8.
The build prompt asked for the indexes: YAML of a Forge custom entity for sprint scope changes, "read by sprint (partition) and ordered by change time (range)", and added that several changes can share a timestamp, so the order must be stable. The tie part comes from the benchmark: its contract asks the ledger table to show equal times ordered by change id. It doesn't say the index has to do it, and asking for it in the index nudges a model toward a tie-breaker there. I ran the request three ways:
what the prompt carried
Max Prime: one range attribute
Flash: one range attribute
the request, which already names the partition and the range
0 of 8
0 of 8
the request plus the contract's storage sentence, as in the benchmark
1 of 8
0 of 7
both of those plus Atlassian's: "This parameter can only have one attribute."
8 of 8
8 of 8
4 rows × 3 columnsHeader row enabled
Every version already called change time the range. So the real split is with and without Atlassian's sentence. The contract's sentence, "indexed by sprint (partition) and change time (range)", names one range attribute without saying a second isn't allowed, and in the probe the models took the hint once in 15 answers. Atlassian's sentence got both down to one range attribute. It didn't get them to a valid manifest. Only 7 of those 16 answers pass Forge's own manifest schema, Max Prime's 5 of 8 and Flash's 2 of 8. Most of the rest wrote partition: sprint as a plain value where Forge wants a list. By its name, Flash's single attribute was a combined key in 5 of 8 answers (changeTimeAndId and the like) and probably in 2 more (ledgerSortKey, changeTimeOrder). Max Prime's was clearly combined in 3 (changeTimeSeq twice, changeTime_changeId) and probably in a fourth (changeKey); three used a plain time attribute, and one slipped the change id in under a separate sort key that Forge doesn't have. Without Atlassian's sentence most answers didn't even use Forge's own partition and range keys; they wrote partitionKey, sortKey, rangeKey, fields or keys.
I can't say Max Prime ignored the rule. I can't show either model knew it. In the benchmark run Flash got it right with the same contract Max Prime had. Given only that contract sentence in the probe, it wrote two attributes 7 times out of 7, so its run may have been the lucky draw rather than the rule.
What it does show is narrower and, for anyone building on Forge, more useful. A quiz fooled me. Both models mostly answered "one", whether the truth was one or many. What changed the answer wasn't more thinking. It was one sentence of small print. Without it they kept to one range attribute once in 31 answers. With it both did every time, though fewer than half of those answers were YAML Forge's schema accepts. The sentence fixed the one mistake an offline lint can't see. The lint would have caught the malformed YAML, but not the three valid answers that quietly dropped the tie-breaker the prompt asked for. That was 100 calls and about 60 cents. I've kept the prompts and every answer with the research for this piece.
Where Qwen3.8 Max Prime won: the Gauntlet payments app, 0.8364 to 0.6846
Gauntlet 7.2 asks for Meridian, a payments console against a mock vendor API with more than 12,000 seeded payments. It has to sync with cursor pagination and ETags. It has to apply webhooks in order while deliveries race, and keep money conserved. It has to survive a SIGKILL halfway through a resync and run a maker/checker approval flow. And it has to draw a 3D field of payment towers that a browser probe inspects pixel by pixel. The 3D work is 46% of the weighted checks.
None of that is niche. Every piece of it is a pattern with thousands of public write-ups: idempotency keys, webhook ordering, outbox tables, WebGL picking. That's the point of the pair.
Here Max Prime was the better programmer. In this run, by a distance. I counted the 15 checks where the two differ. It won 13. It earned 0.9499 to Flash's 0.7126, took no critical, and its only ceiling was the top band, 0.839, for three backend-recovery and replay defects. Flash was capped at 0.699 for five 3D and streaming checks.
Flash's live stream never applied
Streaming is the clearest difference. Gauntlet's vendor pushes payment changes over server-sent events, and the app has to apply each batch to the table and the 3D field. Max Prime's app showed a batch "visible in 63.6 ms". Flash's recorded "no stream batch applied", and the 3D check that watches for the stream's changes in the field saw none either. The stream-latency check scored 0 with it. Separately, so did the check that the app paints a change before the write is confirmed: "state did not paint while the write was provably held".
I'd look at the 3D field next. The rest of Flash's losses sit there. Its inspector is the close-up view of one payment's tower. It framed the tower in 0 of 8 camera poses; in one of them the currency collar sat 27.3 pixels off, against 24 allowed. Label culling scored 0.475. An invariant probe made 16,902 observations of the app's payment states and found 1,471 states the vendor never sent. The committed-event replay failed all three legs. Max Prime passed every one of those except one replay leg.
The approved payment that never showed up
Flash's one critical is familiar. Claude Opus 5.5 took the same one on this board in the prequel. In the scorer's words: "approval completed, but the created payment did not appear in the UI table". The maker created a draft, the checker approved it, the vendor sent it, and the payments table never showed it. The journey check scored 0.7143, and as a critical it multiplied Flash's earned score down from 0.8045 to 0.7126. The 0.699 cap was already binding. On the board it cost Flash 0.0046.
Same probe, two apps. Both sent the 1,250 euro draft. Max Prime's table shows the payment it created, pending. Flash's capture doesn't show its table; the scorer's journey check is what says the payment never appeared there. Scorer captures from the two run pages, Gauntlet 7.2, 5 and 6 October 2026.
The two checks Flash won
Flash beat Max Prime on exactly two Gauntlet checks. Both are backend crash recovery. When the scorer killed the app mid-run, Flash's outbox resumed and nothing was lost. Max Prime's didn't resume: "resumed: False, none_lost: False", scored 0.5. Flash's notifier processed 1,481 events, each exactly once, with full coverage across the crash. Max Prime's covered 0.60 of the crossing window and missed three event ids.
That's a real result. The money buckets and the crash recovery, Flash got right. What it didn't finish was the live and visual layer on top, which is where Gauntlet puts nearly half its weight.
Normal programming vs a niche framework: what the two benchmarks measure
Gauntlet is the internet's kind of app. It's everywhere. A model that has read a great deal of code, and was trained hard to turn that into working software, should do well on it. Max Prime did. Forge measures a niche platform that changes every month, and the model builds it with no way to look anything up. We measured the fence on Flash's own Forge run: direct internet "blocked (curl exit 7)", and a relay that allows exactly one host, openrouter.ai:443. What the model has instead is the contract, pinned type definitions, the pinned manifest schema, Jira's OpenAPI files, a local linter and a dev site. Whatever it knows about Forge beyond that, it brought with it.
The platform moved in the fortnight before the runs
Forge 1.0 uses ten module types plus the Realtime API, and three of those module types changed status in the two weeks before the first Forge runs on 3 October. The prequel has the dates. The one that matters most here: the rovo:skill module "has progressed from the Early Access Program (EAP) to Preview" on 2 October, per the Forge changelog. That's after every Qwen3.8 model on OpenRouter was published, the last being Max Prime on 23 September. So whatever either model learned about it in training was early-access material at best.
Forge doesn't test names, it tests rules
Forge 1.0 is designed to test how the platform behaves, not whether a model remembers names. The contract names every module key, so a model can't fail by not knowing that dashboards:widget exists. It fails by not knowing how the platform behaves.
The 30 published Forge runs bear that out. I counted every failed check across them. Names aren't the problem. Not one run used jira:dashboardGadget. The scorer caught no run on the old storage export from @forge/api, on nodejs18.x or on the removed /rest/api/3/search. (This tutorial covers that last one if it bit you.) One run, Solar Mini4's, did import UI Kit's old @forge/ui, but its manifest was broken before that mattered: it declared no functions and pointed its widget at a resource it never declared, so the app couldn't be opened and that check never looked at its imports. The failures cluster on behaviour. Of the 17 apps that showed the explain button, 10 failed the LLM explain check. Of the 17 whose widget rendered, 8 failed the live Realtime check. The custom-entity index check failed in 10: seven declared no custom entity at all, and three, Max Prime among them, declared one without a valid sprint-partitioned, ranged index.
That's why the index line is a good example of what Forge catches. It isn't a name. It's a limit, written in one sentence on one docs page, that cuts against a habit every database programmer has. And as the probe showed, neither model acted on that sentence until it was handed to them.
Spotting a model that knows the internet but not your platform
This is what the pair is for. It's why we built a second board. One score tells you how good a model is at something. Two scores, one general and one specialist, can hint at where its strength stops carrying over.
Every model on both boards. Above the line, the model did better on Forge than on a payments app; below it, worse. The two Qwen models built on the Flash architecture sit above it, and so does the 27B, at the bottom; Max Prime sits below. One published run per model per board, final scores from the live boards, 8 October 2026, 08:08 EEST.
Thirty models have a run on both boards. Twenty score lower on Forge than on Gauntlet, which is what you'd expect from a niche, moving platform built blind. DeepSeek Pro Latest drops furthest, from 0.6591 to 0, because its Forge app declared no modules at all. Gemini 3.8 Flash drops 0.5028, from 0.6929 to 0.1901. Muse Spark 1.3 drops 0.4514. Claude Sonnet 5.5 drops 0.2476, and Max Prime 0.1411.
Ten go the other way, and the five biggest gains are Claude Opus 5.5 (+0.1841), GPT-6.1 Sol (+0.1791), Pareto 26.10 Preview (+0.1493), Qwen3.8 Omni Flash (+0.1260) and Qwen3.8 Flash (+0.1115). Omni Flash is, in Qwen's words, "Built on the Qwen3.8-Flash-Next architecture". So both Flash-line Qwen models lean Forge, and the flagship from the same lab leans the other way. The unlabelled dot touching Flash's is Jev Router, 0.7978 on Forge, which picks a model and reasoning effort per request, so it isn't one model's run.
I want to be careful with that. Omni Flash's Forge score carries a x0.6 critical, and its Gauntlet score two criticals. Its position is as much about which defects it hit as about what it knows. The boards also differ in settings: Forge pins every model to medium reasoning effort and cuts the internet, while Gauntlet runs each model at its default effort with the network open. And the two scorers aren't on a common scale, so the diagonal is a reference, not a par line. Two points aren't a trend. I'd read it as a hint, not a verdict. But if I were choosing a model to build on a platform like Forge, this is the chart I'd want. Not who's best at coding. Who's best at coding on my platform, relative to how good they are at coding.
What Qwen3.8 Max Prime and Qwen3.8 Flash are, from Qwen's own pages
I read Qwen's own material on both models on 7 October, plus OpenRouter's listings, which is where the model names on our boards come from.
Qwen3.8 Flash
Qwen3.8 Max Prime
what it is
the production version of Qwen3.8-Flash-Next
"a higher-throughput variant of Qwen3.8 Max", per OpenRouter
size
125B, 6B active, plus 51B n-gram embedding and 4B MTP
not stated for Prime; Max is 2.4T total, 95B active
architecture
mixture of experts, "an early preview of the architecture used in Qwen4"
Max: mixture of experts, "Built upon the architectural foundation of Qwen 3.5"
released
blog 26 August 2026
Max blog 3 August; Prime on OpenRouter 23 September
price per million tokens, in / out
$0.15 / $0.47
$4 / $12 (Max: $2 / $6)
knowledge cutoff
not stated
not stated
training tokens
not stated; "1/3 the training tokens" of its predecessor
not stated
post-training
not described
Max: large-scale reinforcement learning, described
distilled from Max
not stated anywhere
n/a
10 rows × 3 columnsHeader row enabled
Max Prime: OpenRouter's description, and Max's training
I couldn't find a Qwen announcement for Max Prime. Qwen Cloud's page for it returned "Page not found" on 7 October, and Alibaba Model Studio's model list, last updated 28 September, doesn't include it. The closest thing to a first-party description is OpenRouter's, where Alibaba is the only provider. It reads: "Qwen3.8 Max Prime is a higher-throughput variant of Qwen3.8 Max from Alibaba's Qwen team, served as a separate SKU at a higher price point." I'm taking that at its word. I treat Prime's training as Max's. That's an assumption. I also can't tell you which Max snapshot it serves.
Max is a different story. It's well documented. Qwen's release post of 3 August calls it "the most capable model in the Qwen family to date", at "2.4T parameters (95B active)". It also says how it was post-trained: "By jointly scaling RL environments and compute, we lift general working competence uniformly across several popular harnesses (QwenWork / Claude Code / Codex / OpenClaw / Hermes)." Its first figure is captioned "Qwen3.8-Max shows steady, consistent gains across dozens of in-house and public working benchmarks as RL training continues to scale up." The open weights are on Hugging Face as Qwen3.8-2.4T-A95B.
That's a model trained very hard, with reinforcement learning, to be good at general software work in common coding harnesses. Gauntlet's parts and rules are the internet's common ones. Forge's rules aren't.
Flash is described as its own model, not a shrunken Max
OpenRouter's qwen/qwen3.8-flash links to Qwen3.8-Flash-Next on Hugging Face. The model card says "Qwen3.8-Flash is the official version based on Qwen3.8-Flash-Next with more production features, e.g., 1M context length by default, official built-in tools." Search results mix the two names up, so to be exact: our board ran the hosted Flash, not the open weights, and a local run of the open model would be a different test that I haven't run.
Qwen's Flash-Next post of 26 August calls it "an early preview of the architecture used in Qwen4". It has "a 125B-parameter main model, supplemented by an additional 51B N-gram embeddings, with 6B parameters activated per token". Training it "takes only about 1/9 as much" as Qwen3.7-Plus. Its tech report adds that it used "1/3 the activated parameters, 1/3 the training tokens, and roughly 1/9 the training FLOPs" of that 397B predecessor.
Nothing I read says Flash was distilled from Max. It's described as its own pre-training run on a new architecture. And nothing says what Flash's post-training looked like; the card says only "Pre-training & Post-training".
Neither has a published knowledge cutoff
OpenRouter returns knowledge_cutoff: null for all six Qwen3.8 listings. Artificial Analysis has none either. Qwen's training data summary says the models are trained on "trillions of tokens of text, image, video and audio". It adds: "The first date a dataset was used during the development of our model pre-dates January 2022." No end date.
So one tempting explanation, that Flash came later and simply saw newer Forge docs, is one I can't check. I don't need it for the probe. On this rule the two models behaved the same, with and without the docs. A stated cutoff wouldn't settle it anyway. Dated Data, by Cheng and colleagues, found that "effective cutoffs often differ from reported cutoffs", partly because new web crawls carry a lot of old pages.
What independent scores say
On public coding scores the two trade wins within a few points. Artificial Analysis puts the open Qwen3.8-Flash-Next and the original Qwen3.8 Max (0803) level on its Intelligence Index, 40 each, and calls the newer Max (0902) "notably slow and very verbose". It has no page for the hosted Flash or for Max Prime. Qwen's own tables, read across two of them, put Max ahead on SWE-bench Pro, 67.7 to 62.5, and Flash ahead on DeepSWE, 58.7 to 56.6.
On Hacker News, the day Flash-Next came out, one commenter put the trade-off in a sentence: "it does not have a world knowledge of larger models. It has most of theirs intelligence." Another answered with the point that matters for Forge: "Models with lots of 'world knowledge' have a good chunk of that knowledge go stale, and there's no real way to refresh it without training a new model."
Illustration, generated locally with FLUX. Max Prime has 95 billion active parameters to Flash's 6 billion, and neither model says when it stopped reading. The big one's book is heavier, not newer.
Catastrophic forgetting in fine-tuning: what our 27B lost
This part I know from the inside. We did it to our own model, so it's evidence about fine-tuning, not about Max Prime. We fine-tuned the open Qwen3.8-27B, the same base that scores 0.0915 and 0.1582 in the table up top, to be better at goose agent work and at Atlassian. The model page has the published result and the prequel has the bill. What I want here is what it taught me about forgetting. It's the closest thing I have to evidence for the question anyone choosing a model for Forge should ask: does training a model hard on general work cost it the niche?
As T10 trained on general agent work, its Atlassian score slid
T10, the round that never shipped, was a LoRA continuing our earlier adapter, T9, on a 50-million-token mix, stopped about 70% of the way through. 78.9% of the mix was synthetic goose agent sessions, mostly general tool-calling work with some Atlassian admin and migration tasks in it. 16.9% was general replay. Atlassian knowledge was 2.2%.
On our held-out 199-question Atlassian multiple-choice test, T9 scored 72.36. T10's first checkpoint, at step 506, scored 72.86. Every later one scored lower: 68.84 at step 1,012, 66.83 at 1,518, 67.34 at 2,024, 66.83 at 2,530 and 67.34 at 3,035, and 64.32 at step 3,541. It slid early and mostly plateaued, though only the final drop is clearly beyond noise. The later checkpoints were better at the goose work they were being trained on. The run stopped on a proof, not a score trend: a recount showed step 2,024 at 134 of 199 against T9's 144, below the gate. The 64.32 was measured after the stop. The prequel plots the curve.
Those numbers have three limits. Only the last drop is statistically significant on its own (p = 0.0037); the middle ones sit at p 0.08 to 0.26. They came from a home probe that read the unmerged adapter, and for one candidate that probe said 4.0 points down where the release, measuring the shipped 8-bit build, said 1.0. And T10's 30 Forge questions didn't drain the same way: 90, 90 and 83.33 against T9's 86.67. The loss was mostly in Service Management and migration questions.
The general drift check never saw it
We ran a guard on a separate GPU through the whole run. One of its checks measured how far the model's predictions on 200 rows of general text drifted from T9's, as a KL divergence. It stayed at about 0.006 the entire time. The KL check stayed flat. The niche knowledge slid.
It rhymes with Gauntlet and Forge. A general measure can stay flat while the specialist one moves, and if you only measure the general one you won't know.
Illustration, generated locally with FLUX. 78.9% of T10's tokens were tool work and 2.2% were Atlassian. It got better at the tools, and the Atlassian notes went out the top.
Pulling the knowledge back pulled a habit back too
T9 had a bad habit. It invented people, greeting an invented forum handle or signing off with an invented name. It learned that from its data. The untouched base model did it 0 times in 120 replies. T10 started from T9's adapter and inherited it.
Short corrective rounds removed the invented people (0 of 160 sampled replies), and on the home probe the knowledge score stayed down at 68.3. The release later measured that same model at only 1.0 point under T9. Rounds that pulled the model back toward T9, with 400 T9-written answers, brought the knowledge score back to 71.4. And the invented people came back with it: 5 of 160 with an anchor and 3 of 160 without. Blending the two adapters half and half bought back only 1.5 points. Our conclusion at the time was that "knowledge and the habit travel together through T9's distribution". The release number above weakens it, because much of the gap we were pulling back was the probe's.
That's a different kind of habit from Max Prime's composite key, a writing tic rather than a design choice. What it taught me is about fine-tuning: what a model knows and how it habitually writes aren't stored in separate drawers. Pull on one and the other moves.
T11: teaching one niche skill cost another
T11, the model we published on 5 October, started over from the base model. Its mix diluted T9's validated Forge app code from 19% to about 4%. A late release gate caught the result: it built a complete, valid Forge app only 9.7 times in 35 requests, against T9's 30. A 388-step booster on 2,127 Forge code rows brought that to 27.3. And it cost something. Forge knowledge questions fell from 73 to 63 on the training box, and Atlassian multiple choice from 78.4 to 76.4.
The general benchmarks barely moved. Against T9, the published T11 scored 96.3 on HumanEval to 95.7, and 84.3 on MMLU to 83.0. General coding held steady while the Forge skills see-sawed with every change to the mix. T11's weak Forge isn't forgetting, to be precise: it started from the base and was never taught enough. T10's drain is the forgetting. This tutorial shows the LoRA setup if you want to try it on a Mac, at a smaller context.
What the research says about forgetting
"Catastrophic forgetting" is the old term for a network losing an earlier skill when trained on a new one. Kirkpatrick and colleagues' elastic weight consolidation paper of 2016 fought it by "selectively slowing down learning on the weights important for those tasks", a cousin of the KL anchor we used. For language models, Luo and colleagues found in 2023 that "catastrophic forgetting is generally observed in LLMs ranging from 1b to 7b parameters" during continual instruction tuning. In that range it got worse as the models got bigger.
Kalajdzievski's scaling laws for forgetting found that LoRA "still suffer[s] from catastrophic forgetting" and that forgetting grows with the number of update steps. Biderman and colleagues' "LoRA Learns Less and Forgets Less" is part of why we chose LoRA: it "better maintains the base model's performance on tasks outside the target domain". Better isn't immune, as we found.
On where knowledge lives, I find Gekhman and colleagues' 2024 paper the clearest. It supports "the view that large language models mostly acquire factual knowledge through pre-training, whereas fine-tuning teaches them to use it more efficiently". Kandpal and colleagues showed that a model's ability to answer a fact-based question "relates to how many documents associated with that question were seen during pre-training". Forge's one-range rule is a long-tail fact if anything is: one sentence on one page.
The old name for post-training costing earlier abilities is the alignment tax. OpenAI's InstructGPT paper used it for regressions on public NLP datasets after RLHF. Lin and colleagues' 2023 follow-up puts it in one line: RLHF "can lead to forgetting pretrained abilities, which is also known as the alignment tax".
Code makes the knowledge side harder because libraries move. VersiCode found version-specific code generation "a significant challenge, even for GPT-4o and other strong frontier models". CodeUpdateArena found that putting the documentation of an API update in the prompt didn't let open code models use the change. Our probe went the other way on one rule. Theirs were smaller open models and synthetic API changes, so the two don't conflict, but docs in the prompt aren't a guaranteed fix. Wang and colleagues' study of deprecated APIs starts from the same place: models "may struggle to use correct and up-to-date Application Programming Interfaces (APIs) due to the rapid and continuous evolution of libraries". LibEvolutionEval and GitChameleon measure the same thing with real version histories; in GitChameleon, GPT-4o reached a pass@10 of 39.9%. Forge is that problem, platform-wide.
Did heavy post-training cost Max Prime its Forge knowledge?
I expected to argue yes. The evidence doesn't get me there. Forgetting needs something known first. The probe can't show Max Prime ever knew this rule: it gave the same "one" for a list where the answer is many, and wrote the composite key cold just as Flash did. What it did was reach for the general habit when it was building.
Could heavy post-training make a habit like that stronger? Some things fit. Qwen says Max's post-training scaled reinforcement learning across popular coding harnesses until "steady, consistent gains" showed up on dozens of working benchmarks. Yue and colleagues found that reinforcement learning on verifiable rewards makes a model more likely to get an answer right first time, while "the observed reasoning abilities originate from and are bounded by the base model". If that holds for Max, RL sharpens what a model already leans toward, and a model that leans toward the internet's way of building an index would get more sure of it. Two big ifs. Yue's runs were verifiable-reward RL on maths, code and visual reasoning benchmarks. The paper itself names scaled, multi-turn agent RL, which is what Qwen describes for Max, as an open question. And I'm assuming Prime was trained like Max.
Against it:
Flash, given only the request, has the same habit. In the probe it wrote two ordering attributes 8 of 8 times too. Whatever made Max Prime do it isn't unique to Max Prime's training.
Our own evidence is supervised fine-tuning. We ran no reinforcement learning at all, and a 2025 paper that tests exactly this comparison, RL's Razor by Shenfeld, Pari and Agrawal, finds that "despite similar performance at a new task, RL preserves prior knowledge and capabilities significantly better" than supervised fine-tuning.
A bigger model should hold more facts, all else equal. Allen-Zhu and Li estimate "2 bits of knowledge per parameter", and Kandpal found "larger models are better at learning long-tail knowledge". On this one rule the probe found no difference to explain: neither model showed it knew it.
Max Prime handled the Forge LLM API and Realtime better than Flash in the run. Its failure was one rule, not general ignorance.
It's one run each, at medium reasoning effort on Forge, and Qwen's own card says lower effort in agentic tasks can mean "more failures". A second Max Prime run might pick a one-attribute range and land a band higher. A second Flash run might not pack its tie-breaker; cold, it didn't.
In the prequel Claude Opus 5.5 won Forge by 0.4259 over Claude Sonnet 5.5. Nothing on our boards says the cheaper model wins Forge as a rule.
So my read is this. These runs, and the probe, are consistent with two capable models that both carry the internet's habit and neither of which acted on Forge's rule until it was in front of them. They don't show forgetting, and they can't show what Qwen's training did. What they do show is narrow. For this one rule, a one-line quiz couldn't tell me whether either model knew it, and one sentence of docs in the prompt fixed the rule, if not the YAML around it.
Qwen3.8 Flash vs 27B and Omni Flash: the rest of the family
Two more Qwen models sit on both boards. I find them useful because they fail in different ways.
Qwen3.8 Omni Flash, Qwen's agentic omni-modal model on the Flash-Next architecture, had the best core work on Forge of the five runs in the table up top, 0.9613 before its critical. Then a double-click posted two comments, the critical multiplied everything by 0.6, and it finished at 0.5529. On Gauntlet it took two criticals, one for money it rendered wrong and one for a lost write, and finished at 0.4269. The Forge double-click is the same trap the prequel described for Sonnet 5.5 in reverse: Sonnet's posted nothing, Omni Flash's posted twice. This tutorial on Forge KVS locks covers the exactly-once pattern if you're building it.
The untouched Qwen3.8 27B scored 0.0915 on Gauntlet, where it never drew a visible 3D scene and took three criticals. On Forge it scored 0.1582; its widget never loaded and its backfill was incomplete. It's the base our own fine-tunes start from. Our first fine-tuned version was trained to write Forge apps, and our restart still builds fewer valid ones than it does, 27.3 of 35 against 30.
And GLM 5.3 Prime, as a control, scored 0.6797 on Gauntlet, almost exactly Flash. On Forge it scored 0.4736. Its manifest failed the linter with 19 errors, which capped it at 0.499. One of them reads "app storage entities attributes property sprintId 'string' must be object". The storage section of a Forge manifest is where models trip.
Qwen3.8 Max Prime vs Flash pricing: $16.46 vs $0.47 a run
run
billed
calls used of 150
wall time
prompt tokens (cached)
output tokens per call
Qwen3.8 Flash, Gauntlet
$0.59
146
82 min
25.2M (24.7M)
1,726
Qwen3.8 Max Prime, Gauntlet
$16.06
94
79 min
19.3M (18.6M)
3,471
Qwen3.8 Flash, Forge
$0.47
148
55 min
21.0M (20.7M)
1,308
Qwen3.8 Max Prime, Forge
$16.46
126
82 min
23.7M (23.0M)
1,386
5 rows × 6 columnsHeader row enabled
All four receipts mark the bill a lower bound. I've kept them that way. On the two Gauntlet runs, 1 and 3 generation ids had no billing record. Max Prime cost 27 times as much as Flash on Gauntlet and 35 times as much on Forge. The list prices explain it. Max Prime is $4 per million input tokens and $12 per million output; Flash is $0.15 and $0.47. Plain Qwen3.8 Max lists at $2 and $6 on Qwen Cloud, so Prime is exactly double. And an agent run is mostly re-reading its own context. Caching carried most of it, 98% of Flash's Gauntlet prompt tokens.
On Gauntlet, Max Prime wrote twice as much per call as Flash and used 52 fewer calls. On Forge the two wrote about the same per call, and Flash made 22 more calls, ending two short of the budget.
What isn't in the table: Flash's earlier attempts. Its first runs on 5 October ended on OpenRouter 429 errors, "Rate limit exceeded: Provider returned error", from the only host that serves it. One Gauntlet attempt got to about 148 calls and $0.77 before the cut and was never scored. The published Flash runs are reruns after goose learned to back off. Max Prime's two published runs were its first.
How the two runs were set up, and what differed
The benchmarks are the same ones the prequel describes in full. One task, 150 model calls, no human. It's scored only by running what was built, offline, in a real browser, and for Forge in an emulator built on Atlassian's own runtime. Failures the model caused are scored, not refused. A model's score is its earned work, multiplied down by any critical defect (no lower than 0.6 each), then capped by the lowest band it failed to clear. The methodology page and our write-up of a scorer that runs the app have the rest.
What was the same: the agent (goose, on the two builds in item 3 below), the budget, the task files, and the host. All entrant calls on all four runs went to Alibaba through OpenRouter, with no host pin, so the provider isn't a variable between them. Each run also made one short Gemini call from the app itself, which isn't counted against the model.
What differed:
Reasoning effort. Forge pins it to medium for every model (both Qwen run receipts read "forge-1.0 pin"), while Gauntlet runs each model at its default. That's the same for both Qwen models, but it's a difference between the boards. Qwen's own card warns that in agentic tasks lower effort "can also lead to insufficient analysis, more failures, and repeated retries".
Internet. Gauntlet's sandbox has the network open; Forge's is fenced. I didn't check whether either model looked anything up on Gauntlet.
App version. Flash's runs were scored on goose 3.0.100 and Max Prime's on 3.0.107, and only Max Prime's Forge run was re-scored on 3.0.108. The scorer changes in between were mostly about charging failures instead of refusing runs, plus the new deploy rule. Flash's one-attribute index passes the new rule, but I haven't re-scored Flash to prove the other changes leave it where it is.
Retries. Flash's published runs are reruns after the rate-limit failures above. Max Prime's are first attempts.
What we got wrong along the way
Our offline Forge kit can't reach the Atlassian check that rejects the invalid index, so the model's own lint was clean, and our scorer only added the explicit deploy rule on 7 October, the morning the Max Prime run was re-scored. That re-score moved its Forge score by 0.0001, and our deploy check called the app deployable until then. Flash's runs needed retries because the only host serving it rate-limited us, and those failed attempts aren't in the published bill. I first read the probe as proof that Max Prime knew the rule; the control question took that away. And on our own model, the knowledge-drain numbers came from a probe that overstated one gap four times over, which we only found at release.
As far as I can tell, none of that changes the order on either board. Several of them change how sure you should be. I'm less sure than the scores look.
The honest verdict on Qwen3.8 Flash vs Max Prime
Max Prime built the better payments app in these runs, by a wide margin, and the more polished Forge app. On a big, ordinary app it built nearly everything, at $16 a run, and it's second of 31 on Gauntlet behind only GPT-6.1 Sol Pro. On Forge it got more of the platform working, the LLM explain feature and the live widget among it, then wrote one line Atlassian's pre-deployment check refuses.
Flash was the better Forge result for the money, by a long way on price and by one check on the board. It cost about 3% of what Max Prime did, got the data layer right, and finished sixth of 30 on Forge. It also never made a call from its own LLM feature, and on Gauntlet it left the live and 3D half unfinished.
Neither is the best Forge model here. GPT-6.1 Sol at 0.9769, Claude Opus 5.5 at 0.9767, GPT-6.1 Sol Pro at 0.9664 and Pareto 26.10 Preview at 0.945 are a band above both.
Which model to use for Atlassian Forge apps
If you build Forge apps with an agent and the budget allows it, I'd use Opus 5.5 or GPT-6.1 Sol; the prequel has their runs. If you want a cheap model for Forge work, Qwen3.8 Flash at under 50 cents a run beat Max Prime on our Forge board at a 35th of the price. Test its LLM and Realtime features by hand. Those are what it didn't get working. Give your agent retries with backoff, too: the one host serving Flash rate-limited us on the benchmark and again on the probe.
Whatever model you use, I'd do three things.
Run a real, online forge lint on whatever the agent writes, against a registered app; npx -y @forge/cli@latest lint gets you the current CLI without touching a global install, and in CI the CLI reads its login from the FORGE_EMAIL and FORGE_API_TOKEN environment variables. The check that caught Max Prime's index is answered by Atlassian's servers, so an offline lint, or an agent's "0 errors", won't see it. I'd have the harness run it and hand the output back rather than give the agent the logged-in CLI, because the same login that lints can deploy and install. If the agent must run it, use a throwaway developer account with no admin on any site that matters.
Put your platform's sharpest rules into the task in the platform's own words. Our contract described one range attribute and, in the probe, the models took the hint once in 15 answers. Atlassian's one sentence, "This parameter can only have one attribute", got both models to one range attribute in all 16 answers, but fewer than half of those were valid manifest YAML, so you want the sentence, the linter, and a test that equal timestamps come back in id order. If you need a tie-breaker in a Forge entity index, pack it into the one range attribute, the way Flash did: a zero-padded timestamp, a separator, and the id, zero-padded too if it's a number.
Look at two numbers when you choose a model: its general score and its score on your platform. The gap is a rough guide to how much of your platform's small print you'll have to supply yourself. Atlassian's Forge MCP Server exists to give coding agents "authoritative Forge and Atlassian Cloud knowledge", and even its own page warns its index "may become out-of-date".
For Qwen in particular, my pick is simple. Max Prime for a general app, if $16 a run is fine. Flash for Forge on a budget, with the docs in its context. And a second run of whichever you pick before you believe either.