The free ACP practice exams just lost 250 questions. That was the point
Gabriela Corocea
Author
9 min readSeptember 8, 2026
Key takeaways
The five pools went from 2,937 questions to 2,687. The count only went one way: 253 questions left across two waves and three came in; questions that could not earn their place were held, not rewritten around.
The stems had a shape problem: 30 to 47 percent opened on a named person at an invented company, 19 to 33 percent appealed to 'Atlassian's documentation' inside the question, and they ran about 40 words against a real median of 22 to 33. The companies and documentation appeals are now zero; the medians came down to 26 to 34 words, a few words above each real paper rather than ten.
Every question was sat blind by a reader who had the key stripped and only the cited Atlassian page to go on. Blind picks agreed with the stored key 98.5 percent of the time; five keys were wrong and are corrected, 36 more items were repaired.
The options had a giveaway too: on 54 to 58 percent of single-answer items the correct option was simply the longest. Options are now terse claims, and the longest-is-right rate sits at chance.
The exams stay free with no account and no password. The only ask is an email for the key, and, if the papers help you pass, a coffee, which the page itself says is never expected.
On Monday the free practice exams for all five Atlassian ACPs went live on leanzero.net/certifications with 2,937 questions. By Tuesday afternoon the pools had been cut to 2,687. The stems, the options and the answer keys had already been reworked on the branch in the days before launch; this post covers both, in order.
That is not a regression. It is what happens when someone who has actually sat the papers looks at the questions and says what is wrong with them. This post is the account of what was wrong, what was done about it, and what the pools look like now. The numbers come from the repository history and from running the suites against the shipped files, so where a figure is only in a commit message rather than reproduced, it says so.
The stems problem
Mihai sat the ACP-120, ACP-520 and ACP-620 papers and came back with one observation about the composed questions: they carried "a lot of emphasis on names and fluff" next to the real ones. He was right, and it was measurable.
Every composed pool was compared, axis by axis, against his own screenshots of the real papers. The gap was on every axis. Between 30 and 47 percent of stems opened on a named persona at an invented company; the real papers use no companies at all, and a first name in only about 12 percent of items. Between 19 and 33 percent of the composed stems appealed to "Atlassian's documentation" inside the question, which real papers never do. Only 0 to 3 percent were direct questions, against 13 to 33 percent on the real thing. And the stems ran around 40 words, against a real median of 22 to 33.
The cause was a single instruction. The composer specification had been calibrated on ACP-120's harvested bank alone, and it told every batch to "use a fresh persona and a fresh business scenario." Every batch obeyed, and every pool inherited the same tic.
A real before and after from the ACP-120 pool shows what the fix looked like in practice. Before: "Deshi administers the company-managed project 'Lighthouse' and wants any visitor on the open internet, including people with no Atlassian account, to browse and search its issues without logging in. Which single change to Lighthouse's permission scheme achieves this?" After: "You administer the company-managed project 'Lighthouse' and want any visitor with no Atlassian account to browse and search its issues without logging in. Which single change to the permission scheme achieves this?" Same discriminator, same key, six fewer words and nobody to keep track of.
Another one went further. "Organization admin Odalys verifies the email domain 'brightloom.com' for her Atlassian organization and assumes every Atlassian account using that domain is now a MANAGED account under her control. Which statement about her assumption is correct?" became "Which statement is correct about an Atlassian account on a domain the organization has verified but not claimed?" The persona was carrying nothing; the question was always about verified versus claimed.
The reshape ran across all five pools: stems rewritten into second-person, bare-situation or direct frames, employers removed, discriminators kept, and then a second pass that removed every documentation appeal. Only the stem, and any persona or company mention that had leaked into the option text, changed. A text-mapped check confirmed that no answer key moved and no quote changed in that pass.
The per-exam table tells the story. ACP-120 went from 43 percent invented companies to zero, documentation appeals from 26 percent to zero, median stem length from 41 words to 33 against a real 29. ACP-520 was the worst offender at 80 percent companies and is now at zero, with the median down from 39 words to 26 against a real 22. ACP-220, ACP-420 and ACP-620 all went from 58, 59 and 61 percent companies to zero.
A sweep of the pools as they are shipped today, with plain regular expressions rather than the campaign's own tool, finds zero company suffixes and zero documentation appeals in every exam, "You..." openers between 12 and 34 percent, and median stem length between 24 and 34 words. First names still open 17 items, 16 of them in ACP-220. The real papers use a first name in about 12 percent of items, so a name on its own is not a defect; the invented employers are what were removed.
The options had their own tell
Stems were not the only shape defect. Real options are terse claims, a median of 4 to 7 words. The composed options explained themselves at 14 to 17 words, and on 54 to 58 percent of the single-answer items the key was simply the longest option. The real papers sit at 22 to 36 percent, and chance is around 20. A candidate who had never opened Jira could have beaten the pools by picking the longest line.
The options were reshaped to terse claims, with the median down to 6 to 8 words, and the longest-option-is-right rate now falls at 17 to 37 percent, which is chance. The multi-select rate had also been copied from ACP-120's bank onto every exam; it now matches each real paper, which means ACP-420 runs no "choose two" items at all and ACP-520 is exactly three options, single answer. Keys, quotes and evidence were untouched in this pass; 35 items were repaired by hand.
Every key, checked blind
Mihai then asked the question that mattered: had all of the items actually been fact-checked? The honest answer at that point was that every quote had been machine-verified verbatim on the live Atlassian page and every keyed option was linked to a proving fact, but a blind check of the key itself had only been sampled, about 150 items. Self-consistency cannot catch a composer that misread a page and then wrote a coherent explanation of its misreading.
So a blind wave ran across the whole ledger of 2,941 items. Eighty-six reader agents each received around 40 items with the key stripped, read the cited Atlassian page, picked an answer and quoted the sentence that decided it. A collate step compared every pick to the stored key and refused to close an exam until every item had a verdict.
Blind picks agreed with the stored key on 2,897 of the 2,941 items, 98.5 percent. Reader confidence across all verdicts was 91 percent high, 6 percent medium, under 1 percent low. Five stored keys were wrong and are now corrected: two in ACP-120, one each in ACP-220, ACP-420 and ACP-620, none in ACP-520. A further 36 items were classed as defective, meaning both a distractor and the key were defensible from the page, and in those the key stayed and the offending option or stem was reworded so that exactly one option is true for a reason the page states.
What this does not prove is worth stating as plainly as what it does. A reader and the composer can share the same misreading of an ambiguous page, and the 6 percent of medium-confidence agreements are inferences rather than stated sentences. It is a much stronger check than a sampled one; it is not infallibility.
Then the cut
Two more waves followed on Tuesday, and those are the ones that took the count down. The first, described in the repository as the coherence wave, read every item whole: options answerable from their stem, paper vocabulary, retired terms removed. Thirteen items were marked unpublishable by the settlers, and the build gate honours that mark by dropping the item and printing its reason. The pools went from 2,937 to 2,929. The 13 is the commit's count over the ledger; only eight of them were in the published pool, which is why the net is eight, not thirteen.
The second, described as the exam-taker UAT and grand-master waves, is summarised in the commit as: every item sat blind by an expert, then every key and distractor fact-checked on the live Atlassian page, quotes re-anchored to the deciding sentence, duplicates and stale keys held. That wave took the pools from 2,929 to 2,687. Per exam, from the blind-key plateau: ACP-120 lost 43 questions, ACP-220 49, ACP-420 30, ACP-520 74 and ACP-620 54.
Two things to be straight about. First, those last two waves are documented only by their commit messages; there is no per-exam table for them the way there is for the reshape and the blind-key wave, so the description above is quoted, not reproduced. Second, the commit message for the final rebuild says 809 and 377 for ACP-120 and ACP-520, and the shipped files hold 808 and 376. The files are what the site serves, so the live counts are 808, 730, 424, 376 and 349.
"Held" is the word the commit uses, and it is the right one. A question that duplicates another, or whose key no longer matches the page, is not fixed by rewording it around the problem. It is held back until it can earn its place, and the pool is smaller and better for it.
What the suites say
Three test suites sit behind the pools. The evidence suite reproduces on the shipped files: 29,090 assertions passed across 2,687 questions. Each item must carry a non-empty evidence array, every entry a quote of at least 15 characters and an https URL on an atlassian.com host that is neither Community nor the retired Confluence docs domain, no duplicated quote within an item, and the headline quote present in the evidence. It checks structure, not live truth; live truth is what the blind wave and the fact-check on the page were for.
The key suite passes 15 of 15: the access key is derived from the email address with an HMAC, deterministically, case- and space-insensitive, and it rejects a wrong address or a mangled key.
The pool suite is the one worth flagging. It traces every published question back to its source batch in a separate repository, and on the machine this was written on that repository is a few days stale, so 30 items match against pre-reshape option text and the suite fails on exactly those three checks. Every other assertion in it, 16,167 of them, passes. That is a provenance check against an outdated local copy, not a defect in the shipped pool, but it is the kind of thing that should be said rather than hidden behind "suite green."
Every question in every pool carries the same eleven fields, and none of them is empty anywhere: the stem, the options, the answer, the pick, a per-option analysis, the verbatim quote, the deciding rule, the documentation references, the evidence array, a topic and an id.
Still free, and what keeps it that way
None of this changes the deal. There is no account and no password. You give an email address, the key is emailed back, and that key unlocks the exams in your browser for six months. The address is recorded only when a key is successfully verified, and the key is derived from the address rather than minted and stored. Your answers, flags and notes stay in your browser and are never sent to a server.
What is new on the exam pages is a small card under the paper controls, and again at the end of the results page. The one on the sitting page is headed "Free, and staying free" and reads: "These papers are a community service. If they help you get certified, a coffee keeps them going (never expected)." The one on the results page asks whether the paper was useful and says a coffee keeps them free for the next candidate. That is the whole business model. The link goes to Buy Me a Coffee, and the wording is deliberate: never expected.
The measure of whether all this worked is not a shape statistic. It is whether a prepared candidate picks the distractor at the rate the real paper makes them. That is not measured yet; the campaign notes say the real check is the next mock against the pool. In the meantime the five pools are live at ACP-120, ACP-220, ACP-420, ACP-520 and ACP-620, and if you sit one and a question feels wrong, that is exactly the kind of feedback that took 250 questions out this week.