A candidate found four bad practice questions. Here is what we did about it
Mihai Perdum
Author
9 min readSeptember 10, 2026
Key takeaways
A candidate flagged four ACP-120 practice questions. All four were bad, and all four had passed four rounds of review. Every one of those rounds read the question next to the Atlassian page. None of them ran anything.
There is a fifth check now: read the question with nothing else open, then try the claim on a real Atlassian Cloud site. Across the five exams it read 2,941 questions, found 155 claims the product contradicts, repaired 566 and withdrew 139.
ACP-220 was reshaped to the shape of Atlassian's own practice paper and read again from scratch, and 40 duplicates went. The live pools are 795, 668, 424, 375 and 344. If a question feels wrong, write to office@leanzero.net, mihai@leanzero.net or gabriela@leanzero.net, or tell us on Discord.
Two days ago we put up a post about the practice exams losing 250 questions, and it ended by asking anyone who sits one to tell us when a question feels wrong. Somebody did, the next day. Four ACP-120 questions, a line under each, and a closing sentence I have been thinking about since: it looked to her like the questions had had minimal human review, or none at all.
If you sat the ACP-120 paper on or before the 9th of September and one of these cost you a mark or a minute of staring at the screen, I am sorry. Here is what the four were, what they told us about our own checks, what has changed since, and the plain way to reach us when you meet the next one.
The four
The first one asked what happens when a saved filter from 2022 with the clause Epic Link = PRINT-100 runs today. Her note: it would fail, because the field name should be in quotes. On a live Jira Cloud site the unquoted clause is a syntax error and the parser tells you why in so many words, Expecting operator but got 'Link'. We had shipped a question whose own query could not run.
The second asked how to get a field onto Bugs across several support projects while keeping it hidden in one company-managed project. She said it was not a good question and answer, and that the Atlassian page behind it is itself badly written. It is. Our stored answer was a doc sentence read as if it were a rule, and the sentence is awkward on the page too. It has been rebuilt as a decision with a real answer.
The third mixed a dropdown option with ID 12 and closedSprints() into one search and asked which two statements about querying that field were correct. Her note was that she had no idea what the question was trying to ask. Once we read it without the page open, neither did we. It was keying a performance footnote about option IDs inside sprint functions. It is gone.
The fourth had a user with the Create attachments permission who still could not attach a screenshot in one project, and asked which condition explained it. Our answer named a project-level attachment setting. Her note: there is no project-level setting for attachments, only site-level settings and project permissions. Correct. Attachments are switched on or off for the whole site. The question now keys that switch, and the project-level setting is what it always was, a wrong answer. A sibling question that leaned on the same non-existent setting went in the same fix.
All four had passed four rounds of review. The last post described them: a coherence pass, a blind key check with the key stripped and only the cited page to go on, and two more on top of those. Four rounds, and all four went straight through.
What the four rounds had in common
Every one of them read the question with the Atlassian page open beside it. None of them ran anything.
That is uncomfortable, because the rounds were doing exactly what they were built to do. When the cited sentence is right there, the question's wording looks fine. "Applied across your site" reads like a rule. "If your space allows attachments" reads like a setting. A reader who has just seen the page fills in whatever the question leaves out, and it never occurs to them to ask whether the query would parse or whether the setting exists at the level the question says it does. The last post even admitted this, that a reader and the writer can share the same misreading of an ambiguous page. We wrote it as a caveat. She showed us it was a whole class of bug.
Who "we" is
Her closing line was about human review, so it is fair to say who is actually doing this. Mostly it is Gabriela. She has been sitting up nights with these banks for weeks now, reading questions one at a time, holding them against the page, fixing the ones that are off and throwing out the ones that cannot be fixed. We do use AI tooling, and I am not going to pretend otherwise: it parses the queries, counts the shapes, runs the mechanical checks and gives us a second cold read of a batch. It is a real help. But the reading and the judgement are still a person's, most of the time a tired one, and that is part of the honest answer to how four got through, because the tooling did not catch them either.
The fifth check
So there is a fifth one now, in two halves that are kept apart on purpose.
First, a cold read. Whoever reads the question, Gabriela or the tooling on a second pass, gets no key, no documentation, no study notes, and is not allowed to open an Atlassian page until they have written one sentence: the candidate has to decide such and such, given such and such. If that sentence cannot be written without guessing, the question is marked unintelligible and left that way. Nobody gets to work backwards from the options to what the stem should have said. That guess is the defect.
Then the product. Every claim the answer rests on gets tried on a live Atlassian Cloud site instead of read on a page. The JQL gets parsed and run. The permission name gets looked up in the site's own list. A "site setting" gets looked up where site settings actually live. And the rule is deliberately lopsided: if the product contradicts the answer, the question goes. It only gets a new answer when the product and a support page agree on what that answer is, never on the product alone, because the real exam is written from the documentation, and where the documentation and the product disagree there is no fair question left to ask.
That rule cost us questions I would rather have kept. The JQL page one ACP-620 question cites says the WAS operators take only a status name, never an ID. On a live site, status WAS 10003 and status WAS "Backlog" return the same count. The question keyed "invalid" on the doc's sentence. It was withdrawn, not corrected, because whichever way we keyed it, somebody who had learned the other version would be marked wrong.
Across the five exams that check went over 2,941 questions and marked 505 of them muddled and two unintelligible. It tried 5,224 product claims. 2,611 held. 155 were contradicted by the product. The remaining 2,458 could not be reached at all, because they live on organisation-admin or Premium screens that a read-only call on our test site cannot see. We repaired 566 questions and withdrew 139. Not one key was changed on the product's word alone, which is the lopsided rule doing its job.
The unquoted Epic Link also turned into a sweep. Every span in every stem, option and explanation that looked like JQL, 288 of them in 125 questions as the sweep logged it on the 9th, went through the live parser. What came back was prose that happened to look like JQL, and the wrong answers that are meant to be invalid syntax in the first place.
ACP-220, further than we said we would go
I sat Atlassian's own ACP-220 practice paper on the 8th. Sixty-three questions, 59 right. Those 63 are the closest thing we have to the real paper's shape, so we measured them instead of guessing. Not one first-name persona, or one in 63 to be exact, and that one carries nothing. Forty-four percent of the stems start with "You". The median stem is 23 words and the median option is four. Four options on 97 percent of questions, three on the rest, never five or six. Not a single choose-two.
Our ACP-220 pool at the time ran 22 percent personas, 27-word stems, six-word options, 21 percent choose-two, and five or six options on one question in five. The last post said a first name in a stem is not a defect on its own, since real papers use one in about 12 percent of questions. On this paper it is zero. So for ACP-220 we went further than we had promised.
Every question was rewritten as one thing, stem and options and explanations together. The mess of three days earlier had come from exactly the opposite: one pass over the stems, then a separate pass over the options, and nobody reading the question whole in between. Where a choose-two became a single answer, one correct option stayed and the other was removed, and because it was true, a removed correct option is never allowed back in as a wrong one.
Then every rewritten question was read cold again, by a second reader that had not touched it. That second read marked 664 clear, 116 muddled, none unintelligible. We repaired 113, withdrew 58, and moved two keys. One of those was a rewrite that had shifted the true claim onto a different letter, which is precisely what the second read is there to catch.
The last thing to come out of that wave was duplicates, from a scan across every batch once the read had closed. Ninety-four published questions, in 46 groups, shared one verbatim documentation sentence: a scenario in one batch, and the same sentence again as "which statement is true" in another. In a single practice session you could meet it twice, and the second time it is free. Forty went. Eight pairs stayed because they ask genuinely different things of one sentence, the rule in one and the diagnosis in the other. A check now refuses to close a release wave if a new pair appears.
Where the numbers are today
The practice exam pages serve 795 questions for ACP-120, 668 for ACP-220, 424 for ACP-420, 375 for ACP-520 and 344 for ACP-620. Two days ago it was 808, 730, 424, 376 and 349. Eighty-one fewer. All but two of those were withdrawn by the checks above; the other two are near-duplicates the site's own builder skips. Separately, seventeen questions whose evidence does not display properly have been held off the site by that builder since before the last post, and one ACP-520 batch of eighteen was held this morning by a stale verification receipt, until we noticed and re-ran it.
I want to be straight about the scale. There are 2,606 live questions. Each one carries the documentation sentence that decides it, verbatim, and an explanation for every option, and each has now been through five rounds. That is a massive amount of work and mistakes happen inside it. This class of mistake was found the same way as the one before it, by a person who sat the questions, not by a check. All but a handful of the 2,458 claims the fifth check could not reach are still sitting on screens we cannot see without an admin logged in, and some of them will be wrong.
If you hit one, it is not you.
Atlassian's own material is not always clear either
Keep this in mind when a practice question feels off. I say it from sitting the things. Gabriela holds ACP-120; I hold ACP-120, ACP-220, ACP-420 and ACP-520. The screenshots I keep of the mocks are often blurry, and our rule when reading one back is to name the word that cannot be read instead of guessing it.
Blur is the small problem. On the ACP-120 practice paper that Atlassian itself publishes, I scored 58 out of 75 answering strictly from the documentation. The results page died before it named the misses. Three separate passes afterwards could not find seventeen wrong answers, because judged against the documentation they were not wrong. Then we found ten places where Atlassian's course teaches something Atlassian's own documentation contradicts. The course says filters that reference a renamed project need updating. The documentation says the new name updates automatically in any filter. On those questions the key follows the course, and knowing the documentation is what costs you the mark. In fairness, one of the ten turned out to be us being wrong, not them.
The ACP-520 practice paper opens with what happens when Guard is enabled. No support page says Guard adds anything. The learning lesson the paper comes from says it adds a default authentication policy. On ACP-220, one question about exporting a single page to PDF sits on two Atlassian pages that disagree with each other, and I answered another one by insisting a setting did not exist, when the live page names it. The second read of ACP-220 withdrew one of our own questions for the same reason: two current support pages disagree on whether a nested group can be deleted from the content tree, so there is no fair question there. Two neighbouring pages on trashing a company-managed project disagree on who is allowed to do it.
So when a question feels wrong, sometimes it is us, sometimes the page, and now and then the page is arguing with itself. Our rule for that last case is to withdraw the question rather than take a side. Yours, on the real paper, has to be the opposite: answer what the course teaches.
How to reach us
The contact page says two people answer, usually the same day. Still true, and the two people are the ones named above. Email office@leanzero.net, mihai@leanzero.net or gabriela@leanzero.net, or post it on our Discord, which is faster. Say which exam, paste the stem or attach a screenshot, and say what you think is wrong. You do not have to be sure. "This feels off" is enough for us to go and run it, and running it is the one thing those four rounds never did.
A report gets the same treatment as everything else: read cold, tried on a live site, then the fix or the withdrawal goes out with the next push. The exams stay free, no account, no password, and a coffee is never expected.
And if you have sat the real thing lately, I would honestly like to know which question made you stare at the screen, and whether it was them or us.
I passed the Claude Certified Architect and ACP-120 within days of each other. One of them changed how I build agentic workflows, and not for the reason I expected. Here is the honest version, and what I am building next.