Notes from the work

One email when a tutorial or migration write-up goes live. Nothing else, and one click to leave.

LeanZero

Two people in Romania doing Atlassian migrations, Forge apps and practical AI work for teams that would rather talk to the person doing the job. Most of what we learn ends up on this site.

Services

  • Atlassian Migrations
  • Atlassian FastShift
  • Atlassian Maintenance
  • Forge App Development
  • AI Development Consultation

Topics

  • Jira
  • Jira Service Management
  • Confluence
  • Bitbucket
  • Atlassian Forge
  • Cloud Migration
  • Local AI
  • AI Coding
  • Certifications
  • Atlassian Team
  • All topics

Company

  • Blog
  • Tutorials
  • Contact

Community

  • Join Discord
  • Support this site

© 2026 LeanZero. All rights reserved.

Privacy PolicyTerms of Service24/7 SupportTrust CenterSecurityDPALegal notice
  1. Home
  2. Topics
  3. Local Ai
Topic hub

Local AI

Local models on Apple Silicon, benchmarked rather than quoted.

Read the articlesAll topics

Running a capable model on your own hardware is mostly an argument about memory — how much you actually have, given that 96GB of marketing is 77.76 GiB you can spend, and what you are willing to give up to fit inside it. The benchmarks here are run, not quoted: MLX against llama.cpp on one machine and one set of weights, MXFP8 against Q8 on the same model, speculative decoding measured instead of assumed. Every number arrives with the command that produced it and the hardware it ran on, because a tokens-per-second figure missing either one is not a measurement. Results that contradicted what I expected are left in. Those tend to be the useful ones.

Read in order

Local Model Benchmarks

7 parts

Measured runs on real hardware: runtimes, quantisation formats and the memory ceiling nobody quotes.

  1. 1Part 1 — ArticleMLX vs GGUF on Apple Silicon: Benchmarking the Same Local Model Two Ways12 min
  2. 2Part 2 — ArticleMXFP8 vs Q8: 10x the weight error, 1% the perplexity14 min
  3. 3Part 3 — ArticleLing 3.0 Flash won't load on a Mac Studio. Qwen3-Coder-Next does, at 73 tok/s.18 min
  4. 4Part 4 — Articlellama.cpp vs MLX on Qwen3.6-27B: MTP is 1.04x here, not 1.85x15 min
  5. 5Part 5 — Article96 GB is 77.76 GiB: the real memory ceiling on an M3 Ultra13 min
  6. 6Part 6 — ArticleQwen3-Coder-Next + MXFP8: The 128GB Local LLM That Runs Predictably10 min
  7. 7Part 7 — ArticleThe best 4-bit format in MLX is the one its own converter sets up to lose9 min

Articles

24
Claude Sonnet 5.5 vs Opus 5.5 and GPT-6.1 Sol: the generalist and the specialist
ArticleLocal AIAtlassian Forge

Claude Sonnet 5.5 vs Opus 5.5 and GPT-6.1 Sol: the generalist and the specialist

Claude Sonnet 5.5 edged Claude Opus 5.5 on our general-programming benchmark, 0.7984 to 0.7926, a margin I read as level. On an Atlassian Forge app, Opus won 0.9767 to 0.5508. Every check behind both results, GPT-6.1 Sol and Sol Pro on the same two boards, how the benchmarks grade a running app, and the full bill for fine-tuning our own 27B model.

Oct 5, 202645 min read
Goose Swarm 3.0: local models on every Mac, joined by LeanZero Link
ArticleLocal AI

Goose Swarm 3.0: local models on every Mac, joined by LeanZero Link

Goose Swarm 3.0 is our local-first build of the goose agent: an MLX engine inside the app, a router that spreads chat turns across local and cloud nodes, LeanZero Link to put your other Mac to work, recurring Agent Work desks, memory that recalls itself, and an OpenAI-compatible endpoint. I installed it on a Mac Studio and went through all of it, including the parts that did not work first time.

Sep 23, 202618 min read
v0.5 of our Atlassian Qwen3.8-27B: beats v0.4, now with GGUF for llama.cpp and Ollama
ArticleLocal AI

v0.5 of our Atlassian Qwen3.8-27B: beats v0.4, now with GGUF for llama.cpp and Ollama

v0.5 of the Qwen3.8-27B we fine-tune for Forge, Jira, Confluence and JSM is out. It beat v0.4 on every gate we measure, with no waiver needed this time, using a training method that fixed the one thing capacity kept breaking. GGUF versions for llama.cpp and Ollama, Q8_0 and Q6_K, are now published alongside the existing MLX release, parity proven across three separate measured links plus a served-model probe.

Sep 20, 202611 min read
Convert a Fine-Tuned MLX Model to GGUF: Proving Quantization Parity with KL Divergence
TutorialLocal AI

Convert a Fine-Tuned MLX Model to GGUF: Proving Quantization Parity with KL Divergence

Converting a fine-tuned MLX model to GGUF for llama.cpp is one merge script and one convert_hf_to_gguf.py away. Nothing in that path tells you the GGUF actually matches the model you fine-tuned. This is the three-link method we used to prove it for a 27B model, with the exact commands, the real KL-divergence numbers, and the flag-collision bug that silently broke the measurement the first time we ran it.

Sep 20, 202616 min read
LeanZero at Atlassian Team '26 Europe Amsterdam: no booth, here to talk Forge migrations
ArticleAtlassian ForgeLocal AI

LeanZero at Atlassian Team '26 Europe Amsterdam: no booth, here to talk Forge migrations

We're going to Atlassian Team '26 Europe in Amsterdam on 6-8 October, no sponsor booth, no speaking slot, just the two of us. Here's what we'd actually talk about: 60,000+ users migrated at a 100% delivery rate, six open-source migration toolkits anyone can take from GitHub, and thirteen published model weights on Hugging Face from training we've written up start to finish, adapters and all.

Sep 17, 20266 min read
How to fine-tune Qwen3.8-27B with LoRA on a Mac: MLX fine-tuning from scratch
TutorialLocal AIAtlassian Forge

How to fine-tune Qwen3.8-27B with LoRA on a Mac: MLX fine-tuning from scratch

Qwen3.8-27B and Qwen3.5-9B were taught Forge, Jira and Confluence on a single Mac Studio with 96 GB. This is the whole procedure with the seven scripts that do it, printed in full: a pinned MLX environment with two patches, an 8-bit base with its speculative-decoding head kept as a sidecar, a chat-JSONL dataset built from a folder of Markdown, a segmented trainer with full-state checkpoints, an adapter expander, a per-module merge into the 8-bit shards. The kit was run end to end for this page on the 9B, in 21 minutes, and its real output follows each step, including the one where the validation loss turns the wrong way.

Sep 15, 202638 min read
v0.4 of our Atlassian model: 21 of 25 apps compile, and the loop that was a diagram
ArticleLocal AIAtlassian Forge

v0.4 of our Atlassian model: 21 of 25 apps compile, and the loop that was a diagram

Three days after v0.3, the Qwen3.8-27B we teach Forge went from 12 of 25 compiling apps to 21 of 25. On the way our own release rule rejected it, correctly by its own text, over a 40-character window repeated 13 times at 32k context. The window was a line of box-drawing characters in an ASCII diagram. Here is what changed, what the metric got wrong, and what we changed so it cannot happen the same way again.

Sep 13, 202612 min read
I taught a 27B model to write Forge apps. As far as I can find, nobody had done that before
ArticleLocal AIAtlassian Forge

I taught a 27B model to write Forge apps. As far as I can find, nobody had done that before

Qwen3.8-27B knows almost nothing about Atlassian Forge: ask it for an app and it invents a manifest format and imports packages that do not exist. Over four days on one Mac Studio I trained it until 14 of 25 manifests passed Atlassian's own validator and 12 of 25 apps compiled, from zero. The weights are public. Here is what it is, what it is not, and why I think it is a first.

Sep 10, 202611 min read
lsof: command not found in a packaged Mac app, and the port reclaim that never ran
ArticleAI Codinggoose Local Edition

lsof: command not found in a packaged Mac app, and the port reclaim that never ran

Our port reclaim logged a warning every single time it ran in the shipped app, and nobody noticed for weeks — because the warning branch was the only branch that could execute there. macOS keeps lsof in /usr/sbin, and a packaged app's PATH does not have it.

Sep 5, 20268 min read
Building a scorer that can't flatter itself
TutorialLocal AIgoose Local Edition

Building a scorer that can't flatter itself

Every number this benchmark ever published was wrong at least once. The transferable part is not the scorer, it is the six disciplines that caught each lie: prove the lever moves before you A/B it, prove the grader in both directions, pre-register the falsifier, positive-control every zero, spell absence and failure differently, and fix the report when the score is right but unreadable.

Aug 26, 202625 min read
goose Local Edition: a benchmark that runs the app your agent built
ArticleLocal AIgoose Local Edition

goose Local Edition: a benchmark that runs the app your agent built

goose Local Edition is our fork of goose with a swarm engine and an execution-based scorer. The scorer boots the built application, drives it over HTTP and in a browser, and prices what a user would actually experience. On the current board our own fleet is last of seventeen at 0.0172 — here is why that number is the point.

Aug 25, 202618 min read
goose swarm: pytest | head -80 exits 0 when nothing ran, and pipefail only trades the lie
ArticleAI Codinggoose Local Edition

goose swarm: pytest | head -80 exits 0 when nothing ran, and pipefail only trades the lie

A worker in my local-model swarm ran pytest against a file that did not exist, twice, and finished the task green. The pipe had eaten the exit code. Measured on bash and zsh, on pytest and cargo — including why turning pipefail on just moves the lie to the other side.

Aug 18, 202614 min read
What 145 MCP Tools Cost Before the Model Reads Your Question
ArticleAI CodingLocal AI

What 145 MCP Tools Cost Before the Model Reads Your Question

MCP went stateless on 2026-07-28. I probed ten real servers: none implement it. Then I measured the thing that actually costs you — 145 tool definitions, 40,784 tokens, and the 30% that is one server repeating itself.

Aug 17, 202614 min read
96 GB is 77.76 GiB: the real memory ceiling on an M3 Ultra
ArticleAI CodingLocal AI

96 GB is 77.76 GiB: the real memory ceiling on an M3 Ultra

Metal will not give you the RAM on the box, and the number it does give is not the 75% everyone repeats. I measured the three ceilings on a 96 GB Mac Studio, then measured what modern hybrid-attention models actually spend against them — including a Gemma 4 cache that quietly holds three times its own sliding window.

Aug 13, 202613 min read
goose compaction: my 27B re-read the file 95% of the time — quoting it in the summary didn't help
ArticleAI Codinggoose Local Edition

goose compaction: my 27B re-read the file 95% of the time — quoting it in the summary didn't help

Our swarm workers kept re-reading files they had already read. Context compaction was summarising the tool output that held the file. Pasting the file into the summary does not fix it; returning the last turns verbatim does. Measured three ways on the same 27B.

Aug 11, 202612 min read
llama.cpp vs MLX on Qwen3.6-27B: MTP is 1.04x here, not 1.85x
ArticleAI CodingLocal AI

llama.cpp vs MLX on Qwen3.6-27B: MTP is 1.04x here, not 1.85x

Multi-token prediction is merged in llama.cpp and still an open PR in mlx-lm. I measured both on an M3 Ultra with the same model. Every default MTP setting was slower than no MTP at all, and the runtime that deletes the MTP head outright is still the fastest thing on the box.

Aug 10, 202615 min read
Ling 3.0 Flash won't load on a Mac Studio. Qwen3-Coder-Next does, at 73 tok/s.
ArticleAI CodingLocal AI

Ling 3.0 Flash won't load on a Mac Studio. Qwen3-Coder-Next does, at 73 tok/s.

Ant Group's 124B/5.1B-active hybrid-linear MoE hit Hugging Face on 2 August. The memory arithmetic fits a 96 GB Mac with room to spare, and mlx-lm still refuses it. I counted exactly which tensors block it — 385 of 62,237 — then measured what that sparsity actually buys on the models that do run.

Aug 5, 202618 min read
MXFP8 vs Q8: 10x the weight error, 1% the perplexity
ArticleAI CodingLocal AI

MXFP8 vs Q8: 10x the weight error, 1% the perplexity

I quantised real Qwen3-Coder weights both ways on an M3 Ultra. MXFP8 reconstructs them about 10x worse than 8-bit affine at identical size — and then costs only 1% perplexity end to end. Both numbers are true, and the gap between them is the interesting part.

Aug 1, 202614 min read
Making the goose swarm predictable: 602 commits, 100 levers, and three bugs I found writing this
ArticleAI Codinggoose Local Edition

Making the goose swarm predictable: 602 commits, 100 levers, and three bugs I found writing this

Three weeks after the swarm shipped its first honest builds, the work stopped being about making small local models smarter and became about making them stop lying to me. Here is what 602 commits bought, what the desktop looks like now, and the three defects I found in my own honesty machinery while writing this post.

Jul 20, 202614 min read
swarm-gym for goose: grading local models by running the code they write
TutorialAI Codinggoose Local Edition

swarm-gym for goose: grading local models by running the code they write

A hands-on walkthrough of swarm-gym — the harness that drives goose local-edition's swarm through real coding tasks and grades the result by running it. Set it up, run both modes, read every output, and tune the swarm from what you find.

Jul 3, 20265 min read
MLX vs GGUF on Apple Silicon: Benchmarking the Same Local Model Two Ways
ArticleAI Codinggoose Local Edition

MLX vs GGUF on Apple Silicon: Benchmarking the Same Local Model Two Ways

We built a harness that makes local coding agents produce real software, grades it by running it, and ran the same model as GGUF and MLX. Here's the harness, its modes and archetypes — and which build to run.

Jul 2, 202612 min read
Inside goose-swarm: How We Turned One Local Model Into a Self-Verifying Fleet
TutorialAI Codinggoose Local Edition

Inside goose-swarm: How We Turned One Local Model Into a Self-Verifying Fleet

The engineering teardown of goose local-edition: how the scheduler, the parallel planner, the CONTRACTS discipline, the model-free judge, and the post-run smoke/AST gates are actually implemented, why each exists — and the self-driving test harness that found the failure behind every one of them.

Jul 1, 20265 min read
The goose swarm, self-verifying: small local models shipping real software
ArticleAI Codinggoose Local Edition

The goose swarm, self-verifying: small local models shipping real software

We forked Block's goose into a multi-agent swarm that decomposes a hard app spec into a task DAG, runs it across three small local models on LM Studio, and makes them verify their own work by actually running it. Over six days and 300-plus commits, the last real limit stopped being the swarm's coordination and became the small model's raw coding ability — and even that ceiling moved.

Jul 1, 20265 min read
Qwen3-Coder-Next + MXFP8: The 128GB Local LLM That Runs Predictably
ArticleAI CodingLocal AI

Qwen3-Coder-Next + MXFP8: The 128GB Local LLM That Runs Predictably

A deep dive into Qwen3-Coder-Next, the 80B MoE model that brings powerful local development within reach. Learn why this 2026 release matters for developers who value control, predictability, and consistent results.

Feb 4, 202610 min read