
ArticleAI CodingLocal AI
llama.cpp vs MLX on Qwen3.6-27B: MTP is 1.04x here, not 1.85x
Multi-token prediction is merged in llama.cpp and still an open PR in mlx-lm. I measured both on an M3 Ultra with the same model. Every default MTP setting was slower than no MTP at all, and the runtime that deletes the MTP head outright is still the fastest thing on the box.