How it works
Writing speed is set by the memory bandwidth
To write each token, the Mac reads every active weight of the model once. So the writing speed (decode) is about the memory bandwidth divided by the size of those weights. An M5 Max reads 614 GB per second: a dense 70B model at 4 bits weighs about 40 GB, so it writes about 13 tokens per second.
The simulator uses 75 to 87% of the bandwidth with MLX depending on the quantization (5 points less with llama.cpp), plus a fixed cost per token measured on each kind of chip and model. A long conversation slows things down: every new token also reads the KV cache, the model's memory of the context.
Reading speed is set by the GPU
Before answering, the model reads your prompt (prefill). That step is limited by compute: about 2 operations per active parameter and per token, plus attention, which grows with the square of the prompt length.
M5 and M6 chips have a Neural Accelerator in each GPU core. MLX uses them (macOS 26.2 and later): prompts are read 3.5 to 4 times faster than on M4. llama.cpp uses them since April 2026 for matrix products, but not yet for attention.
What decides if a model fits
The weights, the KV cache and the runtime's buffers must fit in the memory macOS lets the GPU use. On macOS 26 and 27 it is 74% of 16 or 24 GB, 78% of 32 to 48 GB, 81% of 64 or 96 GB, 84% of 128 GB, 87% of 256 GB and 464 GB of 512 GB. The command sudo sysctl iogpu.wired_limit_mb raises it until the next restart.
Apple's memory sizes are binary: 16 GB is 17.2 billion bytes. Model sizes here are in billions of bytes, like on Hugging Face.
Quantization and mixtures of experts
Quantization stores each weight in fewer bits. 4-bit is the usual choice: the model is 3.5 times smaller than in 16 bits and loses little quality. 8-bit is almost lossless. A bigger model in 4-bit usually beats a smaller one in 8-bit.
A mixture of experts (MoE) like gpt-oss, Qwen3.6 35B A3B or DeepSeek only uses a few experts per token. All of its weights must fit in memory, but each token only reads the active ones: gpt-oss 120B needs 65 GB and writes like a 5B model.
How reliable the numbers are
The constants come from about 300 published measurements: the llama.cpp benchmark thread, MLX and oMLX results, Apple's own figures and reviews of the M5 and M6 Macs. The estimates are usually within 15% of a measurement on the same software. Your speed depends on the app, its version and its settings: Ollama, LM Studio and llama.cpp do not all use the same engine.
Prices are those of the Apple Store of the chosen currency on October 1st, 2026, with base storage. Apple raised them on June 25th, 2026. French MacBooks, and British MacBook Air and Pro, come without a power adapter. The 512 GB Mac Studio ships in late October 2026: its price is not published yet.