Local LLMs on a Mac

Which open models can your Mac run, and how fast? Pick a Mac from Apple's current lineup, from the MacBook Neo to the Mac Studio, and a model: the simulator checks that it fits in memory and estimates how many tokens per second it writes and reads. Every other model and every other Mac are compared too. These are estimates, calibrated on published benchmarks.

Parameters

Mac

Model

Usage

tokens

Gemma 4 26B A4B on the MacBook Pro 14-inch with 48 GB writes

83 tok/s

Much faster than you read. At the end of a 32,768-token context: 54 tok/s.

Reading speed
2,302 tok/s
Prompt processing, with the GPU's Neural Accelerators.
Time to read the context
17 s
A 32,768-token prompt, before the first word of the answer.
Writing at full context
54 tok/s
Once 32,768 tokens are in memory.
Memory needed
17 GB
40 GB usable by the GPU, out of 48 GB.
Memory
17 GB of 40 GB
  • Weights (4.5 bits per weight)15 GB
  • KV cache (32,768 tokens)0.9 GB
  • Runtime buffers1.1 GB

23 GB left for the GPU.

M5 Pro, 18-core CPU, 20-core GPU
CPU
18 cores (6 super + 12 performance)
GPU
20 cores
Memory bandwidth
307 GB/s
GPU compute (FP32)
8.3 TFLOPS
Neural Accelerators (FP16)
33 TFLOPS
Neural Engine
16 cores
Price
$3,299

Price: Apple Store United States, before sales tax, base storage. The currency picks the store.

The cheapest Macs for Gemma 4 26B A4B

In 4-bit, with a 32,768-token context. Prices: Apple Store United States, before sales tax.

Cheapest fast one (30 tok/s)
$1,099
Mac mini, M6 · 12 GPU, 24 GB
53 tok/s

Models on this Mac

23 of 39 models fit, in 4-bit, with a 32,768-token context (or the model's maximum when lower).

ModelMemoryWritingVerdict
Llama 3.2 1B
Meta · 1.2B
2.8 GB309 tok/sFast
Qwen3.5 2B
Alibaba · 2.3B
3.2 GB181 tok/sFast
Gemma 4 E2B
Google · 5.1B
4.9 GB179 tok/sFast
Llama 3.2 3B
Meta · 3.2B
6.6 GB132 tok/sFast
SmolLM3 3B
Hugging Face · 3.1B
5.2 GB137 tok/sFast
Nemotron 3 Nano 4B
NVIDIA · 4B
3.9 GB108 tok/sFast
Qwen3.5 4B
Alibaba · 4.7B
5.2 GB93 tok/sFast
Gemma 4 E4B
Google · 8B
6.7 GB96 tok/sFast
Llama 3.1 8B
Meta · 8B
9.9 GB55 tok/sFast
Granite 4.2 8B
IBM · 8.8B
11 GB50 tok/sFast
Qwen3.5 9B
Alibaba · 9.7B
8.2 GB46 tok/sFast
Gemma 4 12B
Google · 12B
8.6 GB37 tok/sFast
Ministral 3 14B
Mistral AI · 14B
15 GB32 tok/sFast
Phi-4
Microsoft · 14.7B
13 GB30 tok/sFast
gpt-oss 20B
OpenAI · 20.9B, 3.6B active (MoE)
14 GB77 tok/sFast
Devstral Small 2 24B
Mistral AI · 24B
22 GB19 tok/sComfortable
Gemma 4 26B A4B
Google · 25.2B, 3.8B active (MoE)
17 GB83 tok/sFast
Qwen3.8 27B
Alibaba · 27.8B
19 GB16 tok/sComfortable
Muse Glimmer 30B
Meta · 29.6B
21 GB16 tok/sComfortable
GLM-4.7-Flash
Z.ai · 30B, 3B active (MoE)
20 GB54 tok/sFast
Nemotron 3.5 Lightning 30B A3B
NVIDIA · 30B, 3B active (MoE)
19 GB78 tok/sFast
Gemma 4 31B
Google · 30.7B
23 GB15 tok/sComfortable
Qwen3.6 35B A3B
Alibaba · 35B, 3B active (MoE)
22 GB88 tok/sFast

Too big for this Mac: Llama 3.3 70B (52 GB), Qwen3-Coder-Next (47 GB), gpt-oss 120B (66 GB), Mistral Small 4 119B (70 GB), Nemotron 3 Super 120B A12B (70 GB), Qwen3.5 122B A10B (72 GB), Mistral Medium 3.5 128B (91 GB), Qwen3.8 Flash Next (114 GB), MiniMax M2.7 (138 GB), DeepSeek V4 Flash (153 GB), GLM-5.3-Flash (206 GB), Qwen3.5 397B A17B (226 GB), MiniMax M3 (249 GB), DeepSeek V3.2 (382 GB), GLM-5.3 (422 GB), Kimi K2.6 (581 GB).

How it works

Writing speed is set by the memory bandwidth

To write each token, the Mac reads every active weight of the model once. So the writing speed (decode) is about the memory bandwidth divided by the size of those weights. An M5 Max reads 614 GB per second: a dense 70B model at 4 bits weighs about 40 GB, so it writes about 13 tokens per second.

The simulator uses 75 to 87% of the bandwidth with MLX depending on the quantization (5 points less with llama.cpp), plus a fixed cost per token measured on each kind of chip and model. A long conversation slows things down: every new token also reads the KV cache, the model's memory of the context.

Reading speed is set by the GPU

Before answering, the model reads your prompt (prefill). That step is limited by compute: about 2 operations per active parameter and per token, plus attention, which grows with the square of the prompt length.

M5 and M6 chips have a Neural Accelerator in each GPU core. MLX uses them (macOS 26.2 and later): prompts are read 3.5 to 4 times faster than on M4. llama.cpp uses them since April 2026 for matrix products, but not yet for attention.

What decides if a model fits

The weights, the KV cache and the runtime's buffers must fit in the memory macOS lets the GPU use. On macOS 26 and 27 it is 74% of 16 or 24 GB, 78% of 32 to 48 GB, 81% of 64 or 96 GB, 84% of 128 GB, 87% of 256 GB and 464 GB of 512 GB. The command sudo sysctl iogpu.wired_limit_mb raises it until the next restart.

Apple's memory sizes are binary: 16 GB is 17.2 billion bytes. Model sizes here are in billions of bytes, like on Hugging Face.

Quantization and mixtures of experts

Quantization stores each weight in fewer bits. 4-bit is the usual choice: the model is 3.5 times smaller than in 16 bits and loses little quality. 8-bit is almost lossless. A bigger model in 4-bit usually beats a smaller one in 8-bit.

A mixture of experts (MoE) like gpt-oss, Qwen3.6 35B A3B or DeepSeek only uses a few experts per token. All of its weights must fit in memory, but each token only reads the active ones: gpt-oss 120B needs 65 GB and writes like a 5B model.

How reliable the numbers are

The constants come from about 300 published measurements: the llama.cpp benchmark thread, MLX and oMLX results, Apple's own figures and reviews of the M5 and M6 Macs. The estimates are usually within 15% of a measurement on the same software. Your speed depends on the app, its version and its settings: Ollama, LM Studio and llama.cpp do not all use the same engine.

Prices are those of the Apple Store of the chosen currency on October 1st, 2026, with base storage. Apple raised them on June 25th, 2026. French MacBooks, and British MacBook Air and Pro, come without a power adapter. The 512 GB Mac Studio ships in late October 2026: its price is not published yet.

Frequently asked questions

Which Mac do I need for gpt-oss 120B?

More than 64 GB of memory: it needs about 66 GB and 64 GB Macs let the GPU use 56 GB. The cheapest is the Mac Studio M5 Max with 128 GB ($5,099), at about 89 tokens per second. The MacBook Pro M5 Max with 128 GB is as fast. The Mac Studio M5 Ultra with 96 GB ($5,499) writes about 113.

How much memory does a 70B model need?

Llama 3.3 70B in 4-bit needs about 40 GB of weights, plus 11 GB of KV cache for a 32,768-token context. A 64 GB Mac runs it, slowly: about 6 tokens per second on an M5 Pro, 13 on an M5 Max. The M5 Ultra (96 GB at least) writes about 23.

Can a MacBook Neo or a MacBook Air run a local LLM?

Yes, small ones. The MacBook Neo has 8 GB: models up to 4B parameters, like Qwen3.5 4B at about 19 tokens per second. A 16 GB MacBook Air runs models up to 14B (Phi-4 at about 15 tokens per second, Llama 3.1 8B at about 28). They have no fan: long prompts slow down as the chip heats up.

MLX or llama.cpp (Ollama, LM Studio)?

On a Mac, MLX is usually a little faster. It is much faster on hybrid models like Qwen3.5 and Qwen3.6 (1.4 to 1.9 times), and it reads prompts faster on M5 and M6. LM Studio runs both. Ollama runs some models on MLX and the others on llama.cpp.

Why does a 35B model run faster than an 8B one?

Because it is a mixture of experts. Qwen3.6 35B A3B only uses 3 billion parameters per token: with MLX it writes faster than a dense 8B model, but its 35 billion parameters must all fit in memory.

Can I run DeepSeek or Kimi on a Mac?

DeepSeek V4 Flash, yes: in 4-bit it needs about 150 GB and runs on a 256 GB Mac Studio M5 Ultra, at about 48 tokens per second. DeepSeek V3.2 and Kimi need the 512 GB Mac Studio M5 Ultra, which ships in late October 2026. DeepSeek V3.2 in 4-bit needs about 380 GB and writes about 25 tokens per second. Kimi K2.6 (1,000 billion parameters) only fits in 3-bit.

Are my settings sent anywhere?

No. The math runs in your browser. Nothing is sent until you click Save or Share.

Other simulators