# Which Mac for local LLMs? Models and speed (tok/s) for every Mac

> Every Mac on sale, from the MacBook Neo to the Mac Studio: which open models fit in memory (gpt-oss, Qwen, Gemma, Llama, DeepSeek...) and how many tokens per second they write and read. Free.

- Saved simulation
- Name: "GLM-5.3-flash on a Mac Studio"
- Interactive version: https://simulkit.com/en/mac-local-llm/6MwixMWu8m
- Parameters, URL keys and method: https://simulkit.com/en/mac-local-llm.md
- Currency: USD

Which open models can your Mac run, and how fast? Pick a Mac from Apple's current lineup, from the MacBook Neo to the Mac Studio, and a model: the simulator checks that it fits in memory and estimates how many tokens per second it writes and reads. Every other model and every other Mac are compared too. These are estimates, calibrated on published benchmarks.

## Scenarios

| Parameter | flash (active) | standard |
|---|---|---|
| Mac | Mac Studio | Mac Studio |
| Chip | M5 Ultra, 36-core CPU, 80-core GPU | M5 Ultra, 36-core CPU, 80-core GPU |
| Unified memory | 256 GB | 512 GB |
| Model | GLM-5.3-Flash | GLM-5.3 |
| Context | 131,072 tokens | 131,072 tokens |

Other parameters, at their default values:

- Quantization: 4-bit
- Runtime: MLX
- Raise the GPU memory limit: No

## Results

| Result | flash (active) | standard |
|---|---|---|
| Writing speed | 42 tok/s | 22 tok/s |
| Price | $10,799 | – |
| Writing at full context | 40 tok/s | 21 tok/s |
| Reading speed | 875 tok/s | 462 tok/s |
| Time to read the context | 2 min 31 s | 4 min 52 s |
| Memory needed | 207 GB | 432 GB |
| Memory usable by the GPU | 239 GB | 498 GB |
| Models that fit | 35 | 38 |
| Memory bandwidth | 1,228 GB/s | 1,228 GB/s |

## Models on the Mac Studio (M5 Ultra, 36-core CPU, 80-core GPU) with 256 GB

In 4-bit, with a 131,072-token context (or the model's maximum when lower). The GPU can use 239 GB. Estimates.

| Model | Parameters | Memory needed | Writing | Reading | Verdict |
|---|---|---|---|---|---|
| Llama 3.2 1B | 1.2B | 6.1 GB | 193 tok/s | 31,480 tok/s | Fast |
| Qwen3.5 2B | 2.3B | 4.4 GB | 174 tok/s | 17,928 tok/s | Fast |
| Gemma 4 E2B | 5.1B | 5.5 GB | 173 tok/s | 17,073 tok/s | Fast |
| Llama 3.2 3B | 3.2B | 18 GB | 159 tok/s | 12,152 tok/s | Fast |
| SmolLM3 3B | 3.1B | 7.6 GB | 161 tok/s | 12,735 tok/s | Fast |
| Nemotron 3 Nano 4B | 4B | 5.5 GB | 149 tok/s | 10,256 tok/s | Fast |
| Qwen3.5 4B | 4.7B | 8.4 GB | 141 tok/s | 8,705 tok/s | Fast |
| Gemma 4 E4B | 8B | 8.3 GB | 143 tok/s | 8,901 tok/s | Fast |
| Llama 3.1 8B | 8B | 23 GB | 112 tok/s | 4,959 tok/s | Fast |
| Granite 4.2 8B | 8.8B | 27 GB | 107 tok/s | 4,509 tok/s | Fast |
| Qwen3.5 9B | 9.7B | 11 GB | 102 tok/s | 4,235 tok/s | Fast |
| Gemma 4 12B | 12B | 10 GB | 90 tok/s | 3,333 tok/s | Fast |
| Ministral 3 14B | 14B | 31 GB | 82 tok/s | 2,880 tok/s | Fast |
| Phi-4 | 14.7B | 13 GB | 79 tok/s | 2,729 tok/s | Fast |
| gpt-oss 20B | 20.9B, 3.6B active (MoE) | 16 GB | 163 tok/s | 9,450 tok/s | Fast |
| Devstral Small 2 24B | 24B | 38 GB | 56 tok/s | 1,691 tok/s | Fast |
| Gemma 4 26B A4B | 25.2B, 3.8B active (MoE) | 19 GB | 152 tok/s | 8,721 tok/s | Fast |
| Qwen3.8 27B | 27.8B | 26 GB | 50 tok/s | 1,470 tok/s | Fast |
| Muse Glimmer 30B | 29.6B | 22 GB | 50 tok/s | 1,457 tok/s | Fast |
| GLM-4.7-Flash | 30B, 3B active (MoE) | 25 GB | 61 tok/s | 10,217 tok/s | Fast |
| Nemotron 3.5 Lightning 30B A3B | 30B, 3B active (MoE) | 20 GB | 106 tok/s | 8,380 tok/s | Fast |
| Gemma 4 31B | 30.7B | 31 GB | 47 tok/s | 1,299 tok/s | Fast |
| Qwen3.6 35B A3B | 35B, 3B active (MoE) | 24 GB | 131 tok/s | 8,323 tok/s | Fast |
| Llama 3.3 70B | 70.6B | 84 GB | 23 tok/s | 572 tok/s | Comfortable |
| Qwen3-Coder-Next | 80B, 3B active (MoE) | 49 GB | 113 tok/s | 5,224 tok/s | Fast |
| gpt-oss 120B | 116.8B, 5.1B active (MoE) | 69 GB | 113 tok/s | 6,660 tok/s | Fast |
| Mistral Small 4 119B | 119.4B, 6.5B active (MoE) | 72 GB | 67 tok/s | 5,169 tok/s | Fast |
| Nemotron 3 Super 120B A12B | 120B, 12B active (MoE) | 70 GB | 50 tok/s | 1,320 tok/s | Fast |
| Qwen3.5 122B A10B | 122B, 10B active (MoE) | 74 GB | 79 tok/s | 2,509 tok/s | Fast |
| Mistral Medium 3.5 128B | 127.7B | 126 GB | 13 tok/s | 317 tok/s | Comfortable |
| Qwen3.8 Flash Next | 180B, 6B active (MoE) | 116 GB | 95 tok/s | 2,620 tok/s | Fast |
| MiniMax M2.7 | 228.7B, 10B active (MoE) | 163 GB | 67 tok/s | 1,839 tok/s | Fast |
| DeepSeek V4 Flash | 284B, 13B active (MoE) | 153 GB | 49 tok/s | 1,429 tok/s | Fast |
| GLM-5.3-Flash | 320B, 18B active (MoE) | 207 GB | 42 tok/s | 875 tok/s | Fast |
| Qwen3.5 397B A17B | 397B, 17B active (MoE) | 229 GB | 54 tok/s | 929 tok/s | Fast |
| MiniMax M3 | 428B, 23B active (MoE) | 266 GB | – | – | Does not fit |
| DeepSeek V3.2 | 671B, 37B active (MoE) | 390 GB | – | – | Does not fit |
| GLM-5.3 | 744B, 40B active (MoE) | 432 GB | – | – | Does not fit |
| Kimi K2.6 | 1,027B, 32B active (MoE) | 588 GB | – | – | Does not fit |

## GLM-5.3-Flash on every chip

| Chip | Macs | Smallest memory that fits | Writing | Reading | From |
|---|---|---|---|---|---|
| M5 Ultra, 36-core CPU, 80-core GPU | Mac Studio | 256 GB | 42 tok/s | 875 tok/s | $10,799 |
| M5 Ultra, 30-core CPU, 64-core GPU | Mac Studio | 256 GB | 42 tok/s | 700 tok/s | $9,499 |
| M5 Max, 18-core CPU, 40-core GPU | MacBook Pro 14-inch, MacBook Pro 16-inch, Mac Studio | Does not fit | – | – | – |
| M5 Max, 18-core CPU, 32-core GPU | MacBook Pro 14-inch, MacBook Pro 16-inch, Mac Studio | Does not fit | – | – | – |
| M5 Pro, 18-core CPU, 20-core GPU | MacBook Pro 14-inch, MacBook Pro 16-inch, Mac mini | Does not fit | – | – | – |
| M5 Pro, 15-core CPU, 16-core GPU | MacBook Pro 14-inch, Mac mini | Does not fit | – | – | – |
| M6, 12-core CPU, 12-core GPU | Mac mini | Does not fit | – | – | – |
| M5, 10-core CPU, 10-core GPU | MacBook Air 13-inch, MacBook Air 15-inch, MacBook Pro 14-inch | Does not fit | – | – | – |
| M5, 10-core CPU, 8-core GPU | MacBook Air 13-inch | Does not fit | – | – | – |
| M4, 10-core CPU, 10-core GPU | iMac 24-inch | Does not fit | – | – | – |
| M4, 8-core CPU, 8-core GPU | iMac 24-inch | Does not fit | – | – | – |
| A18 Pro, 6-core CPU, 5-core GPU | MacBook Neo 13-inch | Does not fit | – | – | – |

## Every Mac on sale

Starting prices: Apple Store United States, before sales tax.

| Mac | Chip | CPU | GPU | Memory bandwidth | Neural Engine | Memory | From |
|---|---|---|---|---|---|---|---|
| MacBook Neo 13-inch | A18 Pro | 6 cores | 5 cores | 60 GB/s | 16 cores | 8 GB | $699 |
| MacBook Air 13-inch | M5 | 10 cores | 8 cores | 153 GB/s | 16 cores | 16 GB | $1,299 |
| MacBook Air 13-inch | M5 | 10 cores | 10 cores | 153 GB/s | 16 cores | 16 GB, 24 GB, 32 GB | $1,399 |
| MacBook Air 15-inch | M5 | 10 cores | 10 cores | 153 GB/s | 16 cores | 16 GB, 24 GB, 32 GB | $1,499 |
| MacBook Pro 14-inch | M5 | 10 cores | 10 cores | 153 GB/s | 16 cores | 16 GB, 24 GB, 32 GB | $1,999 |
| MacBook Pro 14-inch | M5 Pro | 15 cores | 16 cores | 307 GB/s | 16 cores | 24 GB, 48 GB | $2,499 |
| MacBook Pro 14-inch | M5 Pro | 18 cores | 20 cores | 307 GB/s | 16 cores | 24 GB, 48 GB, 64 GB | $2,699 |
| MacBook Pro 14-inch | M5 Max | 18 cores | 32 cores | 460 GB/s | 16 cores | 36 GB | $4,099 |
| MacBook Pro 14-inch | M5 Max | 18 cores | 40 cores | 614 GB/s | 16 cores | 48 GB, 64 GB, 128 GB | $4,699 |
| MacBook Pro 16-inch | M5 Pro | 18 cores | 20 cores | 307 GB/s | 16 cores | 24 GB, 48 GB, 64 GB | $2,999 |
| MacBook Pro 16-inch | M5 Max | 18 cores | 32 cores | 460 GB/s | 16 cores | 36 GB | $4,399 |
| MacBook Pro 16-inch | M5 Max | 18 cores | 40 cores | 614 GB/s | 16 cores | 48 GB, 64 GB, 128 GB | $4,999 |
| iMac 24-inch | M4 | 8 cores | 8 cores | 120 GB/s | 16 cores | 16 GB, 24 GB | $1,499 |
| iMac 24-inch | M4 | 10 cores | 10 cores | 120 GB/s | 16 cores | 16 GB, 24 GB | $1,699 |
| Mac mini | M6 | 12 cores | 12 cores | 153 to 170 GB/s | 32 cores | 16 GB, 24 GB, 32 GB | $899 |
| Mac mini | M5 Pro | 15 cores | 16 cores | 307 GB/s | 16 cores | 24 GB, 48 GB, 64 GB | $1,699 |
| Mac mini | M5 Pro | 18 cores | 20 cores | 307 GB/s | 16 cores | 24 GB, 48 GB, 64 GB | $1,899 |
| Mac Studio | M5 Max | 18 cores | 32 cores | 460 GB/s | 16 cores | 36 GB | $2,499 |
| Mac Studio | M5 Max | 18 cores | 40 cores | 614 GB/s | 16 cores | 48 GB, 64 GB, 128 GB | $3,099 |
| Mac Studio | M5 Ultra | 30 cores | 64 cores | 1,228 GB/s | 32 cores | 96 GB, 256 GB | $5,499 |
| Mac Studio | M5 Ultra | 36 cores | 80 cores | 1,228 GB/s | 32 cores | 96 GB, 256 GB, 512 GB | $6,799 |

## Writing speed is set by the memory bandwidth

To write each token, the Mac reads every active weight of the model once. So the writing speed (decode) is about the memory bandwidth divided by the size of those weights. An M5 Max reads 614 GB per second: a dense 70B model at 4 bits weighs about 40 GB, so it writes about 13 tokens per second.

The simulator uses 75 to 87% of the bandwidth with MLX depending on the quantization (5 points less with llama.cpp), plus a fixed cost per token measured on each kind of chip and model. A long conversation slows things down: every new token also reads the KV cache, the model's memory of the context.

## Reading speed is set by the GPU

Before answering, the model reads your prompt (prefill). That step is limited by compute: about 2 operations per active parameter and per token, plus attention, which grows with the square of the prompt length.

M5 and M6 chips have a Neural Accelerator in each GPU core. MLX uses them (macOS 26.2 and later): prompts are read 3.5 to 4 times faster than on M4. llama.cpp uses them since April 2026 for matrix products, but not yet for attention.

## What decides if a model fits

The weights, the KV cache and the runtime's buffers must fit in the memory macOS lets the GPU use. On macOS 26 and 27 it is 74% of 16 or 24 GB, 78% of 32 to 48 GB, 81% of 64 or 96 GB, 84% of 128 GB, 87% of 256 GB and 464 GB of 512 GB. The command sudo sysctl iogpu.wired_limit_mb raises it until the next restart.

Apple's memory sizes are binary: 16 GB is 17.2 billion bytes. Model sizes here are in billions of bytes, like on Hugging Face.

## Quantization and mixtures of experts

Quantization stores each weight in fewer bits. 4-bit is the usual choice: the model is 3.5 times smaller than in 16 bits and loses little quality. 8-bit is almost lossless. A bigger model in 4-bit usually beats a smaller one in 8-bit.

A mixture of experts (MoE) like gpt-oss, Qwen3.6 35B A3B or DeepSeek only uses a few experts per token. All of its weights must fit in memory, but each token only reads the active ones: gpt-oss 120B needs 65 GB and writes like a 5B model.

## How reliable the numbers are

The constants come from about 300 published measurements: the llama.cpp benchmark thread, MLX and oMLX results, Apple's own figures and reviews of the M5 and M6 Macs. The estimates are usually within 15% of a measurement on the same software. Your speed depends on the app, its version and its settings: Ollama, LM Studio and llama.cpp do not all use the same engine.

Prices are those of the Apple Store of the chosen currency on October 1st, 2026, with base storage. Apple raised them on June 25th, 2026. French MacBooks, and British MacBook Air and Pro, come without a power adapter. The 512 GB Mac Studio ships in late October 2026: its price is not published yet.

## Frequently asked questions

### Which Mac do I need for gpt-oss 120B?

More than 64 GB of memory: it needs about 66 GB and 64 GB Macs let the GPU use 56 GB. The cheapest is the Mac Studio M5 Max with 128 GB ($5,099), at about 89 tokens per second. The MacBook Pro M5 Max with 128 GB is as fast. The Mac Studio M5 Ultra with 96 GB ($5,499) writes about 113.

### How much memory does a 70B model need?

Llama 3.3 70B in 4-bit needs about 40 GB of weights, plus 11 GB of KV cache for a 32,768-token context. A 64 GB Mac runs it, slowly: about 6 tokens per second on an M5 Pro, 13 on an M5 Max. The M5 Ultra (96 GB at least) writes about 23.

### Can a MacBook Neo or a MacBook Air run a local LLM?

Yes, small ones. The MacBook Neo has 8 GB: models up to 4B parameters, like Qwen3.5 4B at about 19 tokens per second. A 16 GB MacBook Air runs models up to 14B (Phi-4 at about 15 tokens per second, Llama 3.1 8B at about 28). They have no fan: long prompts slow down as the chip heats up.

### MLX or llama.cpp (Ollama, LM Studio)?

On a Mac, MLX is usually a little faster. It is much faster on hybrid models like Qwen3.5 and Qwen3.6 (1.4 to 1.9 times), and it reads prompts faster on M5 and M6. LM Studio runs both. Ollama runs some models on MLX and the others on llama.cpp.

### Why does a 35B model run faster than an 8B one?

Because it is a mixture of experts. Qwen3.6 35B A3B only uses 3 billion parameters per token: with MLX it writes faster than a dense 8B model, but its 35 billion parameters must all fit in memory.

### Can I run DeepSeek or Kimi on a Mac?

DeepSeek V4 Flash, yes: in 4-bit it needs about 150 GB and runs on a 256 GB Mac Studio M5 Ultra, at about 48 tokens per second. DeepSeek V3.2 and Kimi need the 512 GB Mac Studio M5 Ultra, which ships in late October 2026. DeepSeek V3.2 in 4-bit needs about 380 GB and writes about 25 tokens per second. Kimi K2.6 (1,000 billion parameters) only fits in 3-bit.

### Are my settings sent anywhere?

No. The math runs in your browser. Nothing is sent until you click Save or Share.

The URL keys, default values and ranges of every parameter are described in https://simulkit.com/en/mac-local-llm.md
