# Which Mac for local LLMs? Models and speed (tok/s) for every Mac

> Every Mac on sale, from the MacBook Neo to the Mac Studio: which open models fit in memory (gpt-oss, Qwen, Gemma, Llama, DeepSeek...) and how many tokens per second they write and read. Free.

- Interactive version: https://simulkit.com/en/mac-local-llm
- Other languages: [Français](https://simulkit.com/fr/mac-llm-local.md)

Which open models can your Mac run, and how fast? Pick a Mac from Apple's current lineup, from the MacBook Neo to the Mac Studio, and a model: the simulator checks that it fits in memory and estimates how many tokens per second it writes and reads. Every other model and every other Mac are compared too. These are estimates, calibrated on published benchmarks.

## Parameters

### Mac

| Parameter | URL key | Default | Range | Description |
|---|---|---|---|---|
| Mac | `mac` | MacBook Pro 14-inch | `macbook-neo` MacBook Neo 13-inch, `macbook-air-13` MacBook Air 13-inch, `macbook-air-15` MacBook Air 15-inch, `macbook-pro-14` MacBook Pro 14-inch, `macbook-pro-16` MacBook Pro 16-inch, `imac` iMac 24-inch, `mac-mini` Mac mini, `mac-studio` Mac Studio | Every Mac Apple sells new today. |
| Chip | `chip` | M5 Pro, 18-core CPU, 20-core GPU | `a18-pro` A18 Pro, 6-core CPU, 5-core GPU, `m4-8-8` M4, 8-core CPU, 8-core GPU, `m4-10-10` M4, 10-core CPU, 10-core GPU, `m5-10-8` M5, 10-core CPU, 8-core GPU, `m5-10-10` M5, 10-core CPU, 10-core GPU, `m6-12-12` M6, 12-core CPU, 12-core GPU, `m5-pro-15-16` M5 Pro, 15-core CPU, 16-core GPU, `m5-pro-18-20` M5 Pro, 18-core CPU, 20-core GPU, `m5-max-18-32` M5 Max, 18-core CPU, 32-core GPU, `m5-max-18-40` M5 Max, 18-core CPU, 40-core GPU, `m5-ultra-30-64` M5 Ultra, 30-core CPU, 64-core GPU, `m5-ultra-36-80` M5 Ultra, 36-core CPU, 80-core GPU | The chips this Mac is sold with. The memory bandwidth sets the writing speed, the GPU sets the reading speed. |
| Unified memory | `memory` | 48 GB | `8` 8 GB, `16` 16 GB, `24` 24 GB, `32` 32 GB, `36` 36 GB, `48` 48 GB, `64` 64 GB, `96` 96 GB, `128` 128 GB, `256` 256 GB, `512` 512 GB | The memory sizes sold with this chip. It decides which models fit: it cannot be upgraded later. |

### Model

| Parameter | URL key | Default | Range | Description |
|---|---|---|---|---|
| Model | `model` | Gemma 4 26B A4B | `llama-3-2-1b` Llama 3.2 1B, `qwen3-5-2b` Qwen3.5 2B, `gemma-4-e2b` Gemma 4 E2B, `llama-3-2-3b` Llama 3.2 3B, `smollm3-3b` SmolLM3 3B, `nemotron-3-nano-4b` Nemotron 3 Nano 4B, `qwen3-5-4b` Qwen3.5 4B, `gemma-4-e4b` Gemma 4 E4B, `llama-3-1-8b` Llama 3.1 8B, `granite-4-2-8b` Granite 4.2 8B, `qwen3-5-9b` Qwen3.5 9B, `gemma-4-12b` Gemma 4 12B, `ministral-3-14b` Ministral 3 14B, `phi-4` Phi-4, `gpt-oss-20b` gpt-oss 20B, `devstral-small-2` Devstral Small 2 24B, `gemma-4-26b-a4b` Gemma 4 26B A4B, `qwen3-8-27b` Qwen3.8 27B, `muse-glimmer-30b` Muse Glimmer 30B, `glm-4-7-flash` GLM-4.7-Flash, `nemotron-3-5-lightning` Nemotron 3.5 Lightning 30B A3B, `gemma-4-31b` Gemma 4 31B, `qwen3-6-35b-a3b` Qwen3.6 35B A3B, `llama-3-3-70b` Llama 3.3 70B, `qwen3-coder-next` Qwen3-Coder-Next, `gpt-oss-120b` gpt-oss 120B, `mistral-small-4` Mistral Small 4 119B, `nemotron-3-super` Nemotron 3 Super 120B A12B, `qwen3-5-122b-a10b` Qwen3.5 122B A10B, `mistral-medium-3-5` Mistral Medium 3.5 128B, `qwen3-8-flash-next` Qwen3.8 Flash Next, `minimax-m2-7` MiniMax M2.7, `deepseek-v4-flash` DeepSeek V4 Flash, `glm-5-3-flash` GLM-5.3-Flash, `qwen3-5-397b-a17b` Qwen3.5 397B A17B, `minimax-m3` MiniMax M3, `deepseek-v3-2` DeepSeek V3.2, `glm-5-3` GLM-5.3, `kimi-k2-6` Kimi K2.6 | An open-weight model, run on the Mac with LM Studio, Ollama, llama.cpp or MLX. The table below shows every model. |
| Quantization | `quant` | 4-bit | `q3` 3-bit, `q4` 4-bit, `q6` 6-bit, `q8` 8-bit, `full` Original | Bits per weight. 4-bit is the usual choice: about 3.5 times smaller than the original, with a small loss of quality. 8-bit is almost lossless. 3-bit is for the biggest models only. |

### Usage

| Parameter | URL key | Default | Range | Description |
|---|---|---|---|---|
| Context | `context` | 32,768 tokens | 512 tokens to 1,048,576 tokens | Tokens the model keeps in memory: your prompt, documents, the conversation and the answer. A token is about three quarters of an English word. Capped at the model's maximum. |
| Runtime | `runtime` | MLX | `mlx` MLX, `llamacpp` llama.cpp | MLX is Apple's framework (LM Studio, mlx-lm). llama.cpp runs GGUF files (Ollama, LM Studio, llama.cpp). MLX is usually a little faster on a Mac and uses the Neural Accelerators of the M5 and M6 chips. |
| Raise the GPU memory limit | `raiseLimit` | false | `true`, `false` | By default macOS 26 and 27 let the GPU use 74 to 91% of the memory, depending on its size (2/3 of 8 GB). The command sudo sysctl iogpu.wired_limit_mb raises it until the next restart: the simulator leaves 10% to macOS, at least 4 GB. |

## Results with the default parameters

- Writing speed: 83 tok/s
- Price: $3,299
- Writing at full context: 54 tok/s
- Reading speed: 2,302 tok/s
- Time to read the context: 17 s
- Memory needed: 17 GB
- Memory usable by the GPU: 40 GB
- Models that fit: 23
- Memory bandwidth: 307 GB/s

## Models on the MacBook Pro 14-inch (M5 Pro, 18-core CPU, 20-core GPU) with 48 GB

In 4-bit, with a 32,768-token context (or the model's maximum when lower). The GPU can use 40 GB. Estimates.

| Model | Parameters | Memory needed | Writing | Reading | Verdict |
|---|---|---|---|---|---|
| Llama 3.2 1B | 1.2B | 2.8 GB | 309 tok/s | 10,160 tok/s | Fast |
| Qwen3.5 2B | 2.3B | 3.2 GB | 181 tok/s | 5,786 tok/s | Fast |
| Gemma 4 E2B | 5.1B | 4.9 GB | 179 tok/s | 5,510 tok/s | Fast |
| Llama 3.2 3B | 3.2B | 6.6 GB | 132 tok/s | 3,922 tok/s | Fast |
| SmolLM3 3B | 3.1B | 5.2 GB | 137 tok/s | 4,110 tok/s | Fast |
| Nemotron 3 Nano 4B | 4B | 3.9 GB | 108 tok/s | 3,310 tok/s | Fast |
| Qwen3.5 4B | 4.7B | 5.2 GB | 93 tok/s | 2,809 tok/s | Fast |
| Gemma 4 E4B | 8B | 6.7 GB | 96 tok/s | 2,873 tok/s | Fast |
| Llama 3.1 8B | 8B | 9.9 GB | 55 tok/s | 1,600 tok/s | Fast |
| Granite 4.2 8B | 8.8B | 11 GB | 50 tok/s | 1,455 tok/s | Fast |
| Qwen3.5 9B | 9.7B | 8.2 GB | 46 tok/s | 1,367 tok/s | Fast |
| Gemma 4 12B | 12B | 8.6 GB | 37 tok/s | 1,076 tok/s | Fast |
| Ministral 3 14B | 14B | 15 GB | 32 tok/s | 930 tok/s | Fast |
| Phi-4 | 14.7B | 13 GB | 30 tok/s | 881 tok/s | Fast |
| gpt-oss 20B | 20.9B, 3.6B active (MoE) | 14 GB | 77 tok/s | 2,483 tok/s | Fast |
| Devstral Small 2 24B | 24B | 22 GB | 19 tok/s | 546 tok/s | Comfortable |
| Gemma 4 26B A4B | 25.2B, 3.8B active (MoE) | 17 GB | 83 tok/s | 2,302 tok/s | Fast |
| Qwen3.8 27B | 27.8B | 19 GB | 16 tok/s | 474 tok/s | Comfortable |
| Muse Glimmer 30B | 29.6B | 21 GB | 16 tok/s | 470 tok/s | Comfortable |
| GLM-4.7-Flash | 30B, 3B active (MoE) | 20 GB | 54 tok/s | 2,734 tok/s | Fast |
| Nemotron 3.5 Lightning 30B A3B | 30B, 3B active (MoE) | 19 GB | 78 tok/s | 2,195 tok/s | Fast |
| Gemma 4 31B | 30.7B | 23 GB | 15 tok/s | 419 tok/s | Comfortable |
| Qwen3.6 35B A3B | 35B, 3B active (MoE) | 22 GB | 88 tok/s | 2,183 tok/s | Fast |
| Llama 3.3 70B | 70.6B | 52 GB | – | – | Does not fit |
| Qwen3-Coder-Next | 80B, 3B active (MoE) | 47 GB | – | – | Does not fit |
| gpt-oss 120B | 116.8B, 5.1B active (MoE) | 66 GB | – | – | Does not fit |
| Mistral Small 4 119B | 119.4B, 6.5B active (MoE) | 70 GB | – | – | Does not fit |
| Nemotron 3 Super 120B A12B | 120B, 12B active (MoE) | 70 GB | – | – | Does not fit |
| Qwen3.5 122B A10B | 122B, 10B active (MoE) | 72 GB | – | – | Does not fit |
| Mistral Medium 3.5 128B | 127.7B | 91 GB | – | – | Does not fit |
| Qwen3.8 Flash Next | 180B, 6B active (MoE) | 114 GB | – | – | Does not fit |
| MiniMax M2.7 | 228.7B, 10B active (MoE) | 138 GB | – | – | Does not fit |
| DeepSeek V4 Flash | 284B, 13B active (MoE) | 153 GB | – | – | Does not fit |
| GLM-5.3-Flash | 320B, 18B active (MoE) | 206 GB | – | – | Does not fit |
| Qwen3.5 397B A17B | 397B, 17B active (MoE) | 226 GB | – | – | Does not fit |
| MiniMax M3 | 428B, 23B active (MoE) | 249 GB | – | – | Does not fit |
| DeepSeek V3.2 | 671B, 37B active (MoE) | 382 GB | – | – | Does not fit |
| GLM-5.3 | 744B, 40B active (MoE) | 422 GB | – | – | Does not fit |
| Kimi K2.6 | 1,027B, 32B active (MoE) | 581 GB | – | – | Does not fit |

## Gemma 4 26B A4B on every chip

| Chip | Macs | Smallest memory that fits | Writing | Reading | From |
|---|---|---|---|---|---|
| M5 Ultra, 36-core CPU, 80-core GPU | Mac Studio | 96 GB | 152 tok/s | 8,721 tok/s | $6,799 |
| M5 Ultra, 30-core CPU, 64-core GPU | Mac Studio | 96 GB | 152 tok/s | 6,977 tok/s | $5,499 |
| M5 Max, 18-core CPU, 40-core GPU | MacBook Pro 14-inch, MacBook Pro 16-inch, Mac Studio | 48 GB | 128 tok/s | 4,589 tok/s | $3,099 |
| M5 Max, 18-core CPU, 32-core GPU | MacBook Pro 14-inch, MacBook Pro 16-inch, Mac Studio | 36 GB | 109 tok/s | 3,670 tok/s | $2,499 |
| M5 Pro, 18-core CPU, 20-core GPU | MacBook Pro 14-inch, MacBook Pro 16-inch, Mac mini | 24 GB | 83 tok/s | 2,302 tok/s | $1,899 |
| M5 Pro, 15-core CPU, 16-core GPU | MacBook Pro 14-inch, Mac mini | 24 GB | 83 tok/s | 1,838 tok/s | $1,699 |
| M6, 12-core CPU, 12-core GPU | Mac mini | 24 GB | 53 tok/s | 1,449 tok/s | $1,099 |
| M5, 10-core CPU, 10-core GPU | MacBook Air 13-inch, MacBook Air 15-inch, MacBook Pro 14-inch | 24 GB | 49 tok/s | 1,123 tok/s | $1,499 |
| M4, 10-core CPU, 10-core GPU | iMac 24-inch | 24 GB | 40 tok/s | 307 tok/s | $1,899 |
| M4, 8-core CPU, 8-core GPU | iMac 24-inch | 24 GB | 40 tok/s | 246 tok/s | $1,699 |
| M5, 10-core CPU, 8-core GPU | MacBook Air 13-inch | Does not fit | – | – | – |
| A18 Pro, 6-core CPU, 5-core GPU | MacBook Neo 13-inch | Does not fit | – | – | – |

## Every Mac on sale

Starting prices: Apple Store United States, before sales tax.

| Mac | Chip | CPU | GPU | Memory bandwidth | Neural Engine | Memory | From |
|---|---|---|---|---|---|---|---|
| MacBook Neo 13-inch | A18 Pro | 6 cores | 5 cores | 60 GB/s | 16 cores | 8 GB | $699 |
| MacBook Air 13-inch | M5 | 10 cores | 8 cores | 153 GB/s | 16 cores | 16 GB | $1,299 |
| MacBook Air 13-inch | M5 | 10 cores | 10 cores | 153 GB/s | 16 cores | 16 GB, 24 GB, 32 GB | $1,399 |
| MacBook Air 15-inch | M5 | 10 cores | 10 cores | 153 GB/s | 16 cores | 16 GB, 24 GB, 32 GB | $1,499 |
| MacBook Pro 14-inch | M5 | 10 cores | 10 cores | 153 GB/s | 16 cores | 16 GB, 24 GB, 32 GB | $1,999 |
| MacBook Pro 14-inch | M5 Pro | 15 cores | 16 cores | 307 GB/s | 16 cores | 24 GB, 48 GB | $2,499 |
| MacBook Pro 14-inch | M5 Pro | 18 cores | 20 cores | 307 GB/s | 16 cores | 24 GB, 48 GB, 64 GB | $2,699 |
| MacBook Pro 14-inch | M5 Max | 18 cores | 32 cores | 460 GB/s | 16 cores | 36 GB | $4,099 |
| MacBook Pro 14-inch | M5 Max | 18 cores | 40 cores | 614 GB/s | 16 cores | 48 GB, 64 GB, 128 GB | $4,699 |
| MacBook Pro 16-inch | M5 Pro | 18 cores | 20 cores | 307 GB/s | 16 cores | 24 GB, 48 GB, 64 GB | $2,999 |
| MacBook Pro 16-inch | M5 Max | 18 cores | 32 cores | 460 GB/s | 16 cores | 36 GB | $4,399 |
| MacBook Pro 16-inch | M5 Max | 18 cores | 40 cores | 614 GB/s | 16 cores | 48 GB, 64 GB, 128 GB | $4,999 |
| iMac 24-inch | M4 | 8 cores | 8 cores | 120 GB/s | 16 cores | 16 GB, 24 GB | $1,499 |
| iMac 24-inch | M4 | 10 cores | 10 cores | 120 GB/s | 16 cores | 16 GB, 24 GB | $1,699 |
| Mac mini | M6 | 12 cores | 12 cores | 153 to 170 GB/s | 32 cores | 16 GB, 24 GB, 32 GB | $899 |
| Mac mini | M5 Pro | 15 cores | 16 cores | 307 GB/s | 16 cores | 24 GB, 48 GB, 64 GB | $1,699 |
| Mac mini | M5 Pro | 18 cores | 20 cores | 307 GB/s | 16 cores | 24 GB, 48 GB, 64 GB | $1,899 |
| Mac Studio | M5 Max | 18 cores | 32 cores | 460 GB/s | 16 cores | 36 GB | $2,499 |
| Mac Studio | M5 Max | 18 cores | 40 cores | 614 GB/s | 16 cores | 48 GB, 64 GB, 128 GB | $3,099 |
| Mac Studio | M5 Ultra | 30 cores | 64 cores | 1,228 GB/s | 32 cores | 96 GB, 256 GB | $5,499 |
| Mac Studio | M5 Ultra | 36 cores | 80 cores | 1,228 GB/s | 32 cores | 96 GB, 256 GB, 512 GB | $6,799 |

## Writing speed is set by the memory bandwidth

To write each token, the Mac reads every active weight of the model once. So the writing speed (decode) is about the memory bandwidth divided by the size of those weights. An M5 Max reads 614 GB per second: a dense 70B model at 4 bits weighs about 40 GB, so it writes about 13 tokens per second.

The simulator uses 75 to 87% of the bandwidth with MLX depending on the quantization (5 points less with llama.cpp), plus a fixed cost per token measured on each kind of chip and model. A long conversation slows things down: every new token also reads the KV cache, the model's memory of the context.

## Reading speed is set by the GPU

Before answering, the model reads your prompt (prefill). That step is limited by compute: about 2 operations per active parameter and per token, plus attention, which grows with the square of the prompt length.

M5 and M6 chips have a Neural Accelerator in each GPU core. MLX uses them (macOS 26.2 and later): prompts are read 3.5 to 4 times faster than on M4. llama.cpp uses them since April 2026 for matrix products, but not yet for attention.

## What decides if a model fits

The weights, the KV cache and the runtime's buffers must fit in the memory macOS lets the GPU use. On macOS 26 and 27 it is 74% of 16 or 24 GB, 78% of 32 to 48 GB, 81% of 64 or 96 GB, 84% of 128 GB, 87% of 256 GB and 464 GB of 512 GB. The command sudo sysctl iogpu.wired_limit_mb raises it until the next restart.

Apple's memory sizes are binary: 16 GB is 17.2 billion bytes. Model sizes here are in billions of bytes, like on Hugging Face.

## Quantization and mixtures of experts

Quantization stores each weight in fewer bits. 4-bit is the usual choice: the model is 3.5 times smaller than in 16 bits and loses little quality. 8-bit is almost lossless. A bigger model in 4-bit usually beats a smaller one in 8-bit.

A mixture of experts (MoE) like gpt-oss, Qwen3.6 35B A3B or DeepSeek only uses a few experts per token. All of its weights must fit in memory, but each token only reads the active ones: gpt-oss 120B needs 65 GB and writes like a 5B model.

## How reliable the numbers are

The constants come from about 300 published measurements: the llama.cpp benchmark thread, MLX and oMLX results, Apple's own figures and reviews of the M5 and M6 Macs. The estimates are usually within 15% of a measurement on the same software. Your speed depends on the app, its version and its settings: Ollama, LM Studio and llama.cpp do not all use the same engine.

Prices are those of the Apple Store of the chosen currency on October 1st, 2026, with base storage. Apple raised them on June 25th, 2026. French MacBooks, and British MacBook Air and Pro, come without a power adapter. The 512 GB Mac Studio ships in late October 2026: its price is not published yet.

## Frequently asked questions

### Which Mac do I need for gpt-oss 120B?

More than 64 GB of memory: it needs about 66 GB and 64 GB Macs let the GPU use 56 GB. The cheapest is the Mac Studio M5 Max with 128 GB ($5,099), at about 89 tokens per second. The MacBook Pro M5 Max with 128 GB is as fast. The Mac Studio M5 Ultra with 96 GB ($5,499) writes about 113.

### How much memory does a 70B model need?

Llama 3.3 70B in 4-bit needs about 40 GB of weights, plus 11 GB of KV cache for a 32,768-token context. A 64 GB Mac runs it, slowly: about 6 tokens per second on an M5 Pro, 13 on an M5 Max. The M5 Ultra (96 GB at least) writes about 23.

### Can a MacBook Neo or a MacBook Air run a local LLM?

Yes, small ones. The MacBook Neo has 8 GB: models up to 4B parameters, like Qwen3.5 4B at about 19 tokens per second. A 16 GB MacBook Air runs models up to 14B (Phi-4 at about 15 tokens per second, Llama 3.1 8B at about 28). They have no fan: long prompts slow down as the chip heats up.

### MLX or llama.cpp (Ollama, LM Studio)?

On a Mac, MLX is usually a little faster. It is much faster on hybrid models like Qwen3.5 and Qwen3.6 (1.4 to 1.9 times), and it reads prompts faster on M5 and M6. LM Studio runs both. Ollama runs some models on MLX and the others on llama.cpp.

### Why does a 35B model run faster than an 8B one?

Because it is a mixture of experts. Qwen3.6 35B A3B only uses 3 billion parameters per token: with MLX it writes faster than a dense 8B model, but its 35 billion parameters must all fit in memory.

### Can I run DeepSeek or Kimi on a Mac?

DeepSeek V4 Flash, yes: in 4-bit it needs about 150 GB and runs on a 256 GB Mac Studio M5 Ultra, at about 48 tokens per second. DeepSeek V3.2 and Kimi need the 512 GB Mac Studio M5 Ultra, which ships in late October 2026. DeepSeek V3.2 in 4-bit needs about 380 GB and writes about 25 tokens per second. Kimi K2.6 (1,000 billion parameters) only fits in 3-bit.

### Are my settings sent anywhere?

No. The math runs in your browser. Nothing is sent until you click Save or Share.

## Open it with your own numbers

Every parameter can be passed in the query string, using the keys above. The same URL with ".md" (for example /en/simulator.md?key=value) returns the computed results for those parameters. A shared link (/en/simulator/<id>) has a .md version too; a link with "#s=" cannot be read by a server: ask for the Share link. For example:

https://simulkit.com/en/mac-local-llm?model=llama-3-2-1b&mac=macbook-neo
