Which Mac for local LLMs? Models and speed (tok/s) for every Mac

Saved simulation

Name: "GLM-5.3-flash on a Mac Studio"

Scenarios

Parameterflash (active)standard
MacMac StudioMac Studio
ChipM5 Ultra, 36-core CPU, 80-core GPUM5 Ultra, 36-core CPU, 80-core GPU
Unified memory256 GB512 GB
ModelGLM-5.3-FlashGLM-5.3
Context131,072 tokens131,072 tokens

Results

Resultflash (active)standard
Writing speed42 tok/s22 tok/s
Price$10,799–
Writing at full context40 tok/s21 tok/s
Reading speed875 tok/s462 tok/s
Time to read the context2 min 31 s4 min 52 s
Memory needed207 GB432 GB
Memory usable by the GPU239 GB498 GB
Models that fit3538
Memory bandwidth1,228 GB/s1,228 GB/s

Interactive version · Markdown version