How to run Llama, Qwen and DeepSeek models on a Mac

By Tried AI Tools · Published

Short answer

Install Ollama, then run "ollama run llama3.1:8b", "ollama run qwen3.5:9b" or "ollama run deepseek-r1:8b". Pick the size by your Mac's memory, roughly 0.6 GB per billion parameters at 4-bit plus headroom. In LM Studio, search for the same model names and download a 4-bit GGUF or MLX version.

Llama, Qwen and DeepSeek are three of the most popular open model families. All three run on Apple Silicon with free tools. This guide uses Ollama because it takes one command; the LM Studio route is at the end.

Pick a size first

Your Mac's memory Largest comfortable model at 4-bit
8 GB 3B to 4B
16 GB 8B to 9B
32 to 36 GB 14B to 27B
64 GB 32B to 35B
128 GB 70B, or mixture-of-experts models up to ~120B

Details in how much memory you need.

Llama (Meta)

Meta's Llama 3.1 and 3.3 are widely supported and good general-purpose models.

ollama run llama3.2:3b     # 8 GB Macs
ollama run llama3.1:8b     # 16 GB Macs
ollama run llama3.3:70b    # 64 GB with a raised GPU limit, 128 GB recommended

Llama models are released under Meta's own license, which allows most personal and commercial use but has conditions. Read it before building a product on one.

Qwen (Alibaba)

Qwen models are strong at coding, math and many languages, and come in a wide range of sizes. Qwen 3.5 and 3.6 also understand images.

ollama run qwen3.5:4b      # 8 GB Macs
ollama run qwen3.5:9b      # 16 GB Macs
ollama run qwen3.5:27b     # 32 GB and up
ollama run qwen3.6:35b     # 48 GB and up; strong for coding

On Apple Silicon, try the MLX versions too (for example qwen3.6:35b-mlx), which are often faster.

DeepSeek

DeepSeek R1 is a reasoning model: it thinks step by step before answering, which helps on math and logic but makes replies slower. The smaller sizes are distilled versions that fit on a Mac.

ollama run deepseek-r1:8b     # 16 GB Macs
ollama run deepseek-r1:14b    # 24 GB and up
ollama run deepseek-r1:32b    # 48 GB and up

Using LM Studio instead

  1. Download LM Studio and open it.
  2. Open the search tab and type the model name, for example "Qwen3.5 9B".
  3. Choose a 4-bit download. On a Mac, MLX versions are usually the fastest; GGUF versions work everywhere.
  4. Load the model and start chatting.

LM Studio shows whether each download is likely to fit in your Mac's memory before you start.

Tips

Not sure which app to use? See Ollama vs LM Studio vs llama.cpp vs MLX. For the best picks on bigger Macs, see best local AI models for 64 GB and 128 GB Macs.

Frequently asked questions

Is DeepSeek safe to run locally?

Running the open weights locally means your prompts stay on your Mac and are not sent to DeepSeek or anyone else. The model files are data that the runner loads, not programs with network access. As with any model, check its answers on topics where accuracy matters.

Are the small DeepSeek R1 models the real DeepSeek?

Not quite. The 8B to 70B DeepSeek R1 sizes are distilled versions, smaller Qwen and Llama models trained on the full model's reasoning. The full DeepSeek R1 has 671B parameters and needs several hundred gigabytes of memory.