How to Run an LLM Locally: A Beginner's Guide
You can run open-weight AI models like Llama, Qwen, Gemma and Mistral on your own computer, for free and offline. What hardware you need, the easiest tools (Ollama, LM Studio), how model size and quantization work, and what local models can and can't do compared with Claude or GPT.
You don't need a cloud API to use an AI model. Open-weight models — ones whose files anyone can download — run on an ordinary laptop or desktop. It's free after the hardware, works offline, and your data never leaves your machine.
Why run one locally?
- Privacy — nothing is sent to a third party.
- Cost — no per-token bill.
- Offline — works on a plane.
- Learning — you see how models actually behave.
The catch: local models are smaller and less capable than the best cloud models like Claude or GPT. Great for many tasks; not a replacement for frontier models on hard coding or reasoning.
What hardware do you need?
The main limit is memory. The model has to fit in your graphics card's memory (VRAM) or, on Apple Silicon Macs, in unified memory.
Rough guide for quantized models (explained below):
| Model size | Memory needed | Runs on |
|---|---|---|
| 1–4B parameters | 4–6 GB | Almost any modern laptop |
| 7–9B | 8–10 GB | 16 GB Mac, or a GPU with 8+ GB |
| 12–14B | 12–16 GB | 24–32 GB Mac, or 12–16 GB GPU |
| 27–32B | 20–24 GB | 32 GB+ Mac, or 24 GB GPU |
| 70B+ | 40 GB+ | High-end Mac or multiple GPUs |
Without a good GPU it still works on the CPU — just slowly.
"Parameters" roughly measure a model's size. More parameters usually means smarter and slower.
Quantization in one paragraph
Model files store billions of numbers. Quantization stores them with less precision — for example 4 bits instead of 16 — making the file about four times smaller with only a small drop in quality. That's why a "7B" model fits in 8 GB. Names like Q4_K_M describe the quantization level; Q4 is the usual sweet spot.
The easiest tools
Ollama (command line, very simple)
Install from ollama.com, then:
ollama run llama3.2
It downloads the model and opens a chat. It also runs a local API your apps can call. (What is Ollama?)
LM Studio (desktop app)
A graphical app for browsing, downloading and chatting with models, with a built-in local server. Good if you prefer clicking to typing.
Others
llama.cpp (the engine underneath many tools), Jan, GPT4All, and vLLM for serving models on a GPU server.
Which model to start with?
Popular open-weight families include Meta's Llama, Alibaba's Qwen, Google's Gemma, Mistral, and DeepSeek. New versions appear every few months. Start with a small model (3–8B) that fits your memory, try your real task, and move up if it isn't good enough. Coding-specialised variants exist for most families.
What local models are good for
- Summarising, rewriting, classifying text
- Private chats about personal documents
- Simple code explanations and completions
- Experimenting with prompts and embeddings for free (What are embeddings?)
Where they fall short
- Complex multi-file coding and long agent tasks — frontier models are still far ahead
- Up-to-date knowledge (What is a knowledge cutoff?)
- Very long documents (smaller context windows, and long context uses lots of memory)
Using a local model from code
Ollama and LM Studio both offer an OpenAI-compatible API, so the same code that calls a cloud model can call your local one by changing the base URL:
const client = new OpenAI({ baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' })
The summary
- Open-weight models run locally for free, privately and offline.
- Memory decides what fits: 8 GB gets you a capable 7B model.
- Start with Ollama or LM Studio.
- Great for private, simple tasks; use frontier cloud models for hard coding.
EasySpawn runs your app and Claude Code on a persistent cloud server — useful when you want frontier-model coding help while your laptop handles local experiments. See how it works or join the waitlist.
Related: What Is Ollama? · What Is an LLM? · What Is OpenRouter? · Fine-Tuning vs RAG
Keep reading
What Is Ollama? Run AI Models on Your Own Computer
Ollama is a free tool for downloading and running open-weight AI models like Llama, Qwen and Gemma on your own machine, with a simple command line and a local API. How to install it, the commands you'll use, calling it from code, and running it on a server safely.
zsh vs bash: What's the Difference and Which Should You Use?
bash and zsh are both shells — the programs that read your terminal commands. Why macOS switched to zsh, the differences you'll actually notice (config files, completion, globbing, arrays), Oh My Zsh, and why scripts should usually still be written for bash.