Blog
3 min read

How to Run an LLM Locally: A Beginner's Guide

You can run open-weight AI models like Llama, Qwen, Gemma and Mistral on your own computer, for free and offline. What hardware you need, the easiest tools (Ollama, LM Studio), how model size and quantization work, and what local models can and can't do compared with Claude or GPT.

You don't need a cloud API to use an AI model. Open-weight models — ones whose files anyone can download — run on an ordinary laptop or desktop. It's free after the hardware, works offline, and your data never leaves your machine.

Why run one locally?

  • Privacy — nothing is sent to a third party.
  • Cost — no per-token bill.
  • Offline — works on a plane.
  • Learning — you see how models actually behave.

The catch: local models are smaller and less capable than the best cloud models like Claude or GPT. Great for many tasks; not a replacement for frontier models on hard coding or reasoning.

What hardware do you need?

The main limit is memory. The model has to fit in your graphics card's memory (VRAM) or, on Apple Silicon Macs, in unified memory.

Rough guide for quantized models (explained below):

Model size Memory needed Runs on
1–4B parameters 4–6 GB Almost any modern laptop
7–9B 8–10 GB 16 GB Mac, or a GPU with 8+ GB
12–14B 12–16 GB 24–32 GB Mac, or 12–16 GB GPU
27–32B 20–24 GB 32 GB+ Mac, or 24 GB GPU
70B+ 40 GB+ High-end Mac or multiple GPUs

Without a good GPU it still works on the CPU — just slowly.

"Parameters" roughly measure a model's size. More parameters usually means smarter and slower.

Quantization in one paragraph

Model files store billions of numbers. Quantization stores them with less precision — for example 4 bits instead of 16 — making the file about four times smaller with only a small drop in quality. That's why a "7B" model fits in 8 GB. Names like Q4_K_M describe the quantization level; Q4 is the usual sweet spot.

The easiest tools

Ollama (command line, very simple)

Install from ollama.com, then:

ollama run llama3.2

It downloads the model and opens a chat. It also runs a local API your apps can call. (What is Ollama?)

LM Studio (desktop app)

A graphical app for browsing, downloading and chatting with models, with a built-in local server. Good if you prefer clicking to typing.

Others

llama.cpp (the engine underneath many tools), Jan, GPT4All, and vLLM for serving models on a GPU server.

Which model to start with?

Popular open-weight families include Meta's Llama, Alibaba's Qwen, Google's Gemma, Mistral, and DeepSeek. New versions appear every few months. Start with a small model (3–8B) that fits your memory, try your real task, and move up if it isn't good enough. Coding-specialised variants exist for most families.

What local models are good for

  • Summarising, rewriting, classifying text
  • Private chats about personal documents
  • Simple code explanations and completions
  • Experimenting with prompts and embeddings for free (What are embeddings?)

Where they fall short

  • Complex multi-file coding and long agent tasks — frontier models are still far ahead
  • Up-to-date knowledge (What is a knowledge cutoff?)
  • Very long documents (smaller context windows, and long context uses lots of memory)

Using a local model from code

Ollama and LM Studio both offer an OpenAI-compatible API, so the same code that calls a cloud model can call your local one by changing the base URL:

const client = new OpenAI({ baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' })

The summary

  • Open-weight models run locally for free, privately and offline.
  • Memory decides what fits: 8 GB gets you a capable 7B model.
  • Start with Ollama or LM Studio.
  • Great for private, simple tasks; use frontier cloud models for hard coding.

EasySpawn runs your app and Claude Code on a persistent cloud server — useful when you want frontier-model coding help while your laptop handles local experiments. See how it works or join the waitlist.

Related: What Is Ollama? · What Is an LLM? · What Is OpenRouter? · Fine-Tuning vs RAG

Keep reading