What Is Ollama? Run AI Models on Your Own Computer
Ollama is a free tool for downloading and running open-weight AI models like Llama, Qwen and Gemma on your own machine, with a simple command line and a local API. How to install it, the commands you'll use, calling it from code, and running it on a server safely.
Ollama is a free, open-source tool that makes running AI models on your own computer about as easy as installing an app. One command downloads a model; another chats with it. It also runs a small API server so your own programs can use the model.
Install it
Download from ollama.com for macOS or Windows. On Linux:
curl -fsSL https://ollama.com/install.sh | sh
The commands you'll use
ollama run llama3.2 # download (first time) and start chatting
ollama pull qwen2.5-coder # download without chatting
ollama list # models you have
ollama ps # models currently loaded in memory
ollama rm llama3.2 # delete a model
Type /bye to leave a chat. Browse available models in the library on Ollama's website — each page lists sizes like :3b, :8b, :70b. Pick one that fits your memory. (How to run an LLM locally)
Using Ollama from code
When Ollama is running, it serves an API on http://localhost:11434. Its own API:
curl http://localhost:11434/api/generate -d '{
"model": "llama3.2",
"prompt": "Explain DNS in one sentence",
"stream": false
}'
It also offers an OpenAI-compatible endpoint at /v1, so OpenAI SDK code works by changing the base URL:
import OpenAI from 'openai'
const client = new OpenAI({ baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' })
const res = await client.chat.completions.create({
model: 'llama3.2',
messages: [{ role: 'user', content: 'Write a haiku about servers' }],
})
The API key is required by the SDK but ignored by Ollama.
Customising a model
A Modelfile lets you bake in a system prompt and settings:
FROM llama3.2
SYSTEM "You are a concise assistant for a bakery. Answer in two sentences."
PARAMETER temperature 0.3
ollama create bakery-bot -f Modelfile
ollama run bakery-bot
(What is a system prompt?, LLM temperature)
Embeddings too
Ollama can run embedding models for semantic search and RAG, free and locally. (What are embeddings?)
Running Ollama on a server
You can run Ollama on a cloud server so an app can use it. Two things to know:
- Speed depends on the hardware. Without a GPU, a server CPU runs small models slowly — fine for background jobs, frustrating for chat.
- Don't expose port 11434 to the internet. Ollama's API has no authentication by default. Anyone who can reach it can use your server's resources. Keep it on
127.0.0.1and have your app call it locally, or put it behind a reverse proxy with authentication. (UFW firewall basics, Reverse proxies explained)
Ollama vs cloud APIs
| Ollama (local) | Cloud API (Claude, GPT) | |
|---|---|---|
| Cost | Free after hardware | Per token |
| Privacy | Stays on your machine | Sent to the provider |
| Quality | Good for simple tasks | Best available |
| Speed | Depends on your hardware | Fast |
| Setup | Install + download | API key |
Many apps use both: a local model for cheap, private tasks and a cloud model for hard ones.
EasySpawn gives you a persistent server where your app, its database and Claude Code all live together — keep Ollama for local experiments and run production on infrastructure that's always on. See how it works or join the waitlist.
Related: How to Run an LLM Locally · What Is an LLM? · What Is OpenRouter? · What Is RAG?
Keep reading
How to Run an LLM Locally: A Beginner's Guide
You can run open-weight AI models like Llama, Qwen, Gemma and Mistral on your own computer, for free and offline. What hardware you need, the easiest tools (Ollama, LM Studio), how model size and quantization work, and what local models can and can't do compared with Claude or GPT.
zsh vs bash: What's the Difference and Which Should You Use?
bash and zsh are both shells — the programs that read your terminal commands. Why macOS switched to zsh, the differences you'll actually notice (config files, completion, globbing, arrays), Oh My Zsh, and why scripts should usually still be written for bash.