Blog
3 min read

What Is Ollama? Run AI Models on Your Own Computer

Ollama is a free tool for downloading and running open-weight AI models like Llama, Qwen and Gemma on your own machine, with a simple command line and a local API. How to install it, the commands you'll use, calling it from code, and running it on a server safely.

Ollama is a free, open-source tool that makes running AI models on your own computer about as easy as installing an app. One command downloads a model; another chats with it. It also runs a small API server so your own programs can use the model.

Install it

Download from ollama.com for macOS or Windows. On Linux:

curl -fsSL https://ollama.com/install.sh | sh

The commands you'll use

ollama run llama3.2        # download (first time) and start chatting
ollama pull qwen2.5-coder  # download without chatting
ollama list                # models you have
ollama ps                  # models currently loaded in memory
ollama rm llama3.2         # delete a model

Type /bye to leave a chat. Browse available models in the library on Ollama's website — each page lists sizes like :3b, :8b, :70b. Pick one that fits your memory. (How to run an LLM locally)

Using Ollama from code

When Ollama is running, it serves an API on http://localhost:11434. Its own API:

curl http://localhost:11434/api/generate -d '{
  "model": "llama3.2",
  "prompt": "Explain DNS in one sentence",
  "stream": false
}'

It also offers an OpenAI-compatible endpoint at /v1, so OpenAI SDK code works by changing the base URL:

import OpenAI from 'openai'

const client = new OpenAI({ baseURL: 'http://localhost:11434/v1', apiKey: 'ollama' })
const res = await client.chat.completions.create({
  model: 'llama3.2',
  messages: [{ role: 'user', content: 'Write a haiku about servers' }],
})

The API key is required by the SDK but ignored by Ollama.

Customising a model

A Modelfile lets you bake in a system prompt and settings:

FROM llama3.2
SYSTEM "You are a concise assistant for a bakery. Answer in two sentences."
PARAMETER temperature 0.3
ollama create bakery-bot -f Modelfile
ollama run bakery-bot

(What is a system prompt?, LLM temperature)

Embeddings too

Ollama can run embedding models for semantic search and RAG, free and locally. (What are embeddings?)

Running Ollama on a server

You can run Ollama on a cloud server so an app can use it. Two things to know:

  • Speed depends on the hardware. Without a GPU, a server CPU runs small models slowly — fine for background jobs, frustrating for chat.
  • Don't expose port 11434 to the internet. Ollama's API has no authentication by default. Anyone who can reach it can use your server's resources. Keep it on 127.0.0.1 and have your app call it locally, or put it behind a reverse proxy with authentication. (UFW firewall basics, Reverse proxies explained)

Ollama vs cloud APIs

Ollama (local) Cloud API (Claude, GPT)
Cost Free after hardware Per token
Privacy Stays on your machine Sent to the provider
Quality Good for simple tasks Best available
Speed Depends on your hardware Fast
Setup Install + download API key

Many apps use both: a local model for cheap, private tasks and a cloud model for hard ones.


EasySpawn gives you a persistent server where your app, its database and Claude Code all live together — keep Ollama for local experiments and run production on infrastructure that's always on. See how it works or join the waitlist.

Related: How to Run an LLM Locally · What Is an LLM? · What Is OpenRouter? · What Is RAG?

Keep reading