How to Deploy a FastAPI App to Production
Running uvicorn main:app --reload is for development. A production FastAPI deploy needs worker processes, a process manager or container, a reverse proxy with HTTPS, proper settings, migrations, and health checks. Step-by-step options for a VPS and Docker.
FastAPI apps are easy to start — uvicorn main:app --reload and you're running. That command is for development. In production you need more processes, automatic restarts, HTTPS, and settings that don't leak debug information.
Here's a production setup, first on a plain server and then in Docker. (Choosing a framework? Flask vs Django vs FastAPI.)
The production stack
Internet → Caddy/Nginx (HTTPS, port 443) → Uvicorn workers (127.0.0.1:8000) → FastAPI app → Postgres
- Uvicorn is the ASGI server that runs your app.
- Multiple workers let you use more than one CPU core (Python processes are effectively single-core for CPU work).
- A process manager (systemd) or container runtime restarts it on crashes and boot.
- A reverse proxy terminates HTTPS and forwards to Uvicorn. (Reverse proxies explained)
Step 1: Make the app production-ready
Settings from environment variables, validated at startup with pydantic-settings:
# settings.py
from pydantic_settings import BaseSettings
class Settings(BaseSettings):
database_url: str
secret_key: str
environment: str = "production"
cors_origins: list[str] = []
settings = Settings() # fails fast if DATABASE_URL is missing
(Environment variables explained)
A health endpoint that checks the database:
@app.get("/healthz")
async def healthz():
async with engine.connect() as conn:
await conn.execute(text("SELECT 1"))
return {"status": "ok"}
Turn off the interactive docs in production if your API is private: FastAPI(docs_url=None, redoc_url=None) — or protect them.
CORS: list exact origins; never allow_origins=["*"] with credentials. (CORS errors explained)
Pin dependencies with a lock file (uv.lock, poetry.lock, or a pinned requirements.txt). (Python virtual environments)
Step 2 (option A): A plain server with systemd
On the server, in a virtual environment:
cd /srv/api
uv sync --frozen # or: python -m venv .venv && .venv/bin/pip install -r requirements.txt
Create /etc/systemd/system/api.service:
[Unit]
Description=FastAPI app
After=network.target
[Service]
User=deploy
WorkingDirectory=/srv/api
EnvironmentFile=/srv/api/.env
ExecStart=/srv/api/.venv/bin/uvicorn main:app --host 127.0.0.1 --port 8000 --workers 2 --proxy-headers
Restart=always
[Install]
WantedBy=multi-user.target
sudo systemctl enable --now api
curl http://127.0.0.1:8000/healthz
How many workers? Start with the number of CPU cores (for async apps, often 1–2 per core is plenty) and adjust by watching memory and latency. Each worker is a full copy of your app in memory. Some teams use Gunicorn as the process manager with Uvicorn workers (gunicorn -k uvicorn.workers.UvicornWorker); Uvicorn's own --workers is fine for most apps.
--proxy-headers makes FastAPI trust X-Forwarded-For/X-Forwarded-Proto from the proxy, so URLs and client IPs are correct. Only trust them from your proxy (--forwarded-allow-ips).
Step 2 (option B): Docker
FROM python:3.13-slim
WORKDIR /app
COPY pyproject.toml uv.lock ./
RUN pip install uv && uv sync --frozen --no-dev
COPY . .
RUN useradd --create-home app
USER app
EXPOSE 8000
CMD ["/app/.venv/bin/uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000", "--workers", "2", "--proxy-headers"]
Inside a container, listen on 0.0.0.0 (the container's interface), but publish the port only to localhost on the host: -p 127.0.0.1:8000:8000. (Reduce Docker image size, production Dockerfiles — the principles carry over.)
Step 3: HTTPS with a reverse proxy
With Caddy, the whole config:
api.yourdomain.com {
reverse_proxy 127.0.0.1:8000
}
Caddy obtains and renews the certificate. With Nginx, add certbot. (What is Nginx?) For streaming responses, disable proxy buffering on that route. (Streaming LLM responses)
Step 4: Database migrations
Run Alembic migrations as a deploy step, before the new code starts — not on app startup with multiple workers racing:
.venv/bin/alembic upgrade head
sudo systemctl restart api
(Database migrations explained)
Use a connection pool sized for workers × pool size staying under Postgres's connection limit. (Connection pooling)
Step 5: Background work
Don't run long tasks inside request handlers — BackgroundTasks is fine for small fire-and-forget work, but anything slow or important belongs in a proper queue with a separate worker process. (Background jobs)
Pre-launch checklist
- Settings from env, validated; no secrets in code
- Multiple workers under systemd or a container with restart policy
- Uvicorn bound to localhost; HTTPS via proxy
-
/healthzchecks the database - Docs disabled or protected; CORS restricted
- Migrations as a deploy step
- Structured logs and error monitoring (error monitoring)
- Backups of the database
The summary
- Production = Uvicorn workers + process manager/container + HTTPS proxy.
- Validate settings at startup; add a health check; lock down docs and CORS.
- Bind to localhost (or publish container ports to localhost only).
- Run migrations as a separate deploy step; move slow work to a queue.
EasySpawn runs Python APIs on a dedicated server with Postgres alongside, automatic HTTPS on your domain and daily backups — and Claude Code can set up and deploy the FastAPI app for you. See how it works or join the waitlist.
Related: Flask vs Django vs FastAPI · Python Virtual Environments · How to Secure a New VPS · PM2 vs systemd
Keep reading
How to Reduce Docker Image Size: From 1.5 GB to Under 200 MB
Big Docker images are slow to build, push, pull and deploy. The techniques that actually shrink them: slim base images, multi-stage builds, .dockerignore, production-only dependencies, layer ordering, cache cleanup — with Node.js and Python examples and a way to see what's taking space.
PM2 vs systemd: How to Keep a Node.js App Running on a Server
When you run node app.js over SSH and log out, your app dies. How PM2 and systemd keep it running, restart it after crashes and reboots, handle logs and environment variables — with working configs for both, and where Docker fits.