Ollama
Running, managing, and serving local LLMs with Ollama — CLI, Modelfiles, and the REST/OpenAI-compatible API.
Ollama
Running, managing, and serving local LLMs with Ollama — CLI, Modelfiles, and the REST/OpenAI-compatible API.
Overview
Ollama wraps llama.cpp (and GGUF-format models) into a self-contained local server with a model registry, a CLI, and an HTTP API. It pulls quantised models from the Ollama library, manages their lifecycle in memory, and exposes both a native REST API and an OpenAI-compatible surface — making it a drop-in for code already written against the OpenAI SDK.
It runs on macOS, Linux, and Windows, and by default serves on 127.0.0.1:11434. GPU acceleration is automatic where available (NVIDIA CUDA, AMD ROCm, Apple Metal), with CPU fallback otherwise.
Ollama targets easy local and single-node development: pull a model, run it, point your client at localhost. For high-throughput production serving (continuous batching, tensor parallelism, paged attention), reach for vLLM instead — covered in its own sheet.
graph TB
subgraph "Clients"
A[ollama CLI]
B[REST /api/*]
C[OpenAI SDK /v1/*]
D[ollama python pkg]
end
subgraph "Ollama Server :11434"
E[HTTP API]
F[Model Scheduler]
G[llama.cpp Runner]
end
A --> E
B --> E
C --> E
D --> E
E --> F
F -->|load/unload| G
G -->|GGUF blobs| H[(~/.ollama/models)]
F -->|offload layers| I[GPU / CPU]
Install and Service
Install Methods
The two macOS routes install the same binaries but differ in lifecycle: the .dmg is Ollama's recommended path and keeps itself updated through the menu-bar app, whereas the Homebrew formula fits a scripted/dotfiles setup and updates with brew upgrade. Pick one — running both leaves you with two copies fighting over :11434.
# macOS (recommended) — download the .dmg app from ollama.com and drag to /Applications
# macOS (Homebrew) — CLI + server via the formula, for scripted setups
brew install ollama
# Linux — official install script (systemd service + ollama user)
curl -fsSL https://ollama.com/install.sh | sh
# Docker (CPU)
docker run -d -v ollama:/root/.ollama -p 11434:11434 \
--name ollama ollama/ollama
# Docker (NVIDIA GPU — requires nvidia-container-toolkit)
docker run -d --gpus=all -v ollama:/root/.ollama -p 11434:11434 \
--name ollama ollama/ollama
The Background Server
The CLI talks to a running server; commands like ollama run will fail with a connection error if nothing is listening on :11434.
# Run the server in the foreground (manual / dev use)
ollama serve
# On Linux, the install script registers a systemd service that
# starts ollama serve automatically as the 'ollama' user
systemctl status ollama
sudo systemctl restart ollama
journalctl -u ollama -f # follow server logs
# macOS app runs the server as a background agent automatically
GPU Support
| Backend | Platform | Notes |
|---|---|---|
| CUDA | NVIDIA on Linux/Windows | Needs recent driver; Docker needs nvidia-container-toolkit |
| ROCm | AMD on Linux | Supported GPU list is narrower; check official compatibility |
| Metal | Apple Silicon (macOS) | Automatic; unified memory means VRAM ≈ system RAM |
| CPU | Everywhere | Fallback when no GPU detected — much slower for large models |
Confirm what a loaded model is actually using with ollama ps — the PROCESSOR column shows the GPU/CPU split.
Core CLI
Running Models
# Interactive REPL (pulls the model first if not present)
ollama run llama3.1:8b
# One-shot: pass the prompt as an argument, get a single completion
ollama run llama3.1:8b "Summarise the CAP theorem in two sentences."
# Pipe input
cat report.txt | ollama run llama3.1:8b "Summarise this:"
# Multimodal — attach an image by path in the prompt (vision models only)
ollama run llava "Describe this image: ./diagram.png"
Managing Models
ollama pull qwen2.5:7b # download (or update) a model
ollama list # local models, sizes, modified time
ollama ps # running models + VRAM/CPU split + keep-alive
ollama stop llama3.1:8b # unload a running model from memory
ollama rm llama3.1:8b # delete a local model
ollama cp llama3.1:8b my-llama # copy/alias a model locally
ollama show llama3.1:8b # params, template, licence, modelfile
ollama show llama3.1:8b --modelfile # just the Modelfile
ollama create mymodel -f Modelfile # build a model from a Modelfile
ollama push user/mymodel # push to ollama.com (needs an account + key)
Interactive REPL Commands
Inside ollama run, lines beginning with / are commands rather than prompts:
>>> /set parameter temperature 0.2 # override a runtime parameter
>>> /set system "You are a terse assistant."
>>> /show info # model metadata
>>> /show parameters # active parameters
>>> /load mymodel # switch to another model
>>> /save mysession # save current session as a new model
>>> /bye # exit the REPL (Ctrl-D also works)
>>> """ # start a multi-line prompt
... line one
... line two
... """
Model Naming and Tags
Models are addressed as model:tag. Omitting the tag defaults to :latest.
llama3.1:8b # parameter-size tag
qwen2.5:7b-instruct-q4_K_M # size + variant + quantisation
gemma2 # bare name → resolves to :latest
The tag often encodes three things: parameter count (8b, 70b), variant (instruct, text), and quantisation (q4_K_M, q8_0, fp16).
Quantisation vs VRAM
Quantisation trades precision for memory. q4_K_M (4-bit, K-quant medium) is the common default sweet spot — roughly half a byte per weight plus overhead. Rough rule of thumb for a Q4 model: VRAM needed ≈ parameter count in billions × ~0.6–0.7 GB, plus the KV cache (which grows with num_ctx).
| Suffix | Bits | Trade-off |
|---|---|---|
q4_K_M |
~4 | Default sweet spot — good quality/size balance |
q5_K_M |
~5 | Slightly better quality, more VRAM |
q8_0 |
8 | Near-lossless, ~2× the size of Q4 |
fp16 / f16 |
16 | Full half-precision, largest footprint |
If a model won't fit, drop to a smaller quant or a smaller parameter size before reaching for CPU offload.
Modelfile
A Modelfile is the build recipe for a model: a base, runtime parameters, a system prompt, a prompt template, and metadata. Build with ollama create.
Instructions
| Instruction | Purpose |
|---|---|
FROM |
Base model (library name, another local model, or a ./path.gguf) |
PARAMETER |
Default runtime parameter (temperature, num_ctx, stop, …) |
SYSTEM |
Default system message |
TEMPLATE |
Go-template prompt format (.System, .Prompt, .Messages, .Tools) |
MESSAGE |
Seed a canned conversation turn (role + content) |
ADAPTER |
Apply a LoRA/QLoRA adapter on top of FROM |
LICENSE |
Embed licence text |
Common Parameters
| Parameter | Meaning |
|---|---|
temperature |
Sampling randomness (0 = deterministic-ish) |
top_p |
Nucleus sampling cutoff |
top_k |
Top-K sampling cutoff |
num_ctx |
Context window size in tokens (see caveat below) |
num_predict |
Max tokens to generate (-1 = until stop/EOS) |
stop |
Stop sequence (repeat the instruction for several) |
repeat_penalty |
Penalty against token repetition |
seed |
Fix the RNG seed for reproducible output |
Example
# Modelfile
FROM llama3.1:8b
PARAMETER temperature 0.2
PARAMETER top_p 0.9
PARAMETER num_ctx 8192
PARAMETER num_predict 1024
PARAMETER repeat_penalty 1.1
PARAMETER stop "<|eot_id|>"
PARAMETER stop "User:"
SYSTEM """
You are a senior SRE assistant. Answer tersely, in UK English.
Prefer concrete commands over prose.
"""
# Optional: a custom Go template (most base models ship a sensible one)
TEMPLATE """{{ if .System }}<|system|>{{ .System }}{{ end }}
<|user|>{{ .Prompt }}
<|assistant|>"""
LICENSE "Use restricted to internal tooling."
ollama create sre-assistant -f Modelfile
ollama run sre-assistant "Why might a pod be stuck in CrashLoopBackOff?"
Importing a Raw GGUF
Point FROM at a GGUF file downloaded from elsewhere (e.g. Hugging Face):
FROM ./mistral-7b-instruct-q4_k_m.gguf
PARAMETER num_ctx 4096
TEMPLATE """[INST] {{ .Prompt }} [/INST]"""
REST API
The native API lives under /api/* on :11434. This is the operational core. By default, streaming responses are newline-delimited JSON (NDJSON) — one object per line — terminated by a final object with "done": true carrying timing and token-count stats. Set "stream": false to get a single aggregated JSON object instead.
POST /api/generate
Single-prompt completion. Use /api/chat for multi-turn.
curl http://localhost:11434/api/generate -d '{
"model": "llama3.1:8b",
"prompt": "Explain idempotency in one sentence.",
"stream": false,
"options": { "temperature": 0.2, "num_ctx": 8192 },
"keep_alive": "10m"
}'
{
"model": "llama3.1:8b",
"created_at": "2026-06-15T10:00:00Z",
"response": "An idempotent operation produces the same result however many times it is applied.",
"done": true,
"total_duration": 1820000000,
"eval_count": 21,
"eval_duration": 900000000,
"context": [128006, 9125, 128007, ...]
}
Useful fields:
| Field | Purpose |
|---|---|
options |
Per-request overrides (temperature, num_ctx, seed, stop, …) |
format |
"json" for JSON mode, or a full JSON Schema object for structured output |
keep_alive |
How long to keep the model loaded after this request ("10m", -1, 0) |
raw |
true bypasses templating — you supply the fully-formatted prompt |
images |
Array of base64-encoded images for multimodal models |
context |
Opaque token array from a prior response, for stateless continuation |
suffix |
Text after the completion (fill-in-the-middle, for code models) |
POST /api/chat
Multi-turn with a messages array of {role, content}. Roles: system, user, assistant, tool.
curl http://localhost:11434/api/chat -d '{
"model": "llama3.1:8b",
"messages": [
{ "role": "system", "content": "You answer in UK English." },
{ "role": "user", "content": "What is a write-ahead log?" }
],
"stream": false
}'
Tools / function-calling (model must support it, e.g. Llama 3.1, Qwen 2.5):
curl http://localhost:11434/api/chat -d '{
"model": "llama3.1:8b",
"messages": [{ "role": "user", "content": "Weather in Kegworth?" }],
"stream": false,
"tools": [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a place",
"parameters": {
"type": "object",
"properties": { "location": { "type": "string" } },
"required": ["location"]
}
}
}]
}'
The model replies with message.tool_calls; you execute the call and feed the result back as a {"role": "tool", "content": "..."} message.
Structured output — pass a JSON Schema as format:
curl http://localhost:11434/api/chat -d '{
"model": "llama3.1:8b",
"messages": [{ "role": "user", "content": "Give me a UK city and its population." }],
"stream": false,
"format": {
"type": "object",
"properties": {
"city": { "type": "string" },
"population": { "type": "integer" }
},
"required": ["city", "population"]
}
}'
Embeddings
There are two endpoints. Prefer the newer /api/embed (batch via input); /api/embeddings is the older single-prompt form, kept for compatibility.
# Newer batch form — input can be a string or an array of strings
curl http://localhost:11434/api/embed -d '{
"model": "nomic-embed-text",
"input": ["first chunk", "second chunk"]
}'
# → { "embeddings": [[...], [...]], ... }
# Older single-prompt form
curl http://localhost:11434/api/embeddings -d '{
"model": "nomic-embed-text",
"prompt": "first chunk"
}'
# → { "embedding": [...] }
Note the key differences: /api/embed takes input and returns embeddings (array); /api/embeddings takes prompt and returns embedding (single).
Management Endpoints
curl http://localhost:11434/api/tags # GET — list local models
curl http://localhost:11434/api/ps # GET — running models
curl http://localhost:11434/api/show -d '{"model":"llama3.1:8b"}' # POST — details
curl http://localhost:11434/api/pull -d '{"model":"qwen2.5:7b"}' # POST — pull (streams progress)
curl -X DELETE http://localhost:11434/api/delete -d '{"model":"old-model"}'
curl http://localhost:11434/api/create -d '{
"model": "sre-assistant",
"from": "llama3.1:8b",
"system": "You are a terse SRE assistant."
}'
Streaming Format
sequenceDiagram
participant C as Client
participant O as Ollama
C->>O: POST /api/chat (stream: true)
O-->>C: {"message":{"content":"A"},"done":false}
O-->>C: {"message":{"content":" log"},"done":false}
O-->>C: {"message":{"content":"..."},"done":false}
O-->>C: {"done":true,"eval_count":42,"total_duration":...}
Each line is a complete JSON object. Accumulate message.content (chat) or response (generate) across lines; the final "done": true object carries no new text but reports eval_count, prompt_eval_count, total_duration, and friends.
OpenAI-Compatible Endpoint
Ollama exposes an OpenAI-compatible surface under /v1/* — a drop-in for existing OpenAI code. Point the SDK at the local base URL and pass any non-empty api_key (it is ignored but the SDK requires one).
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:11434/v1",
api_key="ollama", # required by the SDK, ignored by Ollama
)
resp = client.chat.completions.create(
model="llama3.1:8b",
messages=[{"role": "user", "content": "What is a deadletter queue?"}],
)
print(resp.choices[0].message.content)
# Streaming
for chunk in client.chat.completions.create(
model="llama3.1:8b",
messages=[{"role": "user", "content": "Count to five."}],
stream=True,
):
print(chunk.choices[0].delta.content or "", end="")
# Embeddings
emb = client.embeddings.create(
model="nomic-embed-text",
input=["chunk one", "chunk two"],
)
Supported routes: /v1/chat/completions, /v1/completions, /v1/embeddings, /v1/models (and /v1/models/{model}), plus the newer /v1/responses and an experimental /v1/images/generations. The native /api/* surface exposes more (e.g. per-request keep_alive, raw mode), so use it when you need Ollama-specific control.
Python
The first-party ollama package wraps the native API.
import ollama
# Chat
resp = ollama.chat(
model="llama3.1:8b",
messages=[{"role": "user", "content": "Define backpressure."}],
)
print(resp["message"]["content"])
# Generate
print(ollama.generate(model="llama3.1:8b", prompt="One word for 'fast':")["response"])
# Embeddings (batch)
vecs = ollama.embed(model="nomic-embed-text", input=["a", "b"])["embeddings"]
# Streaming
for chunk in ollama.chat(
model="llama3.1:8b",
messages=[{"role": "user", "content": "Count to three."}],
stream=True,
):
print(chunk["message"]["content"], end="", flush=True)
Async, and pointing at a non-default host:
import asyncio
from ollama import AsyncClient
async def main():
client = AsyncClient(host="http://gpu-box:11434")
async for chunk in await client.chat(
model="llama3.1:8b",
messages=[{"role": "user", "content": "Hello"}],
stream=True,
):
print(chunk["message"]["content"], end="", flush=True)
asyncio.run(main())
For existing OpenAI-based code, the OpenAI SDK approach above is usually the lower-friction path.
Configuration (Environment Variables)
Configuration is via environment variables read by the server process. On systemd, set them in a drop-in (see below) rather than editing the unit directly.
| Variable | Purpose | Default |
|---|---|---|
OLLAMA_HOST |
Bind address / port | 127.0.0.1:11434 |
OLLAMA_MODELS |
Model storage directory | ~/.ollama/models |
OLLAMA_KEEP_ALIVE |
How long to keep a model loaded after last use | 5m |
OLLAMA_NUM_PARALLEL |
Concurrent requests per loaded model | 1 |
OLLAMA_MAX_LOADED_MODELS |
Max models resident at once | 3 × GPUs (or 3 for CPU) |
OLLAMA_FLASH_ATTENTION |
Enable flash attention (lower KV-cache memory) | off |
OLLAMA_CONTEXT_LENGTH |
Default context window if model/request doesn't set one | 4096 (now VRAM-tiered — see caveat) |
OLLAMA_DEBUG |
Verbose server logging | off |
OLLAMA_KEEP_ALIVE accepts a duration (5m, 1h), -1 to keep loaded indefinitely, or 0 to unload immediately after each request. The same value can be set per-request via keep_alive.
Security:
OLLAMA_HOST=0.0.0.0:11434exposes an unauthenticated server to the network — anyone who can reach the port can run inference, pull/delete models, and read your model list. There is no built-in auth. If you must expose it, put it behind a reverse proxy that enforces TLS and authentication (or a private network / Tailscale), and never bind0.0.0.0on an untrusted segment.
Systemd Drop-in (Linux)
sudo systemctl edit ollama
# Creates /etc/systemd/system/ollama.service.d/override.conf
[Service]
Environment="OLLAMA_HOST=0.0.0.0:11434"
Environment="OLLAMA_KEEP_ALIVE=-1"
Environment="OLLAMA_MODELS=/data/ollama/models"
Environment="OLLAMA_FLASH_ATTENTION=1"
sudo systemctl daemon-reload
sudo systemctl restart ollama
Operational Notes
Model Storage
Models live under ~/.ollama/models (or OLLAMA_MODELS), split into content-addressed blobs/ (the actual GGUF weight files, shared between tags) and manifests/ (which layers make up each model:tag). Because blobs are deduplicated, ollama cp and overlapping tags are cheap. On Docker, this is the /root/.ollama volume — mount it persistently or you re-pull everything on each container recreate.
The Context-Window Caveat
This catches people out. The default context window is modest and is not the model's maximum. As of mid-2026 the server-side default is VRAM-tiered: 4096 tokens under ~24 GiB VRAM, 32K for 24–48 GiB, and 256K at 48 GiB and above (the historical flat default was smaller). Either way, a model advertising 128K context can still truncate below its maximum unless you raise the limit. Set it explicitly:
- Per request:
"options": { "num_ctx": 32768 }(generate/chat) - In a Modelfile:
PARAMETER num_ctx 32768 - Server-wide default:
OLLAMA_CONTEXT_LENGTH=32768
Larger num_ctx costs VRAM (the KV cache grows roughly linearly with context length), so size it to what you actually need.
Keep-Alive and VRAM
A model loaded into VRAM stays resident for keep_alive (default 5m) after the last request, then unloads. This avoids reloading on every call but holds VRAM. Use ollama ps to see what is resident and when it expires. For a long-running service, OLLAMA_KEEP_ALIVE=-1 keeps the model hot; for ad-hoc use that contends for VRAM, 0 frees it immediately.
Concurrency
OLLAMA_NUM_PARALLEL sets how many requests a single loaded model serves concurrently (they share the KV cache, so concurrency × num_ctx drives memory; default 1). OLLAMA_MAX_LOADED_MODELS caps how many distinct models stay resident (default 3 × the number of GPUs, or 3 for CPU inference). Use OLLAMA_MAX_QUEUE (default 512) to bound queued requests before the server returns 503.
GPU vs CPU Offload
Ollama offloads as many model layers to the GPU as fit, running the rest on CPU. Fewer layers fit → slower. Force the layer count with the num_gpu option (number of layers to offload; 0 = pure CPU). Check the actual split with ollama ps — a model showing 100% CPU when you expected GPU means layers didn't fit or the GPU wasn't detected.
Docker and Kubernetes
# GPU container with persistent model volume
docker run -d --gpus=all \
-v ollama:/root/.ollama \
-p 11434:11434 \
-e OLLAMA_KEEP_ALIVE=-1 \
--name ollama ollama/ollama
# Pull a model into the running container
docker exec ollama ollama pull llama3.1:8b
For Kubernetes: schedule onto a GPU node (nvidia.com/gpu resource request plus the device plugin), back /root/.ollama with a PVC so models survive pod restarts, and expose :11434 via a ClusterIP Service (keep it internal — there is no auth). A readiness probe on GET / or GET /api/tags works. Treat one Ollama pod as one inference node — for scale-out throughput, that is the boundary where vLLM becomes the better fit.
Quick Reference
CLI
| Command | Action |
|---|---|
ollama serve |
Start the server (foreground) |
ollama run <model> [prompt] |
Interactive REPL or one-shot completion |
ollama pull <model> |
Download/update a model |
ollama list |
List local models |
ollama ps |
Running models + VRAM/CPU split |
ollama stop <model> |
Unload a running model |
ollama rm <model> |
Delete a local model |
ollama cp <src> <dst> |
Copy/alias a model |
ollama show <model> |
Params, template, licence |
ollama create <name> -f Modelfile |
Build from a Modelfile |
ollama push <user/model> |
Push to ollama.com |
Key API Endpoints
| Endpoint | Method | Purpose |
|---|---|---|
/api/generate |
POST | Single-prompt completion |
/api/chat |
POST | Multi-turn chat + tools |
/api/embed |
POST | Batch embeddings (newer, input) |
/api/embeddings |
POST | Single embedding (older, prompt) |
/api/tags |
GET | List local models |
/api/ps |
GET | Running models |
/api/show |
POST | Model details |
/api/pull |
POST | Pull a model |
/api/delete |
DELETE | Delete a model |
/api/create |
POST | Build a model |
/v1/chat/completions |
POST | OpenAI-compatible chat |
/v1/completions |
POST | OpenAI-compatible completion |
/v1/embeddings |
POST | OpenAI-compatible embeddings |
/v1/models |
GET | OpenAI-compatible model list |
Key Environment Variables
| Variable | Effect |
|---|---|
OLLAMA_HOST |
Bind address (mind 0.0.0.0 — unauthenticated) |
OLLAMA_MODELS |
Model storage path |
OLLAMA_KEEP_ALIVE |
Model unload timeout (5m, -1, 0) |
OLLAMA_CONTEXT_LENGTH |
Default context window |
OLLAMA_NUM_PARALLEL |
Concurrent requests per model |
OLLAMA_MAX_LOADED_MODELS |
Max resident models |
OLLAMA_FLASH_ATTENTION |
Enable flash attention |
OLLAMA_DEBUG |
Verbose logging |
Common Issues and Solutions
| Issue | Cause | Solution |
|---|---|---|
| Model OOM / won't fit VRAM | Quant too large for the GPU | Use a smaller quant (q4_K_M), a smaller parameter size, lower num_ctx, or fewer num_gpu layers |
| Slow / running on CPU | GPU/driver not detected, or layers didn't fit | Check ollama ps for the CPU split; verify driver + CUDA/ROCm and (Docker) nvidia-container-toolkit; reduce model size so layers fit |
| Output truncated at long context | Default num_ctx (VRAM-tiered, e.g. 4096 on smaller GPUs) is below the model max |
Raise num_ctx per request, in a Modelfile, or via OLLAMA_CONTEXT_LENGTH |
| Connection refused | Server not running, or wrong host/port | Start ollama serve / systemctl start ollama; check OLLAMA_HOST and that the client points at the right address |
| Exposed without auth | Bound to 0.0.0.0 with no proxy |
Ollama has no built-in auth — front it with a reverse proxy/TLS/auth or keep it on a private network |
| Embeddings endpoint confusion | Mixing /api/embed and /api/embeddings |
/api/embed uses input → embeddings (array); /api/embeddings uses prompt → embedding (single) |
| Model reloads every request | keep_alive too short, or set to 0 |
Raise keep_alive per request or set OLLAMA_KEEP_ALIVE (-1 keeps it hot) |
| Disk filling up | Many pulled models / large blobs accumulating | ollama list then ollama rm unused models; blobs are shared but large models add up |
ollama run hangs on first use |
Model is downloading | First run pulls the model; watch progress or ollama pull ahead of time |
Related Topics
The following topics complement Ollama and extend a local-LLM workflow:
- vLLM — high-throughput production inference server (continuous batching, paged attention) for when single-node Ollama isn't enough
- FastAPI — wrap Ollama behind your own API, adding auth, rate-limiting, and request shaping
- LangChain / LlamaIndex — RAG and agent frameworks with first-class Ollama integrations
- Open WebUI — self-hosted chat UI that speaks to Ollama out of the box
- Hugging Face / GGUF — sourcing and converting models to import via
FROM ./model.gguf - Kubernetes — scheduling GPU workloads, device plugins, and PVCs for model storage