Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

Ollama

Running, managing, and serving local LLMs with Ollama — CLI, Modelfiles, and the REST/OpenAI-compatible API.

Ollama

Running, managing, and serving local LLMs with Ollama — CLI, Modelfiles, and the REST/OpenAI-compatible API.

Overview

Ollama wraps llama.cpp (and GGUF-format models) into a self-contained local server with a model registry, a CLI, and an HTTP API. It pulls quantised models from the Ollama library, manages their lifecycle in memory, and exposes both a native REST API and an OpenAI-compatible surface — making it a drop-in for code already written against the OpenAI SDK.

It runs on macOS, Linux, and Windows, and by default serves on 127.0.0.1:11434. GPU acceleration is automatic where available (NVIDIA CUDA, AMD ROCm, Apple Metal), with CPU fallback otherwise.

Ollama targets easy local and single-node development: pull a model, run it, point your client at localhost. For high-throughput production serving (continuous batching, tensor parallelism, paged attention), reach for vLLM instead — covered in its own sheet.

Ollama Server :11434Clientsload/unloadGGUF blobsoffload layersollama CLIREST /api/*OpenAI SDK /v1/*ollama python pkgHTTP APIModel Schedulerllama.cpp Runner~/.ollama/modelsGPU / CPUOllama Server :11434Clientsload/unloadGGUF blobsoffload layersollama CLIREST /api/*OpenAI SDK /v1/*ollama python pkgHTTP APIModel Schedulerllama.cpp Runner~/.ollama/modelsGPU / CPU

Install and Service

Install Methods

The two macOS routes install the same binaries but differ in lifecycle: the .dmg is Ollama's recommended path and keeps itself updated through the menu-bar app, whereas the Homebrew formula fits a scripted/dotfiles setup and updates with brew upgrade. Pick one — running both leaves you with two copies fighting over :11434.

# macOS (recommended) — download the .dmg app from ollama.com and drag to /Applications
# macOS (Homebrew) — CLI + server via the formula, for scripted setups
brew install ollama

# Linux — official install script (systemd service + ollama user)
curl -fsSL https://ollama.com/install.sh | sh

# Docker (CPU)
docker run -d -v ollama:/root/.ollama -p 11434:11434 \
  --name ollama ollama/ollama

# Docker (NVIDIA GPU — requires nvidia-container-toolkit)
docker run -d --gpus=all -v ollama:/root/.ollama -p 11434:11434 \
  --name ollama ollama/ollama

The Background Server

The CLI talks to a running server; commands like ollama run will fail with a connection error if nothing is listening on :11434.

# Run the server in the foreground (manual / dev use)
ollama serve

# On Linux, the install script registers a systemd service that
# starts ollama serve automatically as the 'ollama' user
systemctl status ollama
sudo systemctl restart ollama
journalctl -u ollama -f          # follow server logs

# macOS app runs the server as a background agent automatically

GPU Support

Backend Platform Notes
CUDA NVIDIA on Linux/Windows Needs recent driver; Docker needs nvidia-container-toolkit
ROCm AMD on Linux Supported GPU list is narrower; check official compatibility
Metal Apple Silicon (macOS) Automatic; unified memory means VRAM ≈ system RAM
CPU Everywhere Fallback when no GPU detected — much slower for large models

Confirm what a loaded model is actually using with ollama ps — the PROCESSOR column shows the GPU/CPU split.

Core CLI

Running Models

# Interactive REPL (pulls the model first if not present)
ollama run llama3.1:8b

# One-shot: pass the prompt as an argument, get a single completion
ollama run llama3.1:8b "Summarise the CAP theorem in two sentences."

# Pipe input
cat report.txt | ollama run llama3.1:8b "Summarise this:"

# Multimodal — attach an image by path in the prompt (vision models only)
ollama run llava "Describe this image: ./diagram.png"

Managing Models

ollama pull qwen2.5:7b              # download (or update) a model
ollama list                        # local models, sizes, modified time
ollama ps                          # running models + VRAM/CPU split + keep-alive
ollama stop llama3.1:8b            # unload a running model from memory
ollama rm llama3.1:8b             # delete a local model
ollama cp llama3.1:8b my-llama    # copy/alias a model locally
ollama show llama3.1:8b           # params, template, licence, modelfile
ollama show llama3.1:8b --modelfile   # just the Modelfile
ollama create mymodel -f Modelfile    # build a model from a Modelfile
ollama push user/mymodel          # push to ollama.com (needs an account + key)

Interactive REPL Commands

Inside ollama run, lines beginning with / are commands rather than prompts:

>>> /set parameter temperature 0.2   # override a runtime parameter
>>> /set system "You are a terse assistant."
>>> /show info                       # model metadata
>>> /show parameters                 # active parameters
>>> /load mymodel                    # switch to another model
>>> /save mysession                  # save current session as a new model
>>> /bye                             # exit the REPL (Ctrl-D also works)

>>> """                              # start a multi-line prompt
... line one
... line two
... """

Model Naming and Tags

Models are addressed as model:tag. Omitting the tag defaults to :latest.

llama3.1:8b                  # parameter-size tag
qwen2.5:7b-instruct-q4_K_M   # size + variant + quantisation
gemma2                       # bare name → resolves to :latest

The tag often encodes three things: parameter count (8b, 70b), variant (instruct, text), and quantisation (q4_K_M, q8_0, fp16).

Quantisation vs VRAM

Quantisation trades precision for memory. q4_K_M (4-bit, K-quant medium) is the common default sweet spot — roughly half a byte per weight plus overhead. Rough rule of thumb for a Q4 model: VRAM needed ≈ parameter count in billions × ~0.6–0.7 GB, plus the KV cache (which grows with num_ctx).

Suffix Bits Trade-off
q4_K_M ~4 Default sweet spot — good quality/size balance
q5_K_M ~5 Slightly better quality, more VRAM
q8_0 8 Near-lossless, ~2× the size of Q4
fp16 / f16 16 Full half-precision, largest footprint

If a model won't fit, drop to a smaller quant or a smaller parameter size before reaching for CPU offload.

Modelfile

A Modelfile is the build recipe for a model: a base, runtime parameters, a system prompt, a prompt template, and metadata. Build with ollama create.

Instructions

Instruction Purpose
FROM Base model (library name, another local model, or a ./path.gguf)
PARAMETER Default runtime parameter (temperature, num_ctx, stop, …)
SYSTEM Default system message
TEMPLATE Go-template prompt format (.System, .Prompt, .Messages, .Tools)
MESSAGE Seed a canned conversation turn (role + content)
ADAPTER Apply a LoRA/QLoRA adapter on top of FROM
LICENSE Embed licence text

Common Parameters

Parameter Meaning
temperature Sampling randomness (0 = deterministic-ish)
top_p Nucleus sampling cutoff
top_k Top-K sampling cutoff
num_ctx Context window size in tokens (see caveat below)
num_predict Max tokens to generate (-1 = until stop/EOS)
stop Stop sequence (repeat the instruction for several)
repeat_penalty Penalty against token repetition
seed Fix the RNG seed for reproducible output

Example

# Modelfile
FROM llama3.1:8b

PARAMETER temperature 0.2
PARAMETER top_p 0.9
PARAMETER num_ctx 8192
PARAMETER num_predict 1024
PARAMETER repeat_penalty 1.1
PARAMETER stop "<|eot_id|>"
PARAMETER stop "User:"

SYSTEM """
You are a senior SRE assistant. Answer tersely, in UK English.
Prefer concrete commands over prose.
"""

# Optional: a custom Go template (most base models ship a sensible one)
TEMPLATE """{{ if .System }}<|system|>{{ .System }}{{ end }}
<|user|>{{ .Prompt }}
<|assistant|>"""

LICENSE "Use restricted to internal tooling."
ollama create sre-assistant -f Modelfile
ollama run sre-assistant "Why might a pod be stuck in CrashLoopBackOff?"

Importing a Raw GGUF

Point FROM at a GGUF file downloaded from elsewhere (e.g. Hugging Face):

FROM ./mistral-7b-instruct-q4_k_m.gguf
PARAMETER num_ctx 4096
TEMPLATE """[INST] {{ .Prompt }} [/INST]"""

REST API

The native API lives under /api/* on :11434. This is the operational core. By default, streaming responses are newline-delimited JSON (NDJSON) — one object per line — terminated by a final object with "done": true carrying timing and token-count stats. Set "stream": false to get a single aggregated JSON object instead.

POST /api/generate

Single-prompt completion. Use /api/chat for multi-turn.

curl http://localhost:11434/api/generate -d '{
  "model": "llama3.1:8b",
  "prompt": "Explain idempotency in one sentence.",
  "stream": false,
  "options": { "temperature": 0.2, "num_ctx": 8192 },
  "keep_alive": "10m"
}'
{
  "model": "llama3.1:8b",
  "created_at": "2026-06-15T10:00:00Z",
  "response": "An idempotent operation produces the same result however many times it is applied.",
  "done": true,
  "total_duration": 1820000000,
  "eval_count": 21,
  "eval_duration": 900000000,
  "context": [128006, 9125, 128007, ...]
}

Useful fields:

Field Purpose
options Per-request overrides (temperature, num_ctx, seed, stop, …)
format "json" for JSON mode, or a full JSON Schema object for structured output
keep_alive How long to keep the model loaded after this request ("10m", -1, 0)
raw true bypasses templating — you supply the fully-formatted prompt
images Array of base64-encoded images for multimodal models
context Opaque token array from a prior response, for stateless continuation
suffix Text after the completion (fill-in-the-middle, for code models)

POST /api/chat

Multi-turn with a messages array of {role, content}. Roles: system, user, assistant, tool.

curl http://localhost:11434/api/chat -d '{
  "model": "llama3.1:8b",
  "messages": [
    { "role": "system", "content": "You answer in UK English." },
    { "role": "user", "content": "What is a write-ahead log?" }
  ],
  "stream": false
}'

Tools / function-calling (model must support it, e.g. Llama 3.1, Qwen 2.5):

curl http://localhost:11434/api/chat -d '{
  "model": "llama3.1:8b",
  "messages": [{ "role": "user", "content": "Weather in Kegworth?" }],
  "stream": false,
  "tools": [{
    "type": "function",
    "function": {
      "name": "get_weather",
      "description": "Current weather for a place",
      "parameters": {
        "type": "object",
        "properties": { "location": { "type": "string" } },
        "required": ["location"]
      }
    }
  }]
}'

The model replies with message.tool_calls; you execute the call and feed the result back as a {"role": "tool", "content": "..."} message.

Structured output — pass a JSON Schema as format:

curl http://localhost:11434/api/chat -d '{
  "model": "llama3.1:8b",
  "messages": [{ "role": "user", "content": "Give me a UK city and its population." }],
  "stream": false,
  "format": {
    "type": "object",
    "properties": {
      "city": { "type": "string" },
      "population": { "type": "integer" }
    },
    "required": ["city", "population"]
  }
}'

Embeddings

There are two endpoints. Prefer the newer /api/embed (batch via input); /api/embeddings is the older single-prompt form, kept for compatibility.

# Newer batch form — input can be a string or an array of strings
curl http://localhost:11434/api/embed -d '{
  "model": "nomic-embed-text",
  "input": ["first chunk", "second chunk"]
}'
# → { "embeddings": [[...], [...]], ... }

# Older single-prompt form
curl http://localhost:11434/api/embeddings -d '{
  "model": "nomic-embed-text",
  "prompt": "first chunk"
}'
# → { "embedding": [...] }

Note the key differences: /api/embed takes input and returns embeddings (array); /api/embeddings takes prompt and returns embedding (single).

Management Endpoints

curl http://localhost:11434/api/tags                 # GET — list local models
curl http://localhost:11434/api/ps                   # GET — running models
curl http://localhost:11434/api/show -d '{"model":"llama3.1:8b"}'   # POST — details
curl http://localhost:11434/api/pull -d '{"model":"qwen2.5:7b"}'    # POST — pull (streams progress)
curl -X DELETE http://localhost:11434/api/delete -d '{"model":"old-model"}'
curl http://localhost:11434/api/create -d '{
  "model": "sre-assistant",
  "from": "llama3.1:8b",
  "system": "You are a terse SRE assistant."
}'

Streaming Format

OllamaClientOllamaClientPOST /api/chat (stream: true){"message":{"content":"A"},"done":false}{"message":{"content":" log"},"done":false}{"message":{"content":"..."},"done":false}{"done":true,"eval_count":42,"total_duration":...}OllamaClientOllamaClientPOST /api/chat (stream: true){"message":{"content":"A"},"done":false}{"message":{"content":" log"},"done":false}{"message":{"content":"..."},"done":false}{"done":true,"eval_count":42,"total_duration":...}

Each line is a complete JSON object. Accumulate message.content (chat) or response (generate) across lines; the final "done": true object carries no new text but reports eval_count, prompt_eval_count, total_duration, and friends.

OpenAI-Compatible Endpoint

Ollama exposes an OpenAI-compatible surface under /v1/* — a drop-in for existing OpenAI code. Point the SDK at the local base URL and pass any non-empty api_key (it is ignored but the SDK requires one).

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:11434/v1",
    api_key="ollama",  # required by the SDK, ignored by Ollama
)

resp = client.chat.completions.create(
    model="llama3.1:8b",
    messages=[{"role": "user", "content": "What is a deadletter queue?"}],
)
print(resp.choices[0].message.content)

# Streaming
for chunk in client.chat.completions.create(
    model="llama3.1:8b",
    messages=[{"role": "user", "content": "Count to five."}],
    stream=True,
):
    print(chunk.choices[0].delta.content or "", end="")

# Embeddings
emb = client.embeddings.create(
    model="nomic-embed-text",
    input=["chunk one", "chunk two"],
)

Supported routes: /v1/chat/completions, /v1/completions, /v1/embeddings, /v1/models (and /v1/models/{model}), plus the newer /v1/responses and an experimental /v1/images/generations. The native /api/* surface exposes more (e.g. per-request keep_alive, raw mode), so use it when you need Ollama-specific control.

Python

The first-party ollama package wraps the native API.

import ollama

# Chat
resp = ollama.chat(
    model="llama3.1:8b",
    messages=[{"role": "user", "content": "Define backpressure."}],
)
print(resp["message"]["content"])

# Generate
print(ollama.generate(model="llama3.1:8b", prompt="One word for 'fast':")["response"])

# Embeddings (batch)
vecs = ollama.embed(model="nomic-embed-text", input=["a", "b"])["embeddings"]

# Streaming
for chunk in ollama.chat(
    model="llama3.1:8b",
    messages=[{"role": "user", "content": "Count to three."}],
    stream=True,
):
    print(chunk["message"]["content"], end="", flush=True)

Async, and pointing at a non-default host:

import asyncio
from ollama import AsyncClient

async def main():
    client = AsyncClient(host="http://gpu-box:11434")
    async for chunk in await client.chat(
        model="llama3.1:8b",
        messages=[{"role": "user", "content": "Hello"}],
        stream=True,
    ):
        print(chunk["message"]["content"], end="", flush=True)

asyncio.run(main())

For existing OpenAI-based code, the OpenAI SDK approach above is usually the lower-friction path.

Configuration (Environment Variables)

Configuration is via environment variables read by the server process. On systemd, set them in a drop-in (see below) rather than editing the unit directly.

Variable Purpose Default
OLLAMA_HOST Bind address / port 127.0.0.1:11434
OLLAMA_MODELS Model storage directory ~/.ollama/models
OLLAMA_KEEP_ALIVE How long to keep a model loaded after last use 5m
OLLAMA_NUM_PARALLEL Concurrent requests per loaded model 1
OLLAMA_MAX_LOADED_MODELS Max models resident at once 3 × GPUs (or 3 for CPU)
OLLAMA_FLASH_ATTENTION Enable flash attention (lower KV-cache memory) off
OLLAMA_CONTEXT_LENGTH Default context window if model/request doesn't set one 4096 (now VRAM-tiered — see caveat)
OLLAMA_DEBUG Verbose server logging off

OLLAMA_KEEP_ALIVE accepts a duration (5m, 1h), -1 to keep loaded indefinitely, or 0 to unload immediately after each request. The same value can be set per-request via keep_alive.

Security: OLLAMA_HOST=0.0.0.0:11434 exposes an unauthenticated server to the network — anyone who can reach the port can run inference, pull/delete models, and read your model list. There is no built-in auth. If you must expose it, put it behind a reverse proxy that enforces TLS and authentication (or a private network / Tailscale), and never bind 0.0.0.0 on an untrusted segment.

Systemd Drop-in (Linux)

sudo systemctl edit ollama
# Creates /etc/systemd/system/ollama.service.d/override.conf
[Service]
Environment="OLLAMA_HOST=0.0.0.0:11434"
Environment="OLLAMA_KEEP_ALIVE=-1"
Environment="OLLAMA_MODELS=/data/ollama/models"
Environment="OLLAMA_FLASH_ATTENTION=1"
sudo systemctl daemon-reload
sudo systemctl restart ollama

Operational Notes

Model Storage

Models live under ~/.ollama/models (or OLLAMA_MODELS), split into content-addressed blobs/ (the actual GGUF weight files, shared between tags) and manifests/ (which layers make up each model:tag). Because blobs are deduplicated, ollama cp and overlapping tags are cheap. On Docker, this is the /root/.ollama volume — mount it persistently or you re-pull everything on each container recreate.

The Context-Window Caveat

This catches people out. The default context window is modest and is not the model's maximum. As of mid-2026 the server-side default is VRAM-tiered: 4096 tokens under ~24 GiB VRAM, 32K for 24–48 GiB, and 256K at 48 GiB and above (the historical flat default was smaller). Either way, a model advertising 128K context can still truncate below its maximum unless you raise the limit. Set it explicitly:

  • Per request: "options": { "num_ctx": 32768 } (generate/chat)
  • In a Modelfile: PARAMETER num_ctx 32768
  • Server-wide default: OLLAMA_CONTEXT_LENGTH=32768

Larger num_ctx costs VRAM (the KV cache grows roughly linearly with context length), so size it to what you actually need.

Keep-Alive and VRAM

A model loaded into VRAM stays resident for keep_alive (default 5m) after the last request, then unloads. This avoids reloading on every call but holds VRAM. Use ollama ps to see what is resident and when it expires. For a long-running service, OLLAMA_KEEP_ALIVE=-1 keeps the model hot; for ad-hoc use that contends for VRAM, 0 frees it immediately.

Concurrency

OLLAMA_NUM_PARALLEL sets how many requests a single loaded model serves concurrently (they share the KV cache, so concurrency × num_ctx drives memory; default 1). OLLAMA_MAX_LOADED_MODELS caps how many distinct models stay resident (default 3 × the number of GPUs, or 3 for CPU inference). Use OLLAMA_MAX_QUEUE (default 512) to bound queued requests before the server returns 503.

GPU vs CPU Offload

Ollama offloads as many model layers to the GPU as fit, running the rest on CPU. Fewer layers fit → slower. Force the layer count with the num_gpu option (number of layers to offload; 0 = pure CPU). Check the actual split with ollama ps — a model showing 100% CPU when you expected GPU means layers didn't fit or the GPU wasn't detected.

Docker and Kubernetes

# GPU container with persistent model volume
docker run -d --gpus=all \
  -v ollama:/root/.ollama \
  -p 11434:11434 \
  -e OLLAMA_KEEP_ALIVE=-1 \
  --name ollama ollama/ollama

# Pull a model into the running container
docker exec ollama ollama pull llama3.1:8b

For Kubernetes: schedule onto a GPU node (nvidia.com/gpu resource request plus the device plugin), back /root/.ollama with a PVC so models survive pod restarts, and expose :11434 via a ClusterIP Service (keep it internal — there is no auth). A readiness probe on GET / or GET /api/tags works. Treat one Ollama pod as one inference node — for scale-out throughput, that is the boundary where vLLM becomes the better fit.

Quick Reference

CLI

Command Action
ollama serve Start the server (foreground)
ollama run <model> [prompt] Interactive REPL or one-shot completion
ollama pull <model> Download/update a model
ollama list List local models
ollama ps Running models + VRAM/CPU split
ollama stop <model> Unload a running model
ollama rm <model> Delete a local model
ollama cp <src> <dst> Copy/alias a model
ollama show <model> Params, template, licence
ollama create <name> -f Modelfile Build from a Modelfile
ollama push <user/model> Push to ollama.com

Key API Endpoints

Endpoint Method Purpose
/api/generate POST Single-prompt completion
/api/chat POST Multi-turn chat + tools
/api/embed POST Batch embeddings (newer, input)
/api/embeddings POST Single embedding (older, prompt)
/api/tags GET List local models
/api/ps GET Running models
/api/show POST Model details
/api/pull POST Pull a model
/api/delete DELETE Delete a model
/api/create POST Build a model
/v1/chat/completions POST OpenAI-compatible chat
/v1/completions POST OpenAI-compatible completion
/v1/embeddings POST OpenAI-compatible embeddings
/v1/models GET OpenAI-compatible model list

Key Environment Variables

Variable Effect
OLLAMA_HOST Bind address (mind 0.0.0.0 — unauthenticated)
OLLAMA_MODELS Model storage path
OLLAMA_KEEP_ALIVE Model unload timeout (5m, -1, 0)
OLLAMA_CONTEXT_LENGTH Default context window
OLLAMA_NUM_PARALLEL Concurrent requests per model
OLLAMA_MAX_LOADED_MODELS Max resident models
OLLAMA_FLASH_ATTENTION Enable flash attention
OLLAMA_DEBUG Verbose logging

Common Issues and Solutions

Issue Cause Solution
Model OOM / won't fit VRAM Quant too large for the GPU Use a smaller quant (q4_K_M), a smaller parameter size, lower num_ctx, or fewer num_gpu layers
Slow / running on CPU GPU/driver not detected, or layers didn't fit Check ollama ps for the CPU split; verify driver + CUDA/ROCm and (Docker) nvidia-container-toolkit; reduce model size so layers fit
Output truncated at long context Default num_ctx (VRAM-tiered, e.g. 4096 on smaller GPUs) is below the model max Raise num_ctx per request, in a Modelfile, or via OLLAMA_CONTEXT_LENGTH
Connection refused Server not running, or wrong host/port Start ollama serve / systemctl start ollama; check OLLAMA_HOST and that the client points at the right address
Exposed without auth Bound to 0.0.0.0 with no proxy Ollama has no built-in auth — front it with a reverse proxy/TLS/auth or keep it on a private network
Embeddings endpoint confusion Mixing /api/embed and /api/embeddings /api/embed uses input → embeddings (array); /api/embeddings uses prompt → embedding (single)
Model reloads every request keep_alive too short, or set to 0 Raise keep_alive per request or set OLLAMA_KEEP_ALIVE (-1 keeps it hot)
Disk filling up Many pulled models / large blobs accumulating ollama list then ollama rm unused models; blobs are shared but large models add up
ollama run hangs on first use Model is downloading First run pulls the model; watch progress or ollama pull ahead of time

Related Topics

The following topics complement Ollama and extend a local-LLM workflow:

  1. vLLM — high-throughput production inference server (continuous batching, paged attention) for when single-node Ollama isn't enough
  2. FastAPI — wrap Ollama behind your own API, adding auth, rate-limiting, and request shaping
  3. LangChain / LlamaIndex — RAG and agent frameworks with first-class Ollama integrations
  4. Open WebUI — self-hosted chat UI that speaks to Ollama out of the box
  5. Hugging Face / GGUF — sourcing and converting models to import via FROM ./model.gguf
  6. Kubernetes — scheduling GPU workloads, device plugins, and PVCs for model storage