Available for day contractsFrom 21st September I have availability for day and half day contracts. Please contact for more information.

Contact →
mikepreston.org

vLLM

High-throughput LLM serving with vLLM — the OpenAI-compatible server, offline batching, and production deployment.

vLLM

High-throughput LLM serving with vLLM — the OpenAI-compatible server, offline batching, and production deployment.

Overview

vLLM is a high-throughput, memory-efficient inference and serving engine for large language models. Its two defining innovations are PagedAttention (which pages the KV cache like virtual memory, eliminating fragmentation and over-allocation) and continuous (in-flight) batching (which adds and evicts requests from the running batch every step instead of waiting for a fixed batch to drain), and together they keep the GPU saturated under concurrent load.

Where Ollama targets easy single-node local/dev inference on top of llama.cpp/GGUF (see the ollama sheet), vLLM targets production throughput on GPUs with full-precision or AWQ/GPTQ/FP8 weights. It primarily targets NVIDIA GPUs (CUDA), with growing AMD ROCm, CPU, TPU, and Intel Gaudi/XPU support.

vLLM EngineIncoming RequestsSchedulercontinuous batchingModel ExecutorPaged KV CacheGPU HBMSamplerStreamed TokensOpenAI-compatibleAPI server :8000Python LLM classoffline batchvLLM EngineIncoming RequestsSchedulercontinuous batchingModel ExecutorPaged KV CacheGPU HBMSamplerStreamed TokensOpenAI-compatibleAPI server :8000Python LLM classoffline batch

Two surfaces sit on the same engine: the Python LLM class for offline/batched inference, and the vllm serve OpenAI-compatible server for online serving. The rest of this sheet is organised around those two modes.

Installation

vLLM ships as a CUDA wheel built against a specific CUDA/PyTorch combination. On a supported NVIDIA box with recent drivers, the wheel is self-contained.

# uv (preferred) — into a fresh venv.
# --torch-backend=auto inspects the installed CUDA driver and picks the right PyTorch index.
uv venv --python 3.12
source .venv/bin/activate
uv pip install vllm --torch-backend=auto

# or plain pip, pinning the CUDA build explicitly
pip install vllm --extra-index-url https://download.pytorch.org/whl/cu129

Requirements (verify against current docs — these move):

Requirement Notes
Python 3.10–3.13 supported
GPU NVIDIA compute capability 7.0+ (Volta/Turing/Ampere/Hopper/Ada); 8.0+ for bf16/FP8
CUDA Default wheel built for CUDA 12.9; CUDA 12.8 / 13.0 wheels are published via the PyTorch index
Driver Recent NVIDIA driver matching the CUDA runtime
# Pin a specific CUDA backend instead of auto-detecting (e.g. cu128, cu130)
uv pip install vllm --torch-backend cu128

# AMD ROCm / CPU / other backends use dedicated wheel indexes or images —
# e.g. uv pip install vllm --extra-index-url https://wheels.vllm.ai/rocm/ --upgrade
# consult the install docs for the target platform.

Models are pulled from the Hugging Face Hub on first use and cached under ~/.cache/huggingface. Gated or private repos need a token:

export HF_TOKEN=hf_xxxxxxxxxxxxxxxxxxxx          # gated models (Llama, Gemma, etc.)
# legacy var name HUGGING_FACE_HUB_TOKEN is also honoured

Offline / Batched Inference (the LLM class)

For batch jobs — dataset scoring, eval harnesses, synthetic-data generation — drive the engine directly in-process. The LLM class loads the model once and generate() runs an entire list of prompts through continuous batching with no HTTP overhead.

Key Concepts

  • LLM(model=...) constructs an engine; constructor args mirror the server flags (tensor_parallel_size, gpu_memory_utilization, max_model_len, dtype, quantization, …).
  • SamplingParams controls decoding per request (or per batch).
  • generate() takes a list of prompts and returns a list of RequestOutput; ordering is preserved.
  • chat() applies the model's chat template to a list of messages.

Common Patterns

from vllm import LLM, SamplingParams

llm = LLM(
    model="meta-llama/Llama-3.1-8B-Instruct",
    dtype="bfloat16",
    gpu_memory_utilization=0.90,
    max_model_len=8192,
)

params = SamplingParams(
    temperature=0.7,
    top_p=0.95,
    max_tokens=256,
    stop=["\n\n", "<|eot_id|>"],
    n=1,                # number of completions per prompt
)

prompts = [
    "Explain PagedAttention in one sentence.",
    "Write a haiku about GPUs.",
]

outputs = llm.generate(prompts, params)
for out in outputs:
    print(out.prompt)
    print(out.outputs[0].text)   # .outputs is a list of length n

Chat Templates

from vllm import LLM, SamplingParams

llm = LLM(model="meta-llama/Llama-3.1-8B-Instruct")

conversation = [
    {"role": "system", "content": "You are a terse senior engineer."},
    {"role": "user", "content": "Difference between TP and PP?"},
]

# chat() applies the model's tokenizer chat template automatically
outputs = llm.chat(conversation, SamplingParams(temperature=0.3, max_tokens=200))
print(outputs[0].outputs[0].text)

Structured / Guided Decoding

vLLM can constrain output to a JSON schema, a fixed set of choices, a regex, or a grammar via a structured-outputs backend (xgrammar/guidance; backend auto by default). On current versions the per-request control is StructuredOutputsParams attached to SamplingParams via structured_outputs=. This replaces the older GuidedDecodingParams/guided_decoding= API (and the even older guided_json= kwarg), which was removed in vLLM 0.12.0.

from vllm import LLM, SamplingParams
from vllm.sampling_params import StructuredOutputsParams

llm = LLM(model="meta-llama/Llama-3.1-8B-Instruct")

schema = {
    "type": "object",
    "properties": {
        "sentiment": {"type": "string", "enum": ["positive", "negative", "neutral"]},
        "confidence": {"type": "number"},
    },
    "required": ["sentiment", "confidence"],
}

# Constrain to a JSON schema
so = StructuredOutputsParams(json=schema)
params = SamplingParams(temperature=0.0, max_tokens=128, structured_outputs=so)
out = llm.generate("Classify: 'this build is finally green'", params)
print(out[0].outputs[0].text)   # valid JSON matching the schema

# Constrain to a fixed choice
so_choice = StructuredOutputsParams(choice=["yes", "no", "unsure"])

# Constrain to a regex
so_regex = StructuredOutputsParams(regex=r"\d{4}-\d{2}-\d{2}")

# Also available: grammar=..., structural_tag=...

The same constraints are exposed over the API via OpenAI's response_format with a JSON schema (the legacy top-level guided_json / guided_choice / guided_regex request fields are likewise deprecated in favour of structured_outputs).

Online Serving (vllm serve)

The server exposes an OpenAI-compatible HTTP API. vllm serve is the modern CLI entrypoint; the older python -m vllm.entrypoints.openai.api_server --model ... form still works but is deprecated.

# Modern CLI — model is a positional argument
vllm serve meta-llama/Llama-3.1-8B-Instruct \
    --host 0.0.0.0 \
    --port 8000 \
    --max-model-len 8192 \
    --gpu-memory-utilization 0.92

# Deprecated equivalent (older module form)
python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Llama-3.1-8B-Instruct

It listens on :8000 by default and serves a single model per process. To run multiple models, run multiple servers (one per port/GPU set) behind a router.

OpenAI-Compatible API

The server mirrors the OpenAI REST surface, so existing OpenAI SDK code works by repointing base_url.

Endpoints

Endpoint Purpose
POST /v1/chat/completions Chat completions (streaming + non-streaming)
POST /v1/completions Legacy text completions
POST /v1/embeddings Embeddings (embedding/pooling models)
GET /v1/models List served models
POST /tokenize, POST /detokenize Token<->text round-trips
GET /health Liveness/readiness probe
GET /metrics Prometheus metrics
GET /version Server version

curl

curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/Llama-3.1-8B-Instruct",
    "messages": [{"role": "user", "content": "One-line summary of PagedAttention."}],
    "temperature": 0.3,
    "max_tokens": 100
  }'

Python (OpenAI SDK)

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="EMPTY",          # any non-empty string unless --api-key is set
)

resp = client.chat.completions.create(
    model="meta-llama/Llama-3.1-8B-Instruct",
    messages=[{"role": "user", "content": "Explain continuous batching."}],
    stream=True,
)
for chunk in resp:
    delta = chunk.choices[0].delta.content
    if delta:
        print(delta, end="", flush=True)

API Keys, Tools, and Structured Outputs

# Require a bearer token on all requests
vllm serve <model> --api-key sk-my-secret-key

# Enable tool/function calling — parser must match the model family
vllm serve <model> \
    --enable-auto-tool-choice \
    --tool-call-parser hermes        # e.g. hermes, llama3_json, mistral, ...

With --api-key set, clients pass Authorization: Bearer sk-my-secret-key. Tool calling and structured outputs (JSON schema via response_format, or guided_* fields) are supported; the --tool-call-parser must correspond to the model's tool-call format and a matching --chat-template may be required for some models.

Key Server Flags

The operational core. vllm serve <model> --flag ... on the CLI; the same names map to LLM(...) constructor kwargs with underscores (--tensor-parallel-size -> tensor_parallel_size). Defaults below are current-ish but should be verified — they drift across releases.

Parallelism

Flag Default Purpose
--tensor-parallel-size N 1 Shard each layer across N GPUs on one node (TP). Needs NCLL; N must divide the attention head count
--pipeline-parallel-size N 1 Split layers into N stages, typically across nodes (PP)
--data-parallel-size N 1 Replicate the model N times for independent request streams

Memory & Batching

Flag Default Purpose
--gpu-memory-utilization F 0.92 Fraction of GPU memory vLLM may use; the KV-cache headroom knob. Lower it if other processes share the GPU
--max-model-len N model config Max context length (prompt + output). Larger = fewer concurrent sequences; smaller = more concurrency, less OOM risk
--max-num-seqs N ~256 (varies) Max sequences in a running batch (concurrency cap). The V1 scheduler computes a default from model and hardware, commonly 256 — confirm with --help for your build rather than assuming
--max-num-batched-tokens N varies (≈ max(2048, --max-model-len)) Max tokens processed per scheduler step; the throughput/latency lever for prefill. Computed per config, so treat the default as build-dependent
--swap-space GB 0 CPU swap space (GiB) per GPU for preempted KV blocks; 0 disables CPU offload (V1 prefers recompute). The V0 engine defaulted to 4 GiB; the V1 engine (default since v0.8) uses 0

Precision & Cache

Flag Default Purpose
--dtype auto auto/bfloat16/float16/float32/fp8
--quantization none awq/gptq/fp8/bitsandbytes/… (often auto-detected from the repo)
--kv-cache-dtype auto fp8 (e8m0/e4m3) shrinks KV cache, raising concurrency at a small quality cost
--enforce-eager false Disable CUDA graph capture — slower, but lower memory and easier debugging
--enable-chunked-prefill on (V1, where supported) Interleave prefill and decode for steadier latency under load
--enable-prefix-caching on (V1) Reuse KV blocks for shared prompt prefixes (system prompts, few-shot); disable with --no-enable-prefix-caching

Serving & Loading

Flag Default Purpose
--served-model-name NAME repo id Name clients pass as model (decouple from the HF path)
--host / --port 0.0.0.0 / 8000 Bind address and port
--api-key KEY none Require a bearer token
--trust-remote-code false Allow executing custom modelling code from the repo
--download-dir PATH HF cache Override the model download/cache directory
--load-format auto auto/safetensors/pt/dummy/…

Multi-LoRA

Flag Purpose
--enable-lora Turn on dynamic LoRA adapter serving
--max-lora-rank N Max rank across loaded adapters (memory budget)
--lora-modules name=path ... Preload named adapters; request them via the model field
--max-loras N Max adapters resident in a single batch

Performance Concepts

A handful of ideas explain most of vLLM's throughput — and most of its OOMs.

yesnoRequest arrivesPrefillprocess promptChunkedprefill?Interleave prefillchunks with decodePrefill in one stepDecode loop1 token/stepContinuous batching:add/evict each stepPagedAttentionKV blocksyesnoRequest arrivesPrefillprocess promptChunkedprefill?Interleave prefillchunks with decodePrefill in one stepDecode loop1 token/stepContinuous batching:add/evict each stepPagedAttentionKV blocks
  • Continuous batching — the scheduler revises the running batch every decode step, admitting waiting requests and retiring finished ones, instead of waiting for a static batch. This is the throughput win.
  • PagedAttention / KV cache — the KV cache is split into fixed-size blocks allocated on demand (like OS paging), so there's no per-request over-allocation and near-zero fragmentation. More effective KV capacity means more concurrent sequences.
  • Prefix caching (on by default in V1; --no-enable-prefix-caching to disable) — identical prompt prefixes (shared system prompt, few-shot exemplars) reuse already-computed KV blocks, cutting prefill cost for repeated heads.
  • Chunked prefill (on by default in V1 where supported) — long prompts are prefilled in chunks interleaved with ongoing decodes, so a big prompt doesn't stall everyone's token generation; smooths TTFT/TPOT under mixed load.
  • Speculative decoding (--speculative-config '{...}', formerly --speculative-model and friends) — a small draft model (or n-gram/EAGLE/Medusa method) proposes several tokens that the target model verifies in one pass, cutting latency when acceptance is high.
  • Quantisation — AWQ/GPTQ (weight-only 4-bit) and FP8 (weights and/or KV cache) shrink the memory footprint and can raise throughput, trading a little quality. FP8 needs Hopper/Ada-class hardware.

The two knobs behind most OOMs: --gpu-memory-utilization (how much HBM vLLM may claim, hence KV-cache size) and --max-model-len (per-sequence context, hence how many sequences fit). If it OOMs at startup, lower --max-model-len or --gpu-memory-utilization, or use a smaller dtype/quant. If it OOMs under load, the KV cache is exhausted — lower --max-num-seqs or --max-model-len, or enable --kv-cache-dtype fp8.

Multi-GPU & Distributed

Pipeline Parallel (PP) across nodesStage 1layers 0-39Node AStage 2layers 40-79Node BNode A — Tensor Parallel (TP=4)NCCL all-reduceNCCLNCCLEach layer's weightssharded across GPUsGPU 0GPU 1GPU 2GPU 3Pipeline Parallel (PP) across nodesStage 1layers 0-39Node AStage 2layers 40-79Node BNode A — Tensor Parallel (TP=4)NCCL all-reduceNCCLNCCLEach layer's weightssharded across GPUsGPU 0GPU 1GPU 2GPU 3
  • Tensor parallelism (TP) — shards each layer's weights across GPUs that exchange activations every layer via NCCL. Latency-friendly but bandwidth-hungry; keep it within a node (NVLink/PCIe). --tensor-parallel-size must divide the model's attention head count.
  • Pipeline parallelism (PP) — splits the model into sequential layer stages across nodes; tolerates slower interconnects but adds pipeline-fill latency. Use across nodes.
  • Combine them for very large models: --tensor-parallel-size 8 --pipeline-parallel-size 2 = 16 GPUs across 2 nodes.
# Single node, 4 GPUs — tensor parallel
vllm serve meta-llama/Llama-3.1-70B-Instruct \
    --tensor-parallel-size 4

# Multi-node uses Ray as the distributed runtime.
# Start a Ray cluster across the nodes first, then launch with combined TP x PP.
ray start --head            # on the head node
ray start --address=<head>  # on each worker node
vllm serve <big-model> \
    --tensor-parallel-size 8 \
    --pipeline-parallel-size 2

Multi-GPU runs depend on NCCL; in containers give the container enough shared memory (--shm-size) or NCCL will hang or crash. Set NCCL_DEBUG=INFO when diagnosing.

Observability

The server exposes Prometheus metrics at /metrics — wire it into the stack from the prometheus and grafana sheets.

curl -s http://localhost:8000/metrics | grep -E '^vllm:'

Key metric families (names prefixed vllm:):

Metric area Examples Tells you
Throughput prompt/generation tokens per second Saturation and serving rate
Latency time-to-first-token (TTFT), time-per-output-token (TPOT), end-to-end User-facing responsiveness
Queue depth running vs waiting requests Whether you're admission-bound
KV cache GPU cache usage fraction How close you are to KV exhaustion / OOM under load
Requests success/abort counters Error and preemption rates
# Quieten the periodic throughput log lines
vllm serve <model> --disable-log-stats

# Per-request logging (verbose; useful for debugging, noisy for prod)
vllm serve <model> --max-log-len 2048

A minimal Prometheus scrape config:

scrape_configs:
  - job_name: 'vllm'
    static_configs:
      - targets: ['vllm:8000']
    metrics_path: /metrics

Deployment

Docker

The official image is vllm/vllm-openai; its entrypoint is vllm serve, so pass server args directly. Mount the HF cache to avoid re-downloading on every start.

docker run --gpus all \
  -p 8000:8000 \
  --ipc=host \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  -e HF_TOKEN=$HF_TOKEN \
  vllm/vllm-openai:latest \
  --model meta-llama/Llama-3.1-8B-Instruct \
  --max-model-len 8192

--ipc=host (or --shm-size=8g) is important: the default 64 MB /dev/shm is too small for NCCL/PyTorch shared-memory and causes hangs on multi-GPU.

Kubernetes

GPU pods need the NVIDIA device plugin, a GPU request/limit, a node selector/tolerations for GPU nodes, and a cache volume so model load isn't repeated on every reschedule. Probe /health.

apiVersion: apps/v1
kind: Deployment
metadata:
  name: vllm-llama
spec:
  replicas: 1
  selector:
    matchLabels: { app: vllm-llama }
  template:
    metadata:
      labels: { app: vllm-llama }
    spec:
      nodeSelector:
        nvidia.com/gpu.product: NVIDIA-A100-SXM4-80GB
      tolerations:
        - key: nvidia.com/gpu
          operator: Exists
          effect: NoSchedule
      containers:
        - name: vllm
          image: vllm/vllm-openai:latest
          args:
            - "--model=meta-llama/Llama-3.1-8B-Instruct"
            - "--max-model-len=8192"
            - "--gpu-memory-utilization=0.92"
          ports:
            - containerPort: 8000
          env:
            - name: HF_TOKEN
              valueFrom:
                secretKeyRef: { name: hf-token, key: token }
          resources:
            limits:
              nvidia.com/gpu: "1"
          volumeMounts:
            - name: hf-cache
              mountPath: /root/.cache/huggingface
            - name: dshm
              mountPath: /dev/shm
          readinessProbe:
            httpGet: { path: /health, port: 8000 }
            initialDelaySeconds: 60      # model load is slow; be generous
            periodSeconds: 10
          livenessProbe:
            httpGet: { path: /health, port: 8000 }
            initialDelaySeconds: 120
      volumes:
        - name: hf-cache
          persistentVolumeClaim:
            claimName: hf-cache-pvc       # shared, pre-warmed model cache
        - name: dshm
          emptyDir:
            medium: Memory
            sizeLimit: 8Gi

Autoscaling caveats. Cold starts are slow — pulling and loading a multi-GB model plus CUDA graph capture can take minutes — so HPA on request rate reacts far too late. Prefer scaling on a queue/utilisation metric with generous stabilisation windows, keep a warm minimum, and budget for the load time in readiness gates. Kubernetes-native serving stacks (KServe, the vLLM production-stack / llm-d patterns) wrap this with model caching, request-aware routing, and prefix-cache-aware load balancing; for a single fixed model a plain Deployment plus Service is enough.

Quick Reference

Offline (Python)

from vllm import LLM, SamplingParams
llm = LLM(model="org/model", tensor_parallel_size=2, max_model_len=8192)
out = llm.generate(["prompt"], SamplingParams(temperature=0.7, max_tokens=256))
out = llm.chat([{"role": "user", "content": "hi"}], SamplingParams(max_tokens=128))

Serve

vllm serve org/model \
  --tensor-parallel-size 2 \
  --max-model-len 8192 \
  --gpu-memory-utilization 0.92 \
  --quantization awq \
  --enable-prefix-caching \
  --api-key sk-key \
  --served-model-name my-model

Endpoints & Flags

Item Value
Default bind 0.0.0.0:8000
Chat POST /v1/chat/completions
Completions POST /v1/completions
Embeddings POST /v1/embeddings
Models GET /v1/models
Health GET /health
Metrics GET /metrics
OOM knobs --gpu-memory-utilization, --max-model-len
Multi-GPU (1 node) --tensor-parallel-size N
Multi-node --pipeline-parallel-size N + Ray
Quantise --quantization awq|gptq|fp8
Require key --api-key sk-...

Common Issues and Solutions

Issue Cause Solution
CUDA OOM at startup Weights + KV reservation exceed HBM Lower --gpu-memory-utilization, reduce --max-model-len, use a smaller dtype or --quantization, or add GPUs via --tensor-parallel-size
OOM under load KV cache exhausted by concurrent sequences Lower --max-num-seqs or --max-model-len; enable --kv-cache-dtype fp8; cap --max-num-batched-tokens
"model's max seq len ... larger than KV cache can hold" Requested context won't fit the available KV blocks Reduce --max-model-len, or raise --gpu-memory-utilization to give the KV cache more room
TP size doesn't divide attention heads --tensor-parallel-size not a divisor of head count Pick a TP size that divides the model's heads (e.g. 2/4/8); combine with PP if needed
401 pulling a gated HF model No/insufficient token export HF_TOKEN=... and accept the model licence on the Hub first
"requires --trust-remote-code" Custom modelling code in the repo Add --trust-remote-code (only for repos you trust)
Slow first request CUDA graph capture / warmup on first inference Expected; warm the server with a dummy request, or --enforce-eager to skip graph capture (slower steady-state)
Multi-GPU NCCL hang Insufficient shared memory in container Run with --ipc=host or --shm-size=8g; set NCCL_DEBUG=INFO to diagnose
Wrong/garbled tool calls Parser mismatch Match --tool-call-parser to the model family and use the correct chat template
Poor CPU-offload performance CPU offloading is a fallback, not a fast path Size the GPU for the model; offloading trades large latency for fit and isn't a production throughput option

Related Topics

The following complement vLLM for production LLM serving and operations:

  1. Ollama — the easy single-node local/dev counterpart (llama.cpp/GGUF); good for laptops and prototyping where vLLM is overkill
  2. Prometheus — scrape vLLM's /metrics for throughput, latency (TTFT/TPOT), and KV-cache utilisation
  3. Grafana — dashboards over the vLLM Prometheus metrics
  4. Kubernetes — GPU scheduling, device plugin, node selectors/tolerations, and the cache-PVC patterns for serving
  5. FastAPI — wrap or route to vLLM behind your own auth, rate-limiting, and request-shaping layer
  6. Container Security — hardening GPU container images and handling --trust-remote-code / model-supply-chain risk