vLLM
High-throughput LLM serving with vLLM — the OpenAI-compatible server, offline batching, and production deployment.
vLLM
High-throughput LLM serving with vLLM — the OpenAI-compatible server, offline batching, and production deployment.
Overview
vLLM is a high-throughput, memory-efficient inference and serving engine for large language models. Its two defining innovations are PagedAttention (which pages the KV cache like virtual memory, eliminating fragmentation and over-allocation) and continuous (in-flight) batching (which adds and evicts requests from the running batch every step instead of waiting for a fixed batch to drain), and together they keep the GPU saturated under concurrent load.
Where Ollama targets easy single-node local/dev inference on top of llama.cpp/GGUF (see the ollama sheet), vLLM targets production throughput on GPUs with full-precision or AWQ/GPTQ/FP8 weights. It primarily targets NVIDIA GPUs (CUDA), with growing AMD ROCm, CPU, TPU, and Intel Gaudi/XPU support.
graph TB
subgraph "vLLM Engine"
A[Incoming Requests] --> B[Scheduler<br/>continuous batching]
B --> C[Model Executor]
C --> D[(Paged KV Cache<br/>GPU HBM)]
C --> E[Sampler]
E --> F[Streamed Tokens]
F --> B
end
G[OpenAI-compatible<br/>API server :8000] --> A
H[Python LLM class<br/>offline batch] --> A
Two surfaces sit on the same engine: the Python LLM class for offline/batched inference, and the vllm serve OpenAI-compatible server for online serving. The rest of this sheet is organised around those two modes.
Installation
vLLM ships as a CUDA wheel built against a specific CUDA/PyTorch combination. On a supported NVIDIA box with recent drivers, the wheel is self-contained.
# uv (preferred) — into a fresh venv.
# --torch-backend=auto inspects the installed CUDA driver and picks the right PyTorch index.
uv venv --python 3.12
source .venv/bin/activate
uv pip install vllm --torch-backend=auto
# or plain pip, pinning the CUDA build explicitly
pip install vllm --extra-index-url https://download.pytorch.org/whl/cu129
Requirements (verify against current docs — these move):
| Requirement | Notes |
|---|---|
| Python | 3.10–3.13 supported |
| GPU | NVIDIA compute capability 7.0+ (Volta/Turing/Ampere/Hopper/Ada); 8.0+ for bf16/FP8 |
| CUDA | Default wheel built for CUDA 12.9; CUDA 12.8 / 13.0 wheels are published via the PyTorch index |
| Driver | Recent NVIDIA driver matching the CUDA runtime |
# Pin a specific CUDA backend instead of auto-detecting (e.g. cu128, cu130)
uv pip install vllm --torch-backend cu128
# AMD ROCm / CPU / other backends use dedicated wheel indexes or images —
# e.g. uv pip install vllm --extra-index-url https://wheels.vllm.ai/rocm/ --upgrade
# consult the install docs for the target platform.
Models are pulled from the Hugging Face Hub on first use and cached under ~/.cache/huggingface. Gated or private repos need a token:
export HF_TOKEN=hf_xxxxxxxxxxxxxxxxxxxx # gated models (Llama, Gemma, etc.)
# legacy var name HUGGING_FACE_HUB_TOKEN is also honoured
Offline / Batched Inference (the LLM class)
For batch jobs — dataset scoring, eval harnesses, synthetic-data generation — drive the engine directly in-process. The LLM class loads the model once and generate() runs an entire list of prompts through continuous batching with no HTTP overhead.
Key Concepts
LLM(model=...)constructs an engine; constructor args mirror the server flags (tensor_parallel_size,gpu_memory_utilization,max_model_len,dtype,quantization, …).SamplingParamscontrols decoding per request (or per batch).generate()takes a list of prompts and returns a list ofRequestOutput; ordering is preserved.chat()applies the model's chat template to a list of messages.
Common Patterns
from vllm import LLM, SamplingParams
llm = LLM(
model="meta-llama/Llama-3.1-8B-Instruct",
dtype="bfloat16",
gpu_memory_utilization=0.90,
max_model_len=8192,
)
params = SamplingParams(
temperature=0.7,
top_p=0.95,
max_tokens=256,
stop=["\n\n", "<|eot_id|>"],
n=1, # number of completions per prompt
)
prompts = [
"Explain PagedAttention in one sentence.",
"Write a haiku about GPUs.",
]
outputs = llm.generate(prompts, params)
for out in outputs:
print(out.prompt)
print(out.outputs[0].text) # .outputs is a list of length n
Chat Templates
from vllm import LLM, SamplingParams
llm = LLM(model="meta-llama/Llama-3.1-8B-Instruct")
conversation = [
{"role": "system", "content": "You are a terse senior engineer."},
{"role": "user", "content": "Difference between TP and PP?"},
]
# chat() applies the model's tokenizer chat template automatically
outputs = llm.chat(conversation, SamplingParams(temperature=0.3, max_tokens=200))
print(outputs[0].outputs[0].text)
Structured / Guided Decoding
vLLM can constrain output to a JSON schema, a fixed set of choices, a regex, or a grammar via a structured-outputs backend (xgrammar/guidance; backend auto by default). On current versions the per-request control is StructuredOutputsParams attached to SamplingParams via structured_outputs=. This replaces the older GuidedDecodingParams/guided_decoding= API (and the even older guided_json= kwarg), which was removed in vLLM 0.12.0.
from vllm import LLM, SamplingParams
from vllm.sampling_params import StructuredOutputsParams
llm = LLM(model="meta-llama/Llama-3.1-8B-Instruct")
schema = {
"type": "object",
"properties": {
"sentiment": {"type": "string", "enum": ["positive", "negative", "neutral"]},
"confidence": {"type": "number"},
},
"required": ["sentiment", "confidence"],
}
# Constrain to a JSON schema
so = StructuredOutputsParams(json=schema)
params = SamplingParams(temperature=0.0, max_tokens=128, structured_outputs=so)
out = llm.generate("Classify: 'this build is finally green'", params)
print(out[0].outputs[0].text) # valid JSON matching the schema
# Constrain to a fixed choice
so_choice = StructuredOutputsParams(choice=["yes", "no", "unsure"])
# Constrain to a regex
so_regex = StructuredOutputsParams(regex=r"\d{4}-\d{2}-\d{2}")
# Also available: grammar=..., structural_tag=...
The same constraints are exposed over the API via OpenAI's response_format with a JSON schema (the legacy top-level guided_json / guided_choice / guided_regex request fields are likewise deprecated in favour of structured_outputs).
Online Serving (vllm serve)
The server exposes an OpenAI-compatible HTTP API. vllm serve is the modern CLI entrypoint; the older python -m vllm.entrypoints.openai.api_server --model ... form still works but is deprecated.
# Modern CLI — model is a positional argument
vllm serve meta-llama/Llama-3.1-8B-Instruct \
--host 0.0.0.0 \
--port 8000 \
--max-model-len 8192 \
--gpu-memory-utilization 0.92
# Deprecated equivalent (older module form)
python -m vllm.entrypoints.openai.api_server \
--model meta-llama/Llama-3.1-8B-Instruct
It listens on :8000 by default and serves a single model per process. To run multiple models, run multiple servers (one per port/GPU set) behind a router.
OpenAI-Compatible API
The server mirrors the OpenAI REST surface, so existing OpenAI SDK code works by repointing base_url.
Endpoints
| Endpoint | Purpose |
|---|---|
POST /v1/chat/completions |
Chat completions (streaming + non-streaming) |
POST /v1/completions |
Legacy text completions |
POST /v1/embeddings |
Embeddings (embedding/pooling models) |
GET /v1/models |
List served models |
POST /tokenize, POST /detokenize |
Token<->text round-trips |
GET /health |
Liveness/readiness probe |
GET /metrics |
Prometheus metrics |
GET /version |
Server version |
curl
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "meta-llama/Llama-3.1-8B-Instruct",
"messages": [{"role": "user", "content": "One-line summary of PagedAttention."}],
"temperature": 0.3,
"max_tokens": 100
}'
Python (OpenAI SDK)
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="EMPTY", # any non-empty string unless --api-key is set
)
resp = client.chat.completions.create(
model="meta-llama/Llama-3.1-8B-Instruct",
messages=[{"role": "user", "content": "Explain continuous batching."}],
stream=True,
)
for chunk in resp:
delta = chunk.choices[0].delta.content
if delta:
print(delta, end="", flush=True)
API Keys, Tools, and Structured Outputs
# Require a bearer token on all requests
vllm serve <model> --api-key sk-my-secret-key
# Enable tool/function calling — parser must match the model family
vllm serve <model> \
--enable-auto-tool-choice \
--tool-call-parser hermes # e.g. hermes, llama3_json, mistral, ...
With --api-key set, clients pass Authorization: Bearer sk-my-secret-key. Tool calling and structured outputs (JSON schema via response_format, or guided_* fields) are supported; the --tool-call-parser must correspond to the model's tool-call format and a matching --chat-template may be required for some models.
Key Server Flags
The operational core. vllm serve <model> --flag ... on the CLI; the same names map to LLM(...) constructor kwargs with underscores (--tensor-parallel-size -> tensor_parallel_size). Defaults below are current-ish but should be verified — they drift across releases.
Parallelism
| Flag | Default | Purpose |
|---|---|---|
--tensor-parallel-size N |
1 | Shard each layer across N GPUs on one node (TP). Needs NCLL; N must divide the attention head count |
--pipeline-parallel-size N |
1 | Split layers into N stages, typically across nodes (PP) |
--data-parallel-size N |
1 | Replicate the model N times for independent request streams |
Memory & Batching
| Flag | Default | Purpose |
|---|---|---|
--gpu-memory-utilization F |
0.92 | Fraction of GPU memory vLLM may use; the KV-cache headroom knob. Lower it if other processes share the GPU |
--max-model-len N |
model config | Max context length (prompt + output). Larger = fewer concurrent sequences; smaller = more concurrency, less OOM risk |
--max-num-seqs N |
~256 (varies) | Max sequences in a running batch (concurrency cap). The V1 scheduler computes a default from model and hardware, commonly 256 — confirm with --help for your build rather than assuming |
--max-num-batched-tokens N |
varies (≈ max(2048, --max-model-len)) |
Max tokens processed per scheduler step; the throughput/latency lever for prefill. Computed per config, so treat the default as build-dependent |
--swap-space GB |
0 | CPU swap space (GiB) per GPU for preempted KV blocks; 0 disables CPU offload (V1 prefers recompute). The V0 engine defaulted to 4 GiB; the V1 engine (default since v0.8) uses 0 |
Precision & Cache
| Flag | Default | Purpose |
|---|---|---|
--dtype |
auto | auto/bfloat16/float16/float32/fp8 |
--quantization |
none | awq/gptq/fp8/bitsandbytes/… (often auto-detected from the repo) |
--kv-cache-dtype |
auto | fp8 (e8m0/e4m3) shrinks KV cache, raising concurrency at a small quality cost |
--enforce-eager |
false | Disable CUDA graph capture — slower, but lower memory and easier debugging |
--enable-chunked-prefill |
on (V1, where supported) | Interleave prefill and decode for steadier latency under load |
--enable-prefix-caching |
on (V1) | Reuse KV blocks for shared prompt prefixes (system prompts, few-shot); disable with --no-enable-prefix-caching |
Serving & Loading
| Flag | Default | Purpose |
|---|---|---|
--served-model-name NAME |
repo id | Name clients pass as model (decouple from the HF path) |
--host / --port |
0.0.0.0 / 8000 | Bind address and port |
--api-key KEY |
none | Require a bearer token |
--trust-remote-code |
false | Allow executing custom modelling code from the repo |
--download-dir PATH |
HF cache | Override the model download/cache directory |
--load-format |
auto | auto/safetensors/pt/dummy/… |
Multi-LoRA
| Flag | Purpose |
|---|---|
--enable-lora |
Turn on dynamic LoRA adapter serving |
--max-lora-rank N |
Max rank across loaded adapters (memory budget) |
--lora-modules name=path ... |
Preload named adapters; request them via the model field |
--max-loras N |
Max adapters resident in a single batch |
Performance Concepts
A handful of ideas explain most of vLLM's throughput — and most of its OOMs.
flowchart LR
A["Request arrives"] --> B["Prefill<br/>process prompt"]
B --> C{"Chunked<br/>prefill?"}
C -->|yes| D["Interleave prefill<br/>chunks with decode"]
C -->|no| E["Prefill in one step"]
D --> F["Decode loop<br/>1 token/step"]
E --> F
F --> G["Continuous batching:<br/>add/evict each step"]
G --> H["PagedAttention<br/>KV blocks"]
H --> F
- Continuous batching — the scheduler revises the running batch every decode step, admitting waiting requests and retiring finished ones, instead of waiting for a static batch. This is the throughput win.
- PagedAttention / KV cache — the KV cache is split into fixed-size blocks allocated on demand (like OS paging), so there's no per-request over-allocation and near-zero fragmentation. More effective KV capacity means more concurrent sequences.
- Prefix caching (on by default in V1;
--no-enable-prefix-cachingto disable) — identical prompt prefixes (shared system prompt, few-shot exemplars) reuse already-computed KV blocks, cutting prefill cost for repeated heads. - Chunked prefill (on by default in V1 where supported) — long prompts are prefilled in chunks interleaved with ongoing decodes, so a big prompt doesn't stall everyone's token generation; smooths TTFT/TPOT under mixed load.
- Speculative decoding (
--speculative-config '{...}', formerly--speculative-modeland friends) — a small draft model (or n-gram/EAGLE/Medusa method) proposes several tokens that the target model verifies in one pass, cutting latency when acceptance is high. - Quantisation — AWQ/GPTQ (weight-only 4-bit) and FP8 (weights and/or KV cache) shrink the memory footprint and can raise throughput, trading a little quality. FP8 needs Hopper/Ada-class hardware.
The two knobs behind most OOMs: --gpu-memory-utilization (how much HBM vLLM may claim, hence KV-cache size) and --max-model-len (per-sequence context, hence how many sequences fit). If it OOMs at startup, lower --max-model-len or --gpu-memory-utilization, or use a smaller dtype/quant. If it OOMs under load, the KV cache is exhausted — lower --max-num-seqs or --max-model-len, or enable --kv-cache-dtype fp8.
Multi-GPU & Distributed
graph TB
subgraph "Node A — Tensor Parallel (TP=4)"
L["Each layer's weights<br/>sharded across GPUs"]
L --> G0[GPU 0]
L --> G1[GPU 1]
L --> G2[GPU 2]
L --> G3[GPU 3]
G0 <-->|NCCL all-reduce| G1
G1 <-->|NCCL| G2
G2 <-->|NCCL| G3
end
subgraph "Pipeline Parallel (PP) across nodes"
S1["Stage 1<br/>layers 0-39<br/>Node A"] --> S2["Stage 2<br/>layers 40-79<br/>Node B"]
end
- Tensor parallelism (TP) — shards each layer's weights across GPUs that exchange activations every layer via NCCL. Latency-friendly but bandwidth-hungry; keep it within a node (NVLink/PCIe).
--tensor-parallel-sizemust divide the model's attention head count. - Pipeline parallelism (PP) — splits the model into sequential layer stages across nodes; tolerates slower interconnects but adds pipeline-fill latency. Use across nodes.
- Combine them for very large models:
--tensor-parallel-size 8 --pipeline-parallel-size 2= 16 GPUs across 2 nodes.
# Single node, 4 GPUs — tensor parallel
vllm serve meta-llama/Llama-3.1-70B-Instruct \
--tensor-parallel-size 4
# Multi-node uses Ray as the distributed runtime.
# Start a Ray cluster across the nodes first, then launch with combined TP x PP.
ray start --head # on the head node
ray start --address=<head> # on each worker node
vllm serve <big-model> \
--tensor-parallel-size 8 \
--pipeline-parallel-size 2
Multi-GPU runs depend on NCCL; in containers give the container enough shared memory (--shm-size) or NCCL will hang or crash. Set NCCL_DEBUG=INFO when diagnosing.
Observability
The server exposes Prometheus metrics at /metrics — wire it into the stack from the prometheus and grafana sheets.
curl -s http://localhost:8000/metrics | grep -E '^vllm:'
Key metric families (names prefixed vllm:):
| Metric area | Examples | Tells you |
|---|---|---|
| Throughput | prompt/generation tokens per second | Saturation and serving rate |
| Latency | time-to-first-token (TTFT), time-per-output-token (TPOT), end-to-end | User-facing responsiveness |
| Queue depth | running vs waiting requests | Whether you're admission-bound |
| KV cache | GPU cache usage fraction | How close you are to KV exhaustion / OOM under load |
| Requests | success/abort counters | Error and preemption rates |
# Quieten the periodic throughput log lines
vllm serve <model> --disable-log-stats
# Per-request logging (verbose; useful for debugging, noisy for prod)
vllm serve <model> --max-log-len 2048
A minimal Prometheus scrape config:
scrape_configs:
- job_name: 'vllm'
static_configs:
- targets: ['vllm:8000']
metrics_path: /metrics
Deployment
Docker
The official image is vllm/vllm-openai; its entrypoint is vllm serve, so pass server args directly. Mount the HF cache to avoid re-downloading on every start.
docker run --gpus all \
-p 8000:8000 \
--ipc=host \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-e HF_TOKEN=$HF_TOKEN \
vllm/vllm-openai:latest \
--model meta-llama/Llama-3.1-8B-Instruct \
--max-model-len 8192
--ipc=host (or --shm-size=8g) is important: the default 64 MB /dev/shm is too small for NCCL/PyTorch shared-memory and causes hangs on multi-GPU.
Kubernetes
GPU pods need the NVIDIA device plugin, a GPU request/limit, a node selector/tolerations for GPU nodes, and a cache volume so model load isn't repeated on every reschedule. Probe /health.
apiVersion: apps/v1
kind: Deployment
metadata:
name: vllm-llama
spec:
replicas: 1
selector:
matchLabels: { app: vllm-llama }
template:
metadata:
labels: { app: vllm-llama }
spec:
nodeSelector:
nvidia.com/gpu.product: NVIDIA-A100-SXM4-80GB
tolerations:
- key: nvidia.com/gpu
operator: Exists
effect: NoSchedule
containers:
- name: vllm
image: vllm/vllm-openai:latest
args:
- "--model=meta-llama/Llama-3.1-8B-Instruct"
- "--max-model-len=8192"
- "--gpu-memory-utilization=0.92"
ports:
- containerPort: 8000
env:
- name: HF_TOKEN
valueFrom:
secretKeyRef: { name: hf-token, key: token }
resources:
limits:
nvidia.com/gpu: "1"
volumeMounts:
- name: hf-cache
mountPath: /root/.cache/huggingface
- name: dshm
mountPath: /dev/shm
readinessProbe:
httpGet: { path: /health, port: 8000 }
initialDelaySeconds: 60 # model load is slow; be generous
periodSeconds: 10
livenessProbe:
httpGet: { path: /health, port: 8000 }
initialDelaySeconds: 120
volumes:
- name: hf-cache
persistentVolumeClaim:
claimName: hf-cache-pvc # shared, pre-warmed model cache
- name: dshm
emptyDir:
medium: Memory
sizeLimit: 8Gi
Autoscaling caveats. Cold starts are slow — pulling and loading a multi-GB model plus CUDA graph capture can take minutes — so HPA on request rate reacts far too late. Prefer scaling on a queue/utilisation metric with generous stabilisation windows, keep a warm minimum, and budget for the load time in readiness gates. Kubernetes-native serving stacks (KServe, the vLLM production-stack / llm-d patterns) wrap this with model caching, request-aware routing, and prefix-cache-aware load balancing; for a single fixed model a plain Deployment plus Service is enough.
Quick Reference
Offline (Python)
from vllm import LLM, SamplingParams
llm = LLM(model="org/model", tensor_parallel_size=2, max_model_len=8192)
out = llm.generate(["prompt"], SamplingParams(temperature=0.7, max_tokens=256))
out = llm.chat([{"role": "user", "content": "hi"}], SamplingParams(max_tokens=128))
Serve
vllm serve org/model \
--tensor-parallel-size 2 \
--max-model-len 8192 \
--gpu-memory-utilization 0.92 \
--quantization awq \
--enable-prefix-caching \
--api-key sk-key \
--served-model-name my-model
Endpoints & Flags
| Item | Value |
|---|---|
| Default bind | 0.0.0.0:8000 |
| Chat | POST /v1/chat/completions |
| Completions | POST /v1/completions |
| Embeddings | POST /v1/embeddings |
| Models | GET /v1/models |
| Health | GET /health |
| Metrics | GET /metrics |
| OOM knobs | --gpu-memory-utilization, --max-model-len |
| Multi-GPU (1 node) | --tensor-parallel-size N |
| Multi-node | --pipeline-parallel-size N + Ray |
| Quantise | --quantization awq|gptq|fp8 |
| Require key | --api-key sk-... |
Common Issues and Solutions
| Issue | Cause | Solution |
|---|---|---|
| CUDA OOM at startup | Weights + KV reservation exceed HBM | Lower --gpu-memory-utilization, reduce --max-model-len, use a smaller dtype or --quantization, or add GPUs via --tensor-parallel-size |
| OOM under load | KV cache exhausted by concurrent sequences | Lower --max-num-seqs or --max-model-len; enable --kv-cache-dtype fp8; cap --max-num-batched-tokens |
| "model's max seq len ... larger than KV cache can hold" | Requested context won't fit the available KV blocks | Reduce --max-model-len, or raise --gpu-memory-utilization to give the KV cache more room |
| TP size doesn't divide attention heads | --tensor-parallel-size not a divisor of head count |
Pick a TP size that divides the model's heads (e.g. 2/4/8); combine with PP if needed |
| 401 pulling a gated HF model | No/insufficient token | export HF_TOKEN=... and accept the model licence on the Hub first |
| "requires --trust-remote-code" | Custom modelling code in the repo | Add --trust-remote-code (only for repos you trust) |
| Slow first request | CUDA graph capture / warmup on first inference | Expected; warm the server with a dummy request, or --enforce-eager to skip graph capture (slower steady-state) |
| Multi-GPU NCCL hang | Insufficient shared memory in container | Run with --ipc=host or --shm-size=8g; set NCCL_DEBUG=INFO to diagnose |
| Wrong/garbled tool calls | Parser mismatch | Match --tool-call-parser to the model family and use the correct chat template |
| Poor CPU-offload performance | CPU offloading is a fallback, not a fast path | Size the GPU for the model; offloading trades large latency for fit and isn't a production throughput option |
Related Topics
The following complement vLLM for production LLM serving and operations:
- Ollama — the easy single-node local/dev counterpart (llama.cpp/GGUF); good for laptops and prototyping where vLLM is overkill
- Prometheus — scrape vLLM's
/metricsfor throughput, latency (TTFT/TPOT), and KV-cache utilisation - Grafana — dashboards over the vLLM Prometheus metrics
- Kubernetes — GPU scheduling, device plugin, node selectors/tolerations, and the cache-PVC patterns for serving
- FastAPI — wrap or route to vLLM behind your own auth, rate-limiting, and request-shaping layer
- Container Security — hardening GPU container images and handling
--trust-remote-code/ model-supply-chain risk