Blog20 min read

How vLLM Works Internally: A Deep Dive

A systems-level deep dive into vLLM: PagedAttention, the scheduler, KV cache management, continuous batching, tensor parallelism, speculative decoding, and the V1 architecture.

If you serve LLMs in production, you have probably run vLLM. But most engineers treat it as a black box: start the server, send requests, get tokens back.

This post opens the box. We will walk through vLLM's architecture the way you would walk through an operating system kernel: component by component, data structure by data structure, with enough detail that you could reason about its behavior under load.

We will cover:

  • Why naive serving wastes most of your GPU memory
  • PagedAttention: the OS virtual memory trick applied to KV cache
  • The scheduler: three queues, preemption, and chunked prefill
  • How a single request flows from HTTP to streamed tokens
  • Model weight loading and tensor parallelism
  • Speculative decoding, prefix caching, and CUDA graphs
  • The V1 architecture rewrite

The Problem: GPU Memory is the Bottleneck

A 13B parameter model in FP16 consumes about 26 GB just for weights. An A100 has 80 GB. That leaves 54 GB for everything else: activations, KV cache, CUDA context, and framework overhead.

The KV cache is where things get interesting. For each token in each layer, the model stores a key vector and a value vector. For LLaMA-13B, a single sequence at maximum context length (2048 tokens) requires about 1.7 GB of KV cache. That means you can fit roughly 30 concurrent sequences before the GPU is out of memory.

But it gets worse. In practice, existing systems (FasterTransformer, early HuggingFace TGI) pre-allocate KV cache for the maximum possible sequence length, even if the actual output is 10 tokens. The result:

60-80% of KV cache memory is wasted.

This waste comes from three sources:

  1. Reservation waste: pre-allocating for max_seq_len when actual output is much shorter
  2. Internal fragmentation: memory allocators round up, leaving gaps inside allocations
  3. External fragmentation: different-sized allocations create unusable gaps between them

This is the exact same problem that operating systems solved decades ago with virtual memory and paging. vLLM applies the same solution.

PagedAttention: Virtual Memory for KV Cache

PagedAttention (Kwon et al., SOSP 2023) is the core innovation that makes vLLM work. The idea is deceptively simple: instead of storing each sequence's KV cache as one contiguous tensor, break it into fixed-size blocks (pages) that can live anywhere in GPU memory.

The Data Structures

Three concepts map directly from OS virtual memory:

OS ConceptvLLM EquivalentWhat it stores
Virtual pageLogical blockContiguous token positions (e.g., tokens 0-15)
Physical framePhysical blockActual GPU memory: [num_kv_heads, block_size, head_dim] for K and V
Page tableBlock tablePer-sequence mapping: logical block index → physical block index

The default block size is 16 tokens. For a 13B model, each physical block is about 12.8 KB. Small enough to minimize waste in the last block, large enough for efficient GPU memory access.

How the Kernel Works

The PagedAttention CUDA kernel replaces standard attention with a block-aware version. When computing attention for a query token:

  1. Look up the sequence's block table to find physical block locations
  2. Fetch K vectors from potentially non-contiguous physical blocks
  3. Compute attention scores across all blocks
  4. Apply softmax and compute weighted sum of V vectors

There are two kernel variants:

  • PagedAttention V1: One thread group per block. Works well for small batch sizes.
  • PagedAttention V2: Two-pass approach. The first pass computes partial softmax results per block, the second pass reduces across blocks. Better parallelism for large batches.

When multiple beam search candidates share a prefix, they share the same physical blocks. The block table entries point to the same physical memory, tracked by reference count.

When a beam diverges and needs to write a new token to a shared block:

  1. Check if the physical block's reference count > 1
  2. Allocate a new physical block
  3. Copy the old block's contents to the new one
  4. Update this beam's block table to point to the new block
  5. Decrement the old block's reference count

This is exactly how Unix implements fork() with copy-on-write pages. Same problem, same solution, different domain.

The Result

PagedAttention wastes less than 4% of KV cache memory. The only waste is in the last block of each sequence, which may have up to block_size - 1 empty slots.

This translates directly to throughput:

  • 14-24x higher throughput than HuggingFace Transformers
  • 2.2-3.5x higher throughput than HuggingFace Text Generation Inference
  • 2-4x more concurrent sequences at the same GPU memory budget

High-Level Architecture

flowchart TD
    C[Client] -->|HTTP / OpenAI API| A[AsyncLLM]
    A -->|IPC via ZMQ| E[EngineCore]
    E --> SCH[Scheduler]
    E --> KV[KV Cache Manager]
    E --> EX[Model Executor]
    SCH -->|schedule| EX
    KV -->|block allocation| SCH
    EX --> W1[Worker GPU 0]
    EX --> W2[Worker GPU 1]
    EX --> WN[Worker GPU N]
    W1 --> MR1[ModelRunner]
    MR1 --> FP[Forward Pass]
    FP --> SA[Sampler]
    SA -->|tokens| E
    E -->|IPC| A
    A -->|SSE stream| C

In the V1 architecture, vLLM separates into distinct processes to avoid Python's GIL:

  • AsyncLLM (vllm/v1/engine/async_llm.py): The API-facing process. Handles tokenization, detokenization, and request submission. Communicates with EngineCore via asynchronous IPC (ZMQ sockets).

  • EngineCore (vllm/v1/engine/core.py): The inference engine. Runs a tight busy loop: pull requests from input queue, run one scheduler step, execute one forward pass, push outputs. Contains the Scheduler and KV Cache Manager.

  • Scheduler (vllm/v1/core/sched/scheduler.py): Decides which requests run in each step. Maintains waiting and running queues. Enforces token budget and memory constraints.

  • KV Cache Manager (vllm/v1/core/kv_cache_manager.py): Manages the block pool. All KVCacheBlock objects are pre-allocated at startup as a block pool, avoiding Python object creation overhead. Uses a doubly-linked free list for O(1) allocation and deallocation.

  • Model Executor / Worker: One worker per GPU. Each worker owns a ModelRunner that loads the model and executes forward passes. For multi-GPU, a MultiProcExecutor coordinates tensor parallelism via shared-memory message queues.

Life of a Request: End to End

Let us trace a single request from curl to streamed tokens.

curl -N http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/Meta-Llama-3-8B-Instruct",
    "messages": [{"role":"user","content":"Explain paged attention in one paragraph."}],
    "max_tokens": 128,
    "temperature": 0.7,
    "stream": true
  }'
sequenceDiagram
    autonumber
    participant C as Client
    participant API as AsyncLLM
    participant EC as EngineCore
    participant S as Scheduler
    participant KV as KV Cache Mgr
    participant MR as ModelRunner
    participant GPU as GPU

    C->>API: POST /v1/chat/completions
    API->>API: Validate + tokenize prompt
    API->>EC: Send request via IPC
    EC->>S: Enqueue in waiting deque
    Note over S: === Scheduler Step ===
    S->>KV: allocate_slots(prompt_tokens)
    KV-->>S: Block IDs assigned
    S->>MR: Execute prefill batch
    MR->>GPU: Forward pass (all prompt tokens)
    GPU-->>MR: Logits for last position
    MR->>MR: Sample first token
    MR-->>EC: Output token
    EC-->>API: Stream token via IPC
    API-->>C: SSE chunk (first token)
    Note over S: === Decode Loop ===
    loop Until stop condition
        S->>KV: Allocate new block if needed
        S->>MR: Execute decode batch
        MR->>GPU: Forward pass (1 token per seq)
        GPU-->>MR: Logits
        MR->>MR: Sample next token
        MR-->>EC: Output token
        EC-->>API: Stream via IPC
        API-->>C: SSE chunk
    end
    API-->>C: [DONE] + usage stats
    S->>KV: Free all blocks

Step 1: API Ingress and Tokenization

The request hits the FastAPI server. AsyncLLM validates the JSON, applies the model's chat template (Jinja2) to convert the messages array into a single prompt string, and tokenizes it into prompt_token_ids. Tokenization runs in the AsyncLLM process, separate from GPU inference, so it does not block the engine loop.

Step 2: IPC to EngineCore

The tokenized request is sent to EngineCore through AsyncMPClient via ZMQ inter-process communication. A background thread in EngineCore picks it up and places it in the internal input queue.

Step 3: Scheduling

On the next engine step, the scheduler processes the input queue. The new request enters the waiting deque. The scheduler then builds the next batch:

  1. Prioritize running requests (decode): these are already generating and need just 1 token each
  2. Admit waiting requests (prefill): FCFS order, check if enough KV blocks are available
  3. Enforce budget: total tokens in batch must not exceed max_num_batched_tokens

The scheduler calls kv_cache_manager.allocate_slots() for each admitted request. This computes the number of blocks needed, checks availability, and assigns physical blocks from the free queue.

Step 4: Prefill

The prefill phase processes all prompt tokens in one forward pass. This is compute-bound: the GPU runs matrix multiplications through every transformer layer, producing KV cache entries for every prompt token.

The model runner:

  1. Assembles input tensors (token IDs, position IDs, block table mappings)
  2. Runs the forward pass through all transformer layers
  3. At each layer, computes Q, K, V projections; writes K and V to the allocated physical blocks
  4. Returns logits for the last position only

The first output token is sampled from these logits and streamed immediately. This is the time-to-first-token (TTFT) — the metric users feel most.

Step 5: Decode Loop

The request moves to the running list. Each subsequent engine step:

  1. The scheduler includes this request in the decode batch (1 token input)
  2. ModelRunner runs a forward pass — attention reads from all cached KV blocks, writes the new KV entry
  3. If the current block is full (16 tokens used), a new physical block is allocated
  4. The sampler picks the next token using temperature/top-p/top-k
  5. The token is streamed to the client

This repeats until a stop condition fires: EOS token, max_tokens reached, or a stop string matched.

Step 6: Cleanup

When the sequence finishes:

  • Status becomes FINISHED_STOPPED or FINISHED_LENGTH_CAPPED
  • All physical blocks are returned to the free pool
  • Final response chunk with usage statistics is sent
  • The request is removed from the running list

The Scheduler In Detail

The scheduler is the most complex component. It decides, at every step, which requests get GPU time. Get it wrong and you either waste GPU cycles or starve requests.

Two Workload Types

Prefill and decode have fundamentally different computational profiles:

PropertyPrefillDecode
Tokens processedHundreds to thousands1 per sequence
BottleneckCompute (FLOPS)Memory bandwidth (HBM)
DurationMilliseconds to secondsSub-millisecond per token
KV cache behaviorWrite many entriesWrite one entry, read all

This matters because a long prefill can monopolize the GPU and spike latency for all decode requests waiting in the batch.

Continuous Batching

Traditional (static) batching:

  1. Collect N requests
  2. Pad all to the same length
  3. Run forward passes until all sequences finish
  4. Return results

If one sequence generates 10 tokens and another generates 500, the short one wastes GPU time for 490 steps.

Continuous batching (what vLLM does):

  • At every decode step, the scheduler can add new requests and remove finished ones
  • The batch composition changes every iteration
  • Finished sequences immediately free their resources for new work
  • New requests can start prefill while other requests are mid-decode

This is implemented by flattening all active sequences into a single concatenated tensor. Custom PagedAttention kernels use position indices and block tables to ensure each sequence only attends to its own tokens.

Chunked Prefill

Long prompts create a latency spike: a 4096-token prefill might take 100ms, during which no decode tokens are generated for other requests.

Chunked prefill (long_prefill_token_threshold) splits long prompts into multiple steps:

Step 0: Prefill tokens 0-1023 of Request A + Decode for Requests B, C, D
Step 1: Prefill tokens 1024-2047 of Request A + Decode for Requests B, C, D
Step 2: Prefill tokens 2048-3071 of Request A + Decode for Requests B, C, D
Step 3: Prefill tokens 3072-4095 of Request A + Decode for Requests B, C, D
Step 4: Decode for Requests A, B, C, D

Decode latency stays predictable. The cost is slightly higher TTFT for the chunked request.

Preemption

When the GPU runs out of KV blocks and a running sequence needs to grow, the scheduler must preempt. The victim is the lowest-priority (most recently arrived under FCFS) request.

Two strategies:

  • Recompute (default): Free the victim's KV blocks entirely. When it is rescheduled, it goes through prefill again. Simple, no CPU memory needed.
  • Swap: Asynchronously copy the victim's KV blocks from GPU to CPU via cudaMemcpyAsync on a separate CUDA stream. When rescheduled, blocks are swapped back. Avoids recomputation but costs PCIe bandwidth and CPU memory.

The scheduler maintains three queues:

  • waiting: requests that have not started prefill
  • running: requests currently generating tokens
  • swapped: requests whose KV cache has been moved to CPU
flowchart LR
    W[Waiting Queue] -->|admit + prefill| R[Running Queue]
    R -->|preempt: swap| S[Swapped Queue]
    R -->|preempt: recompute| W
    S -->|swap in| R
    R -->|finish| F[Done: free blocks]

KV Cache Memory Management

At startup, vLLM profiles GPU memory usage to determine how many KV blocks can fit:

num_blocks = (total_gpu_memory × gpu_memory_utilization - model_weights - overhead)
             ÷ (2 × block_size × num_layers × num_kv_heads × head_dim × dtype_bytes)

For example, with a 7B model in FP16 on an 80GB A100 at 90% utilization:

  • ~14 GB for weights
  • ~57 GB available for KV cache
  • block_size=16, 32 layers, 8 KV heads, 128 head_dim, 2 bytes
  • Each block: 2 × 16 × 32 × 8 × 128 × 2 = ~2 MB
  • Total: ~28,000 blocks

All physical blocks are pre-allocated as contiguous GPU tensors at startup. The block allocator is just a free list — allocation and deallocation are O(1) pointer operations, no GPU memory allocation calls at serving time.

Prefix Caching (Automatic Prefix Caching)

Many production workloads share prefixes: system prompts, few-shot examples, RAG document preambles. Recomputing the same KV cache for the same tokens is pure waste.

vLLM's Automatic Prefix Caching (APC) works at the block granularity:

  1. When a block is full (all 16 token slots used), compute a hash of the token IDs in that block
  2. Store the mapping: hash → physical_block in a global hash table
  3. For new requests, hash each block-worth of prompt tokens and check the table
  4. Cache hit: reuse the existing physical block (increment reference count, skip prefill for those tokens)
  5. Cache miss: allocate a new block and compute KV normally
  6. Eviction: LRU when memory pressure occurs

This is particularly effective for:

  • System prompts shared across all requests (common in chatbots)
  • Multi-turn conversations where chat history is repeated
  • RAG with common document chunks

Model Weight Loading

Weight loading happens at startup and is a pipeline in itself.

flowchart LR
    A[config.json] --> B[Resolve Architecture]
    B --> C[Load Checkpoint Shards]
    C --> D[Name Mapping]
    D --> E[TP Sharding]
    E --> F[Quantization Decode]
    F --> G[Device Placement]
    G --> H[Model Ready]

The Pipeline

  1. Architecture resolution: vLLM maintains a registry mapping HuggingFace architecture names (e.g., LlamaForCausalLM) to vLLM model classes. The registry lives in vllm/model_executor/models/__init__.py.

  2. Checkpoint loading: Read safetensors or PyTorch checkpoint shards. Map tensor names from HuggingFace format to vLLM's internal module names (each model class defines a load_weights() method that handles this translation).

  3. Tensor parallelism sharding: When using TP, each weight is sliced during loading:

    • Column-parallel (QKV projections, gate/up in MLP): split along out_features. Each GPU gets out_features / tp_size rows.
    • Row-parallel (output projection, down projection): split along in_features. Each GPU gets in_features / tp_size columns.
  4. Quantization: If the model is quantized (GPTQ, AWQ, FP8, etc.), the loader reads packed weights and sets up the appropriate dequantization kernels (Marlin, ExLLama, or native FP8 CUDA cores).

  5. Device placement: Tensors are placed in GPU memory. The model is ready. Weights stay resident for the entire serving lifetime — they are static and shared across all requests.

The fundamental split is:

  • Static: model weights (loaded once, never change)
  • Dynamic: KV cache (allocated and freed per-request per-step)

Tensor Parallelism

For models too large for one GPU, vLLM uses Megatron-LM style tensor parallelism. The model is split across GPUs at the level of individual linear layers.

flowchart TD
    subgraph "Each Transformer Layer"
        I[Input Hidden State] --> QKV[ColumnParallel: Q, K, V Projections]
        QKV --> ATT[Attention: each GPU handles num_heads/TP heads]
        ATT --> OP[RowParallel: Output Projection]
        OP --> AR1[AllReduce]
        AR1 --> GATE[ColumnParallel: Gate + Up Projections]
        GATE --> ACT[SiLU × Element-wise Multiply]
        ACT --> DOWN[RowParallel: Down Projection]
        DOWN --> AR2[AllReduce]
        AR2 --> O[Output Hidden State]
    end

Each transformer layer has exactly two AllReduce operations: one after the attention output projection and one after the MLP down projection. On NVLink systems (A100 SXM, H100 SXM), AllReduce runs at 600+ GB/s bidirectional and adds minimal overhead. On PCIe systems, this becomes the bottleneck.

For multi-node setups, vLLM supports Ray-based distribution via RayGPUExecutor, with NCCL handling cross-node communication.

Performance Optimizations

vLLM layers multiple optimizations that compound:

FlashAttention for Prefill

Prefill processes many tokens and is compute-bound. vLLM uses FlashAttention (Dao et al.) which:

  • Never materializes the full [seq_len, seq_len] attention matrix
  • Uses tiling and recomputation for O(1) extra memory instead of O(n²)
  • Achieves near-peak FLOPS utilization on modern GPUs

For decode (which is memory-bandwidth-bound and needs block table support), vLLM's own PagedAttention kernels are used instead.

CUDA Graphs

Each decode step is very fast (sub-millisecond). At this speed, CPU-side kernel launch overhead becomes significant. CUDA graphs solve this:

  1. During warmup, vLLM captures forward passes for common batch sizes (1, 2, 4, 8, ..., up to max_num_seqs) as CUDA graphs
  2. During serving, instead of launching individual kernels, the entire captured graph is replayed — one CPU call triggers the whole GPU pipeline
  3. The batch size is padded to the nearest captured size

This eliminates per-kernel launch overhead and CPU-GPU synchronization stalls.

Fused Kernels

vLLM implements custom CUDA kernels that fuse multiple operations:

  • RMSNorm + residual addition
  • SiLU activation + element-wise multiply (for gated MLPs)
  • Rotary position embedding application

Each fusion eliminates a kernel launch and a round-trip to GPU global memory.

FP8 KV Cache

On Hopper and Ada GPUs, vLLM can store KV cache in FP8 instead of FP16. This doubles the number of concurrent sequences at the same memory budget with minimal quality degradation.

Speculative Decoding

Autoregressive decoding is inherently sequential: you cannot generate token N+1 without knowing token N. Speculative decoding breaks this bottleneck.

The idea: use a small, fast draft model to propose K candidate tokens. Then verify all K tokens in a single forward pass of the large target model. If the target model agrees, you get K tokens for the cost of one target model inference.

The Algorithm

1. Draft model generates K tokens autoregressively  (fast, small model)
2. Target model runs ONE forward pass on [context + K draft tokens]
3. For each position i = 0 to K-1:
   - If p_target(token_i) >= p_draft(token_i): accept
   - Else: accept with probability p_target(token_i) / p_draft(token_i)
   - If rejected: sample from adjusted distribution max(0, p_target - p_draft), stop
4. Always sample one bonus token from the target model at the last accepted position

The key mathematical property: the output distribution is identical to the target model. Speculative decoding is not an approximation. It is exact.

Configuration example:

python -m vllm.entrypoints.openai.api_server \
  --model meta-llama/Meta-Llama-3-70B \
  --speculative-model meta-llama/Meta-Llama-3-8B \
  --num-speculative-tokens 5

vLLM also supports draft-free speculation methods:

  • N-gram speculation: uses n-gram matching on the prompt itself to predict future tokens, no separate model needed
  • Medusa: multiple decoding heads predict future tokens in parallel
  • EAGLE: a lightweight draft head trained to predict next-token embeddings

V1 Architecture

The V1 rewrite (started late 2024) addresses architectural debt from rapid V0 growth. The key changes:

Process Isolation

V1 runs tokenization, engine core, and API serving in separate processes. This eliminates GIL contention — the most impactful single change for throughput on CPU-limited workloads.

flowchart LR
    subgraph "Process 1: API Server"
        FS[FastAPI + AsyncLLM]
        TOK[Tokenizer]
        DETOK[Detokenizer]
    end
    subgraph "Process 2: EngineCore"
        SCH2[Scheduler]
        KVM[KV Cache Manager]
    end
    subgraph "Process 3+: Workers"
        W[ModelRunner + GPU]
    end
    FS <-->|ZMQ IPC| SCH2
    SCH2 <-->|Shared Memory MQ| W

Simplified Scheduler

The V1 scheduler is leaner. It maintains a waiting_deque and a running_list, uses a simpler priority model, and integrates prefix caching more deeply into the scheduling decision.

Disaggregated Prefill/Decode

V1 supports separating prefill and decode onto different GPU pools:

  • Prefill workers: handle compute-heavy prompt processing at high throughput
  • Decode workers: handle memory-bandwidth-heavy generation with low latency
  • KV cache is transferred between pools via a KV-connector abstraction (shared storage, RDMA, or NVLink)

This allows independent scaling: if your workload has long prompts but short outputs, allocate more prefill GPUs.

Multi-Node Data Parallelism

V1 supports data parallelism with load balancing. A DPLBAsyncMPClient distributes requests across multiple engine instances using a scoring heuristic:

score = len(waiting) × 4 + len(running)

The engine with the lowest score gets the next request.

Debugging and Production Tuning

When latency regresses, decompose it:

MetricWhat it measuresDiagnosis
TTFTTime to first tokenPrefill throughput, queue wait time
ITLInter-token latencyDecode step time, batch size pressure
Queue depthWaiting queue lengthAdmission rate vs processing rate
KV block utilization% of blocks in useMemory pressure, eviction frequency

Key tuning knobs:

  • max_num_seqs: max concurrent sequences. Higher → more throughput, higher ITL.
  • max_num_batched_tokens: token budget per step. Controls prefill chunking and batch size.
  • gpu_memory_utilization: fraction of GPU memory for KV cache. Default 0.9.
  • tensor_parallel_size: number of GPUs. Reduces per-GPU memory, adds communication overhead.
  • enable_prefix_caching: enable APC for workloads with shared prefixes.
  • enable_chunked_prefill: prevent long prefills from spiking decode latency.

The performance model follows a roofline: below a saturation batch size B_sat, step time is dominated by HBM bandwidth (memory-bound). Above B_sat, it becomes compute-bound. The optimal operating point sits at the knee of this curve.

Summary

vLLM is not "faster inference." It is a serving runtime that applies operating system principles to GPU memory management and process scheduling.

The core insights are simple:

  1. Page the KV cache (PagedAttention) — reduce memory waste from 60-80% to under 4%
  2. Schedule continuously — add and remove requests every iteration, not every batch
  3. Separate static and dynamic memory — weights are loaded once; KV cache is allocated per-request
  4. Isolate processes — tokenization, scheduling, and GPU execution run independently

Every design choice flows from one constraint: GPU memory is finite and expensive. vLLM makes it possible to serve more users with fewer GPUs. That is the entire point.


References: Efficient Memory Management for Large Language Model Serving with PagedAttention (Kwon et al., SOSP 2023), vLLM Blog, Anatomy of vLLM, Life of an Inference Request (vLLM V1), vLLM GitHub.