Weights
- What
- The learned numbers: every matrix W and bias b in the network.
- Lives
- The whole run. Loaded once, nudged every step, saved in checkpoints.
- Size
- Parameter count × bytes per parameter. The same at batch size 1 or 1,000.
Training memory, explained
Training a network puts four kinds of tensors on the GPU: weights, activations, gradients and optimizer state. They are created and freed at different moments, and each one grows with something different. This page shows when each exists and lets you size them for your own model.
One parameter, trained with mixed-precision Adam, costs 16 bytes.
Each cell is one byte. This is the cost before a single activation is stored. A 7 billion parameter model needs about 112 GB for these four items alone.
Everything the framework allocates for training falls into one of these buckets, plus a few temporary buffers. The colors below are used for the same thing everywhere on the page.
The word "activation" is easy to misread. An activation function such as ReLU is an operation and holds no memory. The activations are the tensors it produces, and those are what backward needs later. Other buffers also exist: activation gradients that flow backward for a moment, kernel workspaces, the input batch, and at inference time the KV cache. They sit in the calculator under temporary buffers and KV cache.
Layer i computes zi = Wi · ai−1 and then ai = f(zi). During backward, the error arriving at that layer is δi = ∂L/∂zi. Two products come out of it.
Needs the activation ai−1 that layer i saw on the way forward. That is why forward cannot throw its outputs away.
Needs the weights Wi, so they stay resident throughout backward.
The first product has one entry per weight, so gradients match the weights in size. The saved activations have one entry per layer output per sample, so they grow when you add samples or tokens. That difference explains most memory behavior you will see below.
A three-layer network, nine moments. F is forward through a layer, B is backward through a layer. Drag the slider or click a column. A colored cell means the tensor is in memory at the end of that moment.
Choose a transformer and a batch. The chart below uses the same nine moments as above, with the network split into thirds. Results are estimates in GiB (1 GiB = 1,024³ bytes) that follow the formulas in the papers listed at the end.
The largest size each kind reaches, and the rule that produced it for your settings.
Each technique attacks one bucket and pays for it somewhere else.
| Lever | What shrinks | What it costs |
|---|---|---|
| Smaller batch per GPU | Saved activations, in proportion to batch size | Lower GPU utilization, or more gradient accumulation steps |
| Gradient accumulation | Nothing directly. It gives a large effective batch while holding activations at the small batch size | More forward and backward passes per update |
| Activation checkpointing | Saved activations drop to one input per layer plus one layer being recomputed | About one extra forward pass of compute |
| FlashAttention | The s × s attention scores, which otherwise grow with the square of the sequence length | None in practice, it is a fused kernel |
| Sharding (ZeRO-3, FSDP) | Weights, gradients and optimizer state, divided across GPUs | Network traffic between GPUs |
| 8-bit or bf16 optimizer state | Optimizer state, by 2× to 4× | Some risk to training stability |
| LoRA or frozen layers | Gradients and optimizer state exist only for the trainable parameters | Less capacity to change than full fine-tuning |
| CPU or disk offload | Optimizer state and weights move to host memory | Slow transfers over PCIe |
To check an estimate against a real run, reset the peak counter, run one full training step, and read it back.
torch.cuda.reset_peak_memory_stats() # one full step: forward, backward, optimizer.step(), optimizer.zero_grad() print(torch.cuda.max_memory_allocated() / 2**30, "GiB")
The CUDA context (a few hundred MiB) and allocator fragmentation sit on top of that number in nvidia-smi.