Model memory growth calculator – explain training memory expansion

Model memory growth calculator – explain training memory expansion Calculators

This calculator models GPU memory footprint expansion across inference, LoRA parameter-efficient fine-tuning, and full Adam optimizer training workloads. It breaks memory allocation into model weights, gradients, optimizer states, and activation tensors. Machine learning engineers use it to understand why model training requires far more GPU memory than inference serving.

Loading calculator...

A model that executes inference smoothly in 14 GB of VRAM requires over 100 GB of memory for full fine-tuning. Tracking memory growth across weights, gradients, and Adam optimizer states explains this massive footprint expansion.

How to use it

Enter your target model’s parameter count in billions. Select your target workload mode from Inference (weights only), LoRA fine-tune (~1%), or Full fine-tune (Adam).

Input estimated activation memory footprint in gigabytes based on your expected batch size and sequence length.

Remember that activation memory scales directly with batch size and sequence length; use gradient checkpointing to keep activation footprints low.

The dashboard displays total required GPU memory in gigabytes, bytes per parameter ratio, static weight footprint, gradient plus optimizer state memory, and activation memory.

Fields explained

Parameters (billions) – total parameter count of the target model in billions. Default value is 7.0, step size 0.1.

Workload – target compute workload type selection. Options: Inference (weights only) (inference), LoRA fine-tune (~1%) (lora), Full fine-tune (Adam) (full). Default selection is Full fine-tune (Adam).

Activation memory (GB) – estimated memory allocated for intermediate layer activation tensors. Default value is 6.0, step size 0.5.

Reading the results

Memory ComponentFP16 Byte Footprint per ParameterWorkload Footprint Scope
Weights2 bytes / param (FP16)Static parameter weights loaded into VRAM across all workload modes.
Gradients2 bytes / param (FP16)Gradient tensors stored during backward passes (active in training modes).
Optimizer States12 bytes / param (Adam FP32)First/second momentum vectors (8B) plus master weights (4B) in Adam optimizer.
ActivationsBatch & sequence dependentIntermediate layer activations stored for backpropagation gradient steps.

Full training using mixed-precision Adam requires approximately 16 bytes per parameter before accounting for activation tensors. LoRA fine-tuning freezes base model weights, restricting optimizer state memory strictly to trainable adapter weights.

Attempting full fine-tuning on a 7B model without multi-GPU tensor parallelism or ZeRO optimizer sharding causes immediate out-of-memory crashes.

Understanding memory growth components helps engineers select parameter-efficient training methods. Switching from full fine-tuning to LoRA on a 7B model cuts memory requirements from 118 GB down to 20 GB.

The formula

Inference mode allocates 2 bytes per parameter for FP16 weights plus activation memory. Full fine-tuning adds 2 bytes for FP16 gradients and 12 bytes for Adam optimizer states (8 bytes for momentum vectors plus 4 bytes for FP32 master weights), totaling 16 bytes per parameter before activations. LoRA fine-tuning freezes base weights, applying gradient and optimizer memory strictly to trainable adapter parameters (~1% of total parameters).

The mathematical representation for memory components across workload modes is:

WeightsGB = ParametersInBillions × 2

InferenceTotalGB = WeightsGB + ActivationGB

FullTrainingTotalGB = (ParametersInBillions × 16) + ActivationGB

LoraTrainingTotalGB = WeightsGB + (ParametersInBillions × 0.01 × 14) + ActivationGB

The mathematical representation for per-parameter memory ratio is:

BytesPerParameter = TotalGB / ParametersInBillions

Workload ModeEffective Bytes / Parameter7B Model Memory (+6GB Activations)
Inference (FP16)2 bytes + activations20.0 GB total VRAM required
LoRA Fine-Tune (~1%)~2.14 bytes + activations21.0 GB total VRAM required
Full Adam Fine-Tune16 bytes + activations118.0 GB total VRAM required

DeepSpeed ZeRO-3 optimizer sharding divides optimizer states, gradients, and model weights across multiple GPUs, reducing per-GPU memory overhead proportionally.

For a baseline setup with a 7B model, full Adam fine-tuning, and 6GB activation memory: Weights equal 7 × 2 = 14.0 GB. Gradients equal 7 × 2 = 14.0 GB. Optimizer states equal 7 × 12 = 84.0 GB. Combined gradients and optimizer states total 98.0 GB. Total memory required equals 14.0 + 98.0 + 6.0 = 118.0 GB (16.9 bytes/param).

Worked examples

Full Fine-Tuning a 70B Foundation Model

A research team evaluates full Adam training for a 70B model with 20GB activation memory. Parameters: 70B, full workload, 20GB activations. Weights: 70 × 2 = 140.0 GB. Gradients: 140.0 GB. Optimizer states: 70 × 12 = 840.0 GB. Gradients plus optimizer: 980.0 GB. Full fine-tuning a 70B model requires 1,140.0 GB of VRAM, mandating a cluster of 16 × 80GB A100 GPUs.

LoRA Fine-Tuning a 70B Model on Single Node

Switching the 70B model to LoRA fine-tuning freezes base weights. Parameters: 70B, lora workload, 20GB activations. Base weights: 140.0 GB. LoRA adapter gradients and optimizer states: 70 × 0.01 × 14 = 9.8 GB. Total memory: 140.0 + 9.8 + 20.0 = 169.8 GB. LoRA allows fine-tuning 70B models across just 3 × 80GB GPUs instead of 16.

Inference Serving for an 8B Model

Serving an 8B model for inference with 4GB activation memory. Parameters: 8B, inference workload, 4GB activations. Weights: 8 × 2 = 16.0 GB. Gradients + optimizer: 0 GB. Activations: 4.0 GB. Total memory: 16.0 + 4.0 = 20.0 GB. Fits comfortably on a single 24GB RTX 3090/4090 GPU card.

LoRA Fine-Tuning an 8B Model on Consumer Hardware

Fine-tuning an 8B model using QLoRA (4-bit quantized base weights). Parameters: 8B, lora workload, 3GB activations. Quantized base weights: 8 × 0.5 = 4.0 GB. LoRA adapter states: ~1.1 GB. Activations: 3.0 GB. Total memory: ~8.1 GB. QLoRA enables fine-tuning 8B models on single consumer GPUs with 12GB VRAM.

Common mistakes

Assuming a model that fits in VRAM during inference can be fine-tuned without additional memory is a major mistake. Full Adam training requires eight times more memory per parameter than FP16 inference serving due to optimizer momentum vectors.

Ignoring activation memory growth during batch size increases causes out-of-memory crashes. Activations scale linearly with batch size and sequence length; storing intermediate activations without gradient checkpointing rapidly consumes available VRAM.

Failing to use parameter-efficient fine-tuning (PEFT) methods like LoRA or QLoRA for domain adaptation wastes massive GPU compute budget on full weight updates.

Attempting full Adam fine-tuning without gradient checkpointing or ZeRO memory sharding causes immediate out-of-memory crashes during backward passes.

Enable activation gradient checkpointing to trade minor compute overhead for substantial activation memory savings.

FAQ

Why does full Adam fine-tuning require 16 bytes per parameter?

Mixed-precision Adam training stores FP16 model weights (2 bytes), FP16 gradients (2 bytes), FP32 first momentum vectors (4 bytes), FP32 second momentum vectors (4 bytes), and FP32 master weight copies (4 bytes), totaling 16 bytes per parameter.

Optimizer state memory dominates total VRAM consumption during full training.

How does LoRA (Low-Rank Adaptation) reduce training memory?

LoRA freezes base model weights and inserts small trainable rank-decomposition adapter matrices into transformer layers. Because only ~1% of parameters update, optimizer states and gradients track only the adapter weights.

LoRA reduces training memory footprints down near inference baseline levels.

What is activation memory and how can it be reduced?

Activation memory stores intermediate tensor outputs generated during the forward pass, which are needed later to calculate gradients during backpropagation. Activations scale with batch size and sequence length.

Gradient checkpointing discards intermediate activations during the forward pass and re-computes them as needed during backpropagation, saving 60 to 80 percent of activation memory.

What is QLoRA and how does it enable fine-tuning on consumer GPUs?

QLoRA quantizes frozen base model weights down to 4-bit precision (NormalFloat4) while keeping 16-bit LoRA adapter weights. This reduces static model weight memory by 75 percent.

QLoRA allows fine-tuning 7B and 13B models on single consumer GPUs with 12GB to 24GB of VRAM.

What is DeepSpeed ZeRO and how does it shard memory?

DeepSpeed Zero Redundancy Optimizer (ZeRO) partitions training memory across multi-GPU clusters. ZeRO-1 shards optimizer states, ZeRO-2 shards gradients, and ZeRO-3 shards model weights across GPUs.

ZeRO-3 eliminates memory redundancy, enabling full training of massive models across GPU clusters.

Disclaimer

This calculator provides model memory footprint growth estimates based on standard mathematical formulas for mixed-precision FP16 training and Adam optimizer states. Actual GPU memory requirements vary based on specific deep learning framework runtimes (PyTorch, DeepSpeed, Megatron), activation checkpointing settings, sequence length settings, KV-cache allocations, and CUDA driver memory overhead.

The interactive calculator on this page serves as the primary tool for testing memory growth scenarios and hardware planning. Machine learning engineers should run benchmark training steps on target hardware to measure exact memory allocation peaks before initiating long-term fine-tuning jobs.

Rate article
Ai review
Add a comment