This calculator ranks graphic processing units by true financial cost per million tokens rather than hourly rental price alone. It evaluates hourly instance pricing, VRAM capacity, peak token throughput, and realistic operational utilization percentages. Infrastructure engineers use it to select optimal GPU hardware for self-hosted LLM inference clusters.
Loading calculator...
Selecting GPU hardware based strictly on lowest hourly rental price leads to poor unit economics. A more expensive GPU delivering significantly higher token throughput often produces a far lower cost per million tokens.
How to use it
Set realistic operational utilization percentage to account for traffic fluctuations. Peak hardware throughput is rarely sustained continuously in production; 30% to 50% utilization represents realistic production traffic loads.
Review the customizable GPU comparison table containing model names, VRAM capacities in gigabytes, hourly instance rental prices in dollars, and peak token throughput ratings in tokens per second.
Measure peak token generation speeds on your specific model checkpoint and batch size to enter accurate throughput figures into the comparison matrix.
Add new custom GPU options or delete existing rows as pricing and hardware specs change. The tool automatically ranks hardware options by unit cost per million tokens, highlighting the most cost-effective GPU choice.
Fields explained
Realistic utilization (%) – operational capacity utilization factor accounting for real-world traffic fluctuations. Default value is 40, step size 1, range 0 to 100.
GPU Name – descriptive text identifier for the target GPU hardware instance (such as RTX 4090, L4, A100 80GB, or H100 80GB).
VRAM (GB) – onboard video memory capacity in gigabytes available to host model weights and KV-cache context. Default values range from 24 to 80 GB.
$/hr – hourly instance rental rate in dollars charged by cloud hosting providers. Default values range from $0.70 to $4.50/hr.
Peak tok/s – maximum theoretical or benchmarked output token generation speed across all batched streams. Default values range from 700 to 5,200 tok/s.
Reading the results
| Comparison Output | Financial Representation | Infrastructure Procurement Decision |
|---|---|---|
| $/1M tok | Effective hosting expense incurred to generate one million tokens at target utilization. | Select the GPU instance delivering the lowest true cost per token. |
| Cheapest per token summary | Highlights the top-ranked GPU name, unit token cost, and hourly token volume. | Validate node deployment choices for production cluster architectures. |
Calculating unit cost per token reveals how high-throughput hardware offsets higher hourly rental prices. Operating GPUs at low utilization increases effective unit token expenses significantly.
Choosing a GPU with insufficient VRAM capacity prevents loading model weights into memory, resulting in zero usable throughput regardless of hourly price.
Evaluating effective token throughput guides hardware procurement choices. High-throughput H100 instances deliver unit token costs up to 40 percent lower than budget L4 GPUs.
The formula
Effective token throughput scales peak tokens per second by realistic utilization percentage. Hourly token volume multiplies effective tok/s by 3,600 seconds. Unit cost per million tokens divides hourly GPU rental price by hourly token volume in millions. The results table sorts GPU candidates in ascending order of cost per million tokens.
The mathematical representation for effective throughput and hourly token volume is:
EffectiveTokSec = PeakTokSec × (UtilizationPct / 100)
HourlyTokens = EffectiveTokSec × 3600
HourlyTokensInMillions = HourlyTokens / 1,000,000
The mathematical representation for unit token cost ranking is:
CostPerMillionTokens = HourlyPrice / HourlyTokensInMillions
| GPU Architecture | Typical Hourly Rate | Primary Workload Suitability |
|---|---|---|
| Consumer / Entry (RTX 4090 / L4) | $0.70 to $0.90/hr | Low-traffic endpoints, small 7B quantized models, dev testing |
| Enterprise Frontier (A100 / H100) | $2.50 to $4.50/hr | High-concurrency production, 70B+ models, large batch serving |
Monthly instance costs calculate as hourly rental price multiplied by 720 hours (24 hours × 30 days) for always-on dedicated hosting.
For a baseline setup at 40% utilization comparing an L4 ($0.80/hr, 700 tok/s peak) against an H100 ($4.50/hr, 5,200 tok/s peak): L4 effective throughput equals 700 × 0.40 = 280 tok/s (1,008,000 tok/hr), yielding $0.794 per 1M tokens. H100 effective throughput equals 5,200 × 0.40 = 2,080 tok/s (7,488,000 tok/hr), yielding $0.601 per 1M tokens. The H100 proves 24% cheaper per token despite costing 5.6× more per hour.
Worked examples
High-Traffic Production Cluster Sizing
An infrastructure team compares RTX 4090 ($0.70/hr, 900 tok/s) against A100 80GB ($2.50/hr, 2,400 tok/s) at 50% utilization. RTX 4090 effective throughput: 450 tok/s (1.62M tok/hr), cost = $0.432/1M tokens. A100 effective throughput: 1,200 tok/s (4.32M tok/hr), cost = $0.579/1M tokens. At 50 percent utilization, the RTX 4090 delivers lower unit costs, provided model weights fit within its 24GB VRAM limit. The team selects RTX 4090 nodes for 8B model workloads.
Large 70B Parameter Model Deployment
Deploying a 70B parameter model requires 80GB VRAM. Comparing L4 (24GB VRAM – cannot fit model) against A100 80GB ($2.50/hr, 2,000 tok/s) and H100 80GB ($4.50/hr, 4,800 tok/s) at 35% utilization. A100 effective throughput: 700 tok/s (2.52M tok/hr), cost = $0.992/1M tokens. H100 effective throughput: 1,680 tok/s (6.048M tok/hr), cost = $0.744/1M tokens. H100 delivers a 25% lower cost per million tokens while meeting VRAM constraints.
Low-Traffic Dev Environment Evaluation
A dev environment operates at low 15% utilization. RTX 4090 ($0.70/hr, 900 tok/s) vs H100 ($4.50/hr, 5,200 tok/s). RTX 4090 effective throughput: 135 tok/s (0.486M tok/hr), cost = $1.440/1M tokens. H100 effective throughput: 780 tok/s (2.808M tok/hr), cost = $1.603/1M tokens. At low utilization, the lower hourly price of the RTX 4090 delivers better financial efficiency.
Batch Offline Processing Comparison
An offline batch pipeline runs continuously at 80% utilization. L4 ($0.80/hr, 700 tok/s) vs H100 ($4.50/hr, 5,200 tok/s). L4 effective throughput: 560 tok/s (2.016M tok/hr), cost = $0.397/1M tokens. H100 effective throughput: 4,160 tok/s (14.976M tok/hr), cost = $0.301/1M tokens. High utilization maximizes H100 memory bandwidth advantages, achieving minimal cost per token.
Common mistakes
Comparing GPUs based on theoretical peak TFLOPS instead of measured inference token throughput leads to incorrect hardware selection. Autoregressive LLM decoding is memory-bandwidth bound, making VRAM bandwidth far more predictive of generation speed than raw compute TFLOPS.
Ignoring VRAM memory capacity constraints causes deployment failures. Selecting a cheap GPU with 24GB VRAM to serve a 70B parameter model fails because model weights and KV-cache context exceed available hardware memory.
Assuming 100 percent GPU utilization in financial models creates unrealistic cost projections. Real-world user traffic fluctuates across daily cycles, resulting in average operational utilization between 30 and 50 percent.
Procuring GPU instances without verifying VRAM requirements results in immediate out-of-memory errors when launching target model checkpoints.
Benchmark model throughput on candidate GPU hardware using production-equivalent batch sizes before signing hosting contracts.
FAQ
Why is cost per token a better metric than price per GPU hour?
Price per GPU hour measures raw instance rental cost without accounting for work accomplished. Cost per token evaluates how many generated tokens an instance delivers per dollar spent.
High-end GPUs costing more per hour often generate tokens so much faster that unit cost per token is lower.
How does memory bandwidth affect GPU generation speed?
Autoregressive text generation requires reading model weights from GPU memory into compute cores for every single generated token. High memory bandwidth (such as 3.35 TB/s on H100 vs. 300 GB/s on L4) speeds up token generation directly.
Memory bandwidth is the primary hardware bottleneck for LLM inference decoding.
What operational utilization percentage should I use for planning?
For interactive user-facing applications with variable daily traffic, use 30 to 40 percent utilization. For steady-state batch processing or queued background workloads, use 70 to 80 percent utilization.
Adjust utilization settings in the calculator to match your application’s traffic profile.
Can consumer GPUs like the RTX 4090 be used for production inference?
Yes. Consumer GPUs offer excellent memory bandwidth per dollar for smaller models (7B to 14B parameters). However, they lack high-bandwidth interconnects (like NVLink) and enterprise cloud SLA guarantees.
Consumer GPUs work well for single-node deployments serving small quantized models.
How does model quantization affect GPU hardware requirements?
Quantizing a model (such as converting 16-bit floats to 4-bit integers) reduces model weight size by up to 75 percent. This allows larger models to fit inside smaller GPU VRAM footprints (like 24GB).
Quantization enables running larger models on cheaper GPU instances, reducing unit hosting costs.
Disclaimer
This calculator provides GPU ranking and cost per token estimates based on user-entered hourly rental rates, VRAM specs, peak throughput ratings, and utilization factors. Actual inference performance depends on specific cloud provider instance configurations, model parameter precision, batch size settings, attention framework optimizations (like vLLM), and real-time network traffic patterns.
The interactive calculator on this page serves as the primary tool for testing hardware scenarios and ranking GPU options. Infrastructure teams should run benchmark load tests on candidate GPU instances using production model checkpoints before finalizing hardware procurement contracts.







