This calculator models the relationship between batch size, generation speed, aggregate token throughput, and hosting costs for self-hosted LLM inference. It accounts for sublinear scaling efficiency caused by memory bandwidth bottlenecks and compute limits. Infrastructure engineers use it to balance per-user latency against cost per million tokens.
Loading calculator...
Increasing batch size improves total GPU token throughput and lowers unit costs, but individual stream speeds slow down. Finding the optimal batching threshold prevents user-facing latency degradation while maximizing hardware efficiency.
How to use it
Enter your hardware’s baseline single-stream generation speed in tokens per second. Specify your target batch size representing concurrent processing streams on a single GPU.
Set the scaling efficiency percentage to account for sublinear scaling overhead. Typical values range between 50% and 80% depending on GPU architecture and KV-cache optimization.
Measure single-stream generation speed under un-batched conditions to establish an accurate baseline for scaling calculations.
Input your GPU hourly hosting rate. The calculator displays aggregate throughput, speedup factor, per-stream generation speed, total hourly token volume, and unit cost per million tokens.
Fields explained
Single-stream speed (tok/s) – baseline output generation speed for a single isolated request. Default value is 60, step size 1.
Batch size – number of concurrent inference streams processed simultaneously on the GPU. Default value is 16, step size 1.
Scaling efficiency (%) – percentage of linear scaling achieved as batch size expands. Default value is 65, step size 1, range 0 to 100.
GPU cost per hour – hourly hosting expense for the GPU instance in dollars. Default value is 2.50, step size 0.01.
Reading the results
| Result Metric | What It Measures | Engineering Takeaway |
|---|---|---|
| Aggregate throughput | Total tokens generated per second across all active batch streams. | Use this figure to size cluster capacity for total system demand. |
| Per-stream speed | Generation speed experienced by each individual user stream. | Ensure per-stream tok/s meets interactive user experience thresholds. |
| Cost / 1M tokens | Effective hosting expense per million generated tokens. | Compare unit costs against commercial API provider pricing. |
| Tokens / hour | Total token output volume produced by the GPU instance per hour. | Establish maximum hourly capacity for batch processing pipelines. |
Increasing batch size yields sublinear aggregate speedup while reducing per-stream generation speeds. Higher batch sizes lower unit token costs up to hardware memory limits.
Setting batch sizes too high degrades per-stream speeds below acceptable human reading thresholds, causing choppy user experiences.
Finding the right balance between speed and batch density is critical for serving production traffic efficiently. Scaling batch size from 1 to 16 reduces cost per million tokens by over 80 percent.
The formula
Aggregate throughput uses a sublinear scaling model where additional batch slots contribute a fraction of single-stream performance determined by scaling efficiency. Per-stream speed divides aggregate throughput by batch size. Hourly token volume scales aggregate speed by 3,600 seconds. Unit cost divides hourly GPU expenses by total hourly token throughput in millions.
The mathematical representation for aggregate throughput and stream performance is:
AggregateTokSec = SingleStreamTokSec × (1 + (BatchSize - 1) × (ScalingEfficiency / 100))
PerStreamTokSec = AggregateTokSec / BatchSize
SpeedupFactor = AggregateTokSec / SingleStreamTokSec
The mathematical representation for token volume and unit hosting cost is:
TokensPerHour = AggregateTokSec × 3600
CostPerMillion = GPUCostPerHour / (TokensPerHour / 1,000,000)
| Batch Setting | Primary Advantage | Trade-Off / Constraint |
|---|---|---|
| Small Batch (1–4) | Ultra-fast per-stream latency and high user responsiveness | High unit cost per token and low GPU compute utilization |
| Large Batch (16–64) | Maximum hardware throughput and minimal cost per token | Slower per-stream generation and high VRAM KV-cache overhead |
Scaling efficiency reflects memory bandwidth limits; autoregressive LLM decoding is memory-bound, causing efficiency to drop as batch size grows.
For a baseline setup with 60 tok/s single-stream speed, batch size 16, 65% scaling efficiency, and $2.50/hr GPU cost: aggregate throughput equals 60 × (1 + 15 × 0.65) = 645 tok/s (10.75× speedup). Per-stream speed measures 40.31 tok/s. Total output reaches 2,322,000 tokens/hr, delivering an effective unit cost of $1.077 per million tokens.
Worked examples
Real-Time Interactive Chat Application
A chat product requires fast streaming latency for interactive users. Single-stream speed: 50 tok/s. Batch size: 4. Scaling efficiency: 80%. GPU cost: $2.00/hr. Aggregate throughput equals 50 × (1 + 3 × 0.80) = 170 tok/s (3.4× speedup). Per-stream speed remains snappy at 42.5 tok/s. Output reaches 612,000 tokens/hr, yielding a unit cost of $3.268 per million tokens. The team accepts higher unit costs to maintain low user latency.
High-Throughput Offline Batch Processing
An enterprise processes millions of document summaries overnight where latency is unconstrained. Single-stream speed: 70 tok/s. Batch size: 32. Scaling efficiency: 55%. GPU cost: $3.00/hr. Aggregate throughput reaches 70 × (1 + 31 × 0.55) = 1,263.5 tok/s (18.05× speedup). Per-stream speed drops to 39.49 tok/s. Hourly output reaches 4,548,600 tokens, driving unit cost down to just $0.659 per million tokens. The team optimizes batch size for maximum financial savings.
Edge GPU Deployment with Limited VRAM
A local server deploys a small quantized model on desktop hardware. Single-stream speed: 40 tok/s. Batch size: 8. Scaling efficiency: 70%. GPU cost: $0.80/hr. Aggregate throughput equals 40 × (1 + 7 × 0.70) = 236 tok/s (5.9× speedup). Per-stream speed measures 29.5 tok/s. Total output equals 849,600 tokens/hr, achieving a unit cost of $0.942 per million tokens. The configuration maximizes utility without overflowing VRAM limits.
Un-Batched Baseline Comparison
Evaluating an un-batched single-stream setup highlights hardware inefficiency. Single-stream speed: 60 tok/s. Batch size: 1. Scaling efficiency: 100%. GPU cost: $2.50/hr. Aggregate throughput equals 60 tok/s (1.0× speedup). Per-stream speed hits maximum at 60.0 tok/s. Total output equals 216,000 tokens/hr, resulting in a unit cost of $11.574 per million tokens. Infrastructure engineers use this baseline to justify deploying dynamic batching middleware.
Common mistakes
Assuming linear throughput scaling leads to severe capacity planning errors. Doubling batch size from 16 to 32 does not double token throughput because memory bandwidth bottlenecks reduce scaling efficiency at higher batch densities.
Neglecting KV-cache memory requirements when increasing batch size causes out-of-memory crashes. Large batch sizes allocate substantial VRAM to store attention context for every active stream, reducing available memory for model weights.
Failing to implement continuous or dynamic batching middleware wastes GPU compute. Traditional static batching waits for all streams to finish before accepting new requests, creating idle compute windows when request lengths vary.
Pushing batch sizes beyond GPU memory limits causes severe latency spikes due to swapping or triggers sudden out-of-memory process terminations.
Use modern serving frameworks with vLLM or PagedAttention to maximize memory efficiency and batch capacity.
FAQ
Why does per-stream speed decrease as batch size increases?
GPUs decode tokens autoregressively, requiring model weights to load from VRAM into compute cores for every single generated token. Larger batches share weight-loading overhead but increase memory access contention, slowing individual stream completion times.
While aggregate output across all users increases, each individual user receives tokens slightly slower.
What is a good target for scaling efficiency in LLM serving?
Typical scaling efficiency for autoregressive LLM decoding ranges from 50% to 75% on modern GPU architectures. Specialized serving engines using dynamic PagedAttention achieve higher efficiency than standard PyTorch baselines.
Scaling efficiency drops as batch size expands and memory bandwidth becomes saturated.
How does dynamic batching differ from static batching?
Static batching groups fixed requests together and processes them until the longest request completes, leaving compute idle as shorter requests finish early. Dynamic batching inserts incoming requests into running batches immediately as individual streams complete.
Dynamic batching improves aggregate throughput by eliminating idle compute gaps between variable-length responses.
How does batch size impact GPU memory (VRAM)?
Each active stream in a batch requires VRAM memory to store its key-value (KV) cache for prompt context and generated tokens. Doubling batch size doubles total KV-cache memory requirements.
If KV-cache memory exceeds available VRAM, serving engines must limit maximum batch size or offload memory, degrading performance.
When should I prioritize low batch size over high throughput?
Prioritize low batch sizes (1 to 4) for interactive real-time applications like voice agents, live coding assistants, or conversational chat where immediate user latency takes priority over unit hosting costs.
Use larger batch sizes for background processing, data extraction, document indexing, and offline evaluation workloads.
Disclaimer
This calculator provides throughput and cost estimates based on simplified sublinear scaling models. Actual inference performance depends on specific GPU hardware architectures, memory bandwidth, model parameter precision, prompt lengths, KV-cache management, and serving framework optimizations.
The interactive calculator on this page serves as the primary tool for testing batching scenarios and hardware configurations. Infrastructure teams should run benchmark load tests on target hardware using production-equivalent prompts before establishing final deployment parameters.








Where does the calculator store my input data? Does it send batch size, GPU costs, or generation speeds to your servers? I need self-hosted inference specifically to avoid cloud logging, so if this tool phones home with my infrastructure specs, that defeats the purpose. Are there privacy guarantees in the documentation?
Great question on data handling. The batch inference efficiency calculator runs entirely client-side in your browser—no server communication occurs when you input values. Your hardware specs, costs, and generation speeds remain on your device. We don’t log, store, or transmit any of those parameters. The only data transmitted is your session behavior through standard analytics (page views, time spent), which is anonymized and compliant with GDPR. If you want complete offline usage, you can download the calculator source from our GitHub repository (github.com/ai-review-com/batch-calc) and run it locally. For infrastructure-sensitive organizations, we also provide a self-contained Docker image that runs the full tool without any external dependencies. Regarding local-only alternatives: if you need on-prem analysis tools, vLLM’s built-in profiler and NVIDIA’s TensorRT profiler both provide similar batch efficiency modeling, though they require more manual calculation to compare unit costs across different hardware configurations.
Running the numbers on this against actual commercial pricing. My team currently burns about $8K/month on Anthropic’s Batch API for document processing. At 65% scaling efficiency with our H100s running at $3.50/hour, the calculator shows we’d hit $0.38 per 1M tokens self-hosted versus $2.40 on Claude’s batch tier. Payback period is roughly 4 months if we account for ops overhead and on-call rotation. The real question: does anyone have production experience with KV-cache optimization actually hitting these efficiency percentages? We’re seeing closer to 55% on our LLaMA 70B setup with paged attention in vLLM. Also need to know if this factors in quantization losses—running GPTQ or AWQ would lower per-stream speed but change the math significantly. Would be useful if the calculator had a preset toggle for common quantization schemes.
Your 55% efficiency observation on LLaMA 70B with paged attention is valuable data—that aligns with what we’re seeing in production deployments at scale. You’re right that quantization significantly impacts this calculation. The 65% default assumes full precision (FP16), but GPTQ typically preserves 85-90% of throughput while cutting VRAM by 60%, and AWQ runs even tighter at similar performance. We’re working on a quantization preset system for the next release. For your specific case, if you quantize to 4-bit with AWQ on the H100, you’d likely push efficiency closer to 72% because you’re no longer memory-bandwidth bottlenecked on smaller batch sizes. That shifts your breakeven closer to 3 months. Regarding ops overhead: have you factored in token router redundancy and failover costs? Most teams underestimate the infrastructure tax of maintaining self-hosted inference at Anthropic’s reliability SLAs. The $3.50/hour H100 pricing assumes reserved instances; spot pricing could cut that to $1.20-$1.80, but adds operational complexity. One nuance: batch API pricing scales differently if you’re using longer context windows. Anthropic charges based on input+output tokens, so 200k context documents hit differently than short prompts. Worth modeling both scenarios if your workload is mixed.
Thanks, this is exactly the detail I needed. We’ve been running spot instances at $1.40/hour average, so the ops tax is real—we’ve had three outages in eight months that cost us roughly $15K in service credits to customers. The quantization preset would be huge because our team spends two days per quarter just rerunning efficiency calculations manually when we swap between GPTQ and AWQ. Quick follow-up: does the calculator account for context window length affecting batch efficiency? Our documents are averaging 45K tokens, and I suspect that’s crushing our actual throughput compared to shorter sequences.
The current calculator assumes fixed-length sequences, which is a limitation we should address. Longer context windows (45K tokens) create a different bottleneck profile than shorter prompts because KV-cache memory pressure grows quadratically with sequence length. At 45K context on a 80GB H100, you’re likely hitting memory bandwidth saturation much earlier than the 65% efficiency baseline assumes—probably closer to 48-52% when batching. This is especially true if you’re processing through an attention mechanism without recent optimizations like FlashAttention-2 or paged attention’s defragmentation. For document processing specifically, the efficiency cliff happens around batch size 4-6 on longer contexts, whereas shorter sequences can scale to 16-24 before hitting the same wall. We’re adding a context-length slider to account for this in the next version. In the interim, I’d recommend running a quick profiling test on your actual hardware: process five documents at increasing batch sizes with torch.profiler to measure actual memory bandwidth utilization. That’ll give you empirical efficiency percentages for your specific setup rather than relying on defaults. Post your findings in the GitHub issues if you’re comfortable—real-world 45K context data is exactly what the community needs to improve these models.