Inference Cost Comparison: API vs self-hosting at your volume

Inference Cost Comparison: API vs self-hosting at your volume Calculators

Running a model behind a managed API and running it on a GPU you rent by the hour produce two very different bills, and the crossover between them moves as your request volume, token sizes, and prices change. This calculator takes one workload description and prices it both ways, so you can see which path is cheaper at your actual numbers instead of guessing from a blog post that assumed someone else’s traffic. It is built for solo developers, indie founders, and small teams deciding whether to keep paying per token or reserve a GPU.

Loading calculator...

The managed API path charges you for every token you send and receive, so its monthly cost scales in a straight line with volume. The self-hosted path charges you for GPU time whether the machine is busy or idle, so it behaves like a fixed cost that only steps up when you saturate a card and need a second one. Low traffic favours the API, heavy steady traffic favours self-hosting, and the interesting question is where your workload sits relative to that line.

You do not need to commit to either answer to use the tool. Enter what you already know, read the two monthly figures side by side, and note the break-even volume the calculator reports. That single number tells you how much your traffic would have to grow before the decision flips.

How to use the Inference Cost Comparison

Start with your monthly request count. If you are pre-launch and have no traffic yet, estimate from a comparable feature or use a round planning figure like 50,000 requests and revisit once real usage appears. Every other field scales against this number, so it anchors the whole comparison.

Next, fill in the two token fields: average input tokens per request and average output tokens per request. Input tokens are everything you send, including the system prompt, retrieved context, and the user message. Output tokens are what the model generates back. These two numbers usually differ by a lot, and because output tokens are priced higher on almost every provider, the output figure often drives more of your bill than the input figure does.

The API price fields ask for the cost per million tokens, separately for input and output, in USD. Copy these straight from your provider’s pricing page for the specific model you plan to call. A Budget tier model and a Frontier tier model can differ by 30 times on the same workload, so the model you pick here matters more than any other single entry.

Price the model you will actually ship, not the cheapest one on the page. Teams routinely prototype on a small model, quote its price internally, then deploy a larger one and get a bill three times what they planned for.

The self-hosted side needs three inputs. GPU hourly cost is what your cloud charges for the card you would rent, per hour. Throughput is how many of your requests one such GPU can finish per hour at your token sizes, which you get from a short load test or the model card. GPU utilization is the share of paid time the card spends doing useful work, since a reserved GPU bills 24 hours a day even when your traffic is quiet at 3 a.m.

Read the results as a pair. The calculator shows the API monthly cost, the self-hosted monthly cost, which one is lower, and the monthly saving between them. It also reports the break-even request volume, which is the traffic level where the two paths cost the same. Below it the API wins, above it a single reserved GPU wins, until you outgrow one card.

Change one input at a time to build intuition. Double the output tokens and watch the API cost jump while the self-hosted cost holds steady. Drop utilization from 60 percent to 25 percent and watch the self-hosted cost climb because the card now serves fewer of your requests per paid hour.

Calculator fields explained

Monthly requests – the number of inference calls you expect in a month. Default 100,000. Unit is requests per month. This is the volume every cost figure scales against.

Average input tokens per request – the mean size of everything you send the model, prompt plus context plus user text. Default 1,500 tokens. Larger contexts and retrieval-augmented prompts push this up quickly.

Average output tokens per request – the mean length of the model’s reply. Default 500 tokens. Chat and summarisation stay low here, while long-form generation and agent chains run high.

API input price (USD per million tokens) – what your provider charges per million input tokens for the chosen model. Default 3.00. Read it from the provider’s current pricing, since these change without much notice.

API output price (USD per million tokens) – the per-million price for generated tokens. Default 15.00. Output almost always costs more than input, often three to five times more.

GPU hourly cost (USD/hour) – the rental rate for the GPU you would self-host on. Default 2.50. A mid-range card sits near this figure, while top-end cards run higher.

Throughput (requests/hour) – how many of your requests one GPU completes per hour at your token sizes. Default 1,800. Bigger prompts and longer outputs lower this number.

GPU utilization (%) – the fraction of paid GPU time spent serving requests rather than sitting idle. Default 60. A reserved card bills continuously, so idle time is money you paid for and did not use.

Understanding the results

ResultWhat it meansHow to act on it
API monthly costTotal per-token spend for the volume you enteredCompare directly against the self-hosted figure below it
Self-hosted monthly costFixed GPU rental for enough cards to serve the volumeCheck how many GPUs it assumed before trusting it
Cheaper option and savingWhich path costs less and by how much per monthWeigh the saving against the engineering effort to switch
Cost per 1,000 requestsUnit price on each path, easier to reason aboutUse it to price your own product features
Cost per requestThe smallest unit, for margin mathSet your per-call price above this to stay profitable
GPU-hours per monthReserved hours the self-hosted path pays forConfirm it matches the GPUs your traffic actually needs
Break-even request volumeWhere API and one GPU cost the sameJudge how far your growth is from flipping the decision

The hero comparison is the two monthly figures. If the API number is lower, your traffic has not yet grown into the fixed cost of a reserved GPU, and paying per token is the efficient choice. If the self-hosted number is lower, you are sending enough steady volume through one card to beat per-token pricing.

The break-even volume is the number to remember. It divides the fixed monthly cost of one GPU by your API cost per request. At the default inputs a single GPU costs about 1,825 dollars a month and each request costs 1.2 cents on the API, so the two paths meet near 152,000 requests. Send fewer than that and the API wins; send more and one reserved card wins.

Output tokens usually move your bill more than input tokens do. Because output is priced several times higher, a workload that generates long replies can cost more on the API even with a modest input prompt, which pulls the break-even volume down and makes self-hosting attractive sooner.

A self-hosted figure that looks cheap at 100 percent utilization is fiction. Real serving traffic is bursty, so plan around 50 to 70 percent and treat anything higher as a stretch that needs autoscaling or batching to reach.

Watch the GPU count baked into the self-hosted number. Once your volume exceeds one card’s capacity, the calculator adds a whole GPU, and the self-hosted cost jumps in a step rather than a smooth curve. A workload just above a card’s ceiling pays for two cards while using barely more than one, which is the worst place to sit.

Very low volume makes the API look almost free per month while the self-hosted path still bills a full GPU. At 10,000 requests on a cheap model you might spend under 10 dollars on the API against 800 or more for a reserved card. Nobody should self-host at that scale.

Very high volume inverts this completely. At three million requests the per-token API bill can run tens of thousands of dollars a month while four reserved GPUs serve the same load for a fraction of it. The gap at scale is the strongest argument for owning your inference.

Calculation formulas

The API path is a straight per-token calculation. Cost per request comes from the two token counts and their prices:

api_cost_per_request = (input_tokens / 1,000,000 × input_price) + (output_tokens / 1,000,000 × output_price)

api_monthly_cost = monthly_requests × api_cost_per_request

The self-hosted path treats a GPU as a fixed monthly cost running around the clock. A month is taken as 730 hours (365 days divided by 12, times 24):

gpu_monthly_fixed = gpu_hourly_cost × 730

capacity_per_gpu = throughput × 730 × (utilization / 100)

gpus_needed = ceil(monthly_requests / capacity_per_gpu)

selfhosted_monthly_cost = gpus_needed × gpu_monthly_fixed

The break-even volume compares the API against a single reserved GPU:

break_even_requests = gpu_monthly_fixed / api_cost_per_request

Work through the defaults step by step. Input cost per request is 1,500 divided by a million times 3.00, which is 0.0045 dollars. Output cost is 500 divided by a million times 15.00, which is 0.0075 dollars. Together the API costs 0.012 dollars per request, so 100,000 requests come to 1,200 dollars a month. The GPU fixed cost is 2.50 times 730, or 1,825 dollars, and one card handles 1,800 times 730 times 0.60, which is 788,400 requests, comfortably above the volume, so a single GPU costs 1,825 dollars. The API wins here, and break-even sits at 1,825 divided by 0.012, about 152,083 requests.

The 730-hour month is a convention. If you rent GPUs by spot or preemptible pricing and shut them down off-peak, your real self-hosted cost drops below this fixed figure, which is one lever for making self-hosting cheaper without changing the model.

The reference tables below give realistic starting values for the price and hardware fields.

Model tierInput (USD / 1M tokens)Output (USD / 1M tokens)Typical use
Budget0.10 to 0.600.30 to 1.80Classification, extraction, high volume
Standard0.50 to 3.001.50 to 6.00General chat, summarisation
Mid3.00 to 8.0010.00 to 24.00Reasoning, longer generation
Frontier10.00 to 20.0040.00 to 80.00Hard reasoning, agents, code
GPU classTypical hourly rent (USD)Rough throughput (req/hr)Best for
Entry (24 GB)0.50 to 1.202,000 to 4,000Small quantized models
Mid (40 to 48 GB)1.50 to 3.001,200 to 2,5007B to 13B models
High (80 GB)3.00 to 5.00800 to 1,80030B to 70B models
Multi-GPU node8.00 to 30.00variesLarge or high-concurrency serving

Practical examples

Each scenario states the inputs, runs the formulas, and ends with the decision the numbers point to.

Solo side project, tiny volume

Inputs: 10,000 requests, 800 input, 300 output, 0.50 input price, 1.50 output price, 1.20 GPU/hr, 2,000 req/hr, 50 percent utilization. API per request is 0.0004 plus 0.00045, so 0.00085 dollars. API monthly is 8.50 dollars. One GPU costs 876 dollars fixed. The API wins by a mile, and break-even is over a million requests. Use the API and forget self-hosting exists at this scale.

Growing indie app

Inputs: 50,000 requests, 1,200 input, 400 output, 3.00 and 15.00 prices, 2.50 GPU/hr, 1,800 req/hr, 60 percent. API per request is 0.0096 dollars, so 480 dollars a month against 1,825 for a reserved card. The API is still cheaper. Break-even sits near 190,000 requests, so this app has room to roughly quadruple before the question changes.

Right at the crossover

Inputs: 200,000 requests at the default token sizes and prices. API per request is 0.012 dollars, giving 2,400 dollars a month. One GPU costs 1,825 dollars and still has spare capacity. Self-hosting now wins by 575 dollars a month. This is the first scenario where reserving a card pays off.

The crossover is not a cliff, it is a slope. A few thousand requests either side of break-even barely changes the decision, so do not over-engineer a migration to save 50 dollars a month. Wait until the gap is clearly worth the effort.

Production scale, multiple GPUs

Inputs: 3,000,000 requests at default sizes and prices, 2.50 GPU/hr, 1,800 req/hr, 60 percent. API monthly is 36,000 dollars. One card serves 788,400 requests, so you need four cards, costing 7,300 dollars. Self-hosting saves about 28,700 dollars every month here. At this volume, owning inference is the obvious call.

Output-heavy generation

Inputs: 100,000 requests, 500 input, 2,000 output, 3.00 and 15.00 prices, 2.50 GPU/hr, 1,000 req/hr, 60 percent. The long outputs push API per request to 0.0315 dollars, so 3,150 dollars a month. One GPU at 1,825 dollars wins by 1,325. Heavy output generation flips the decision toward self-hosting far earlier than chat workloads do.

Frontier model, expensive tokens

Inputs: 80,000 requests, 2,000 input, 800 output, 15.00 and 75.00 prices, 4.00 GPU/hr, 900 req/hr, 55 percent. API per request is 0.09 dollars, so 7,200 dollars a month, while a single high-end card costs 2,920 dollars. Self-hosting wins by 4,280 dollars, and break-even is only about 32,000 requests. Expensive per-token models make owning the hardware attractive at modest volume.

Low utilization penalty

Inputs: 400,000 requests at default sizes and prices, 2.50 GPU/hr, 1,800 req/hr, but utilization at 25 percent. Each card now serves only 328,500 requests, so you need two GPUs at 3,650 dollars against 4,800 for the API. Self-hosting still wins, but poor utilization doubled the GPU bill. At 60 percent one card would have covered it for 1,825 dollars.

Cheap model at high volume

Inputs: 600,000 requests, 1,000 input, 300 output, 0.50 and 1.50 prices, 1.20 GPU/hr, 2,500 req/hr, 80 percent. API per request is only 0.00095 dollars, so 570 dollars a month, while one GPU costs 876. The API wins despite the high volume, because cheap tokens keep per-request cost tiny. Break-even sits above 920,000 requests. A cheap model can stay competitive on the API well past the point where a pricier model would have flipped.

Tips and best practices

Measure throughput on your own hardware and token sizes before trusting the self-hosted number. A model card quotes throughput under ideal batching, and your real figure with production prompt lengths is often half that. Feeding an optimistic throughput into the calculator makes self-hosting look cheaper than it will be.

Separate the input and output token averages carefully. Because output is priced three to five times higher, mixing them into one average distorts the API cost. If most of your traffic is short prompts with long answers, the output field is where your money goes.

Run the comparison at three volumes: your current traffic, your six-month target, and a pessimistic double of that. The path that wins today may lose next quarter, and knowing the break-even in advance means you switch on schedule rather than after a surprise bill.

Treat utilization as the honesty knob on the self-hosted side. Bursty consumer traffic rarely clears 40 percent on a single reserved card without batching. Steady backend jobs and overnight batch runs can reach 80 percent or more, which is where self-hosting shines.

Batch your non-urgent inference. Queuing background jobs and running them at high GPU utilization can push a self-hosted card from 30 to 80 percent, cutting the effective cost per request by more than half without renting anything extra.

Keep the API path as your baseline even if you self-host. When traffic dips below break-even, or during an outage on your own cluster, per-token pricing is a clean fallback. Many small teams run a hybrid where steady load hits their GPU and spillover overflows to the API.

Re-run the numbers whenever a provider changes prices, and they change often. A 40 percent output price cut can move break-even by tens of thousands of requests, which might pull you back from a self-hosting migration you were about to start.

Include the hidden costs of self-hosting in your own mental model even though this tool prices only the GPU. Someone has to patch the server, handle a failed card at 2 a.m., and keep the serving stack current. For a solo developer that time has real value, and it often justifies staying on the API past the raw break-even.

Quote your product pricing from the cost-per-request figure, not the monthly total. If a call costs you 1.2 cents, charging a customer 1 cent per call loses money on every request no matter how you dress up the monthly plan.

Common mistakes to avoid

Assuming 100 percent GPU utilization

The single biggest error is entering a self-hosted cost that assumes the card is always busy. Real traffic ebbs, and a reserved GPU bills through every idle minute.

A comparison run at 100 percent utilization can understate the true self-hosted bill by two or three times once real, bursty traffic is served. Never present that number to anyone making a spend decision.

Use a utilization you have measured or can defend, usually somewhere between 50 and 70 percent for interactive workloads, and treat higher figures as goals that need batching to hit.

Pricing the wrong model

Teams prototype on a small, cheap model, record its per-token price, then ship a larger one and never update the numbers. Shipping a model three tiers up can triple your bill overnight. Always price the exact model that will handle production traffic.

Ignoring output token weight

Plugging a single blended token price into both fields hides where the cost lives. Output tokens usually cost several times more than input, so a workload with long generations is far more expensive than an input-heavy one at the same total token count. Enter the two prices separately, every time.

Forgetting the step function on GPUs

Self-hosted cost does not rise smoothly. The moment your volume crosses one card’s capacity, the calculator adds a whole second GPU, and the cost jumps.

Landing your traffic just above a single card’s ceiling is the worst outcome: you pay for two GPUs while using barely more than one, wasting close to a full card’s cost every month.

If you sit near a capacity boundary, try raising utilization through batching or trimming token sizes before committing to a second card.

Comparing at a single volume

A comparison run only at today’s traffic answers the wrong question. The decision that matters is where you will be in six months. Run the tool at several volumes so you know your break-even and can plan the switch instead of reacting to it.

Leaving out real-world serving overhead

The GPU rental is not the whole self-hosted cost. Load balancers, storage for model weights, monitoring, and the engineering hours to keep it all running are real. This tool prices the compute; you add the rest before deciding, especially on a small team where labour is the scarcest resource.

Trusting vendor throughput blindly

Published throughput numbers assume ideal batch sizes and short sequences. Your production throughput with long contexts can be half the quoted figure, which doubles the GPUs you actually need. Load test with your real prompt distribution before locking in the self-hosted estimate.

When to use this calculator

Reach for this comparison when you are about to sign up for a new provider or reserve a GPU and want the cost consequence before you commit. The moment a per-token bill starts feeling uncomfortable is exactly when the break-even volume becomes worth knowing.

Use it again whenever your traffic changes materially or a provider updates prices. A workload that clearly belonged on the API at launch can cross into self-hosting territory after a growth spurt, and the only way to catch that on time is to re-run the numbers on a schedule rather than waiting for the invoice to shock you.

It also earns its place during a product pricing exercise. The cost-per-request figure on whichever path you choose sets the floor under what you can charge, and pricing a feature without that floor is how teams end up losing money on their most popular endpoints.

The cheapest inference is the request you never send. Before comparing API against self-hosting, check whether caching, shorter prompts, or a smaller model removes the cost entirely.

Skip the tool when your volume is trivially small. If you are serving a few hundred requests a month on a cheap model, the API is obviously right and the comparison only confirms the obvious. Spend that time shipping features instead.

  • Multi-Model API Cost Comparator
  • Self-Hosting Break-Even Calculator
  • Monthly AI API Budget Calculator
  • GPU Type Comparison Calculator
  • Batch Inference Efficiency Calculator
  • Token Generation Speed Simulator
  • Peak Load AI Cost Estimator

Glossary

Inference – running a trained model to produce an output from an input, as opposed to training the model in the first place.

Input tokens – the tokens you send to the model, including system prompt, retrieved context, and the user message.

Output tokens – the tokens the model generates in its reply, usually priced higher than input tokens.

Cost per million tokens – the standard unit providers price by, quoted separately for input and output.

Managed API – a hosted endpoint where you pay per token and the provider runs the hardware.

Self-hosting – running the model on GPUs you rent or own, paying for compute time rather than tokens.

Does self-hosting always beat the API at scale? Not always. A cheap model with tiny per-token cost can stay competitive on the API well past a million requests, because the fixed GPU cost has less headroom to undercut.

Throughput – how many requests one GPU completes per hour at your token sizes.

GPU utilization – the share of paid GPU time actually spent serving requests rather than idling.

Break-even volume – the request count at which the API and a single reserved GPU cost the same.

Reserved GPU – a card rented continuously by the hour, billed whether busy or idle.

Blended cost – a single combined price that averages input and output, useful for quick math but misleading for output-heavy work.

Capacity – the number of requests one GPU can serve per month at a given throughput and utilization.

Unit economics – the cost and revenue of a single request or customer, the basis for judging whether a feature is profitable.

Frequently asked questions

Why does output cost more than input?

Generating tokens is more computationally expensive than reading them, because the model produces output one token at a time in sequence while it can process input in parallel.

On most providers output is priced three to five times higher than input. At the default 3.00 input and 15.00 output prices, a reply of 500 tokens costs more than a prompt of 1,500 tokens, which is why the output field usually dominates the bill.

What month length does the calculator use?

It uses 730 hours, which is 365 days divided by 12 months times 24 hours. This keeps the self-hosted cost consistent across months of different lengths.

If you shut GPUs down during off-peak hours using spot or preemptible instances, your real cost falls below this figure, since you are no longer paying for a full 730 hours per card.

Does the tool account for multiple GPUs?

Yes. It divides your monthly volume by one card’s capacity and rounds up, so a workload needing 3.8 cards is charged for 4.

This is why the self-hosted cost rises in steps. Crossing from 788,400 to 788,401 requests at the default throughput adds a whole second GPU and a full card’s cost, so watch for capacity boundaries.

How accurate is the throughput field?

It is only as accurate as the number you enter, and that number varies with your token sizes, batching, and quantization.

Published throughput assumes ideal conditions, so measure it yourself with production-length prompts. A real figure half the quoted one doubles the GPUs you need and can flip the decision back to the API.

Should a solo developer ever self-host?

Usually not at low volume, where a reserved GPU costs hundreds of dollars a month while the API costs a few dollars.

Self-hosting starts to make sense for a solo developer only with steady heavy traffic or an expensive per-token model, and even then the maintenance burden of running a server alone often tips the balance back toward the API.

What if my traffic is very bursty?

Bursty traffic hurts self-hosting because a reserved card sits idle between spikes while still billing every hour.

A workload that peaks hard for two hours a day and idles the rest often serves under 20 percent utilization on a reserved card, making the API the cheaper and simpler choice despite a high peak rate.

If your traffic is spiky, either stay on the API or add autoscaling and batching to lift utilization before committing to owned hardware.

Can I use both paths together?

Yes, and many small teams do. Steady baseline load runs on a reserved GPU while spikes overflow to the API.

This hybrid captures most of the self-hosting saving on predictable traffic while keeping the API as elastic overflow, so you never pay for a second idle card just to cover an occasional peak.

How often should I re-run this?

Re-run whenever traffic changes noticeably or a provider updates prices, which happens several times a year across the major vendors.

A single output price cut can shift break-even by tens of thousands of requests. Checking quarterly is enough for most small teams, with an extra run any time you are about to commit to a GPU reservation.

Does it include data transfer or storage costs?

No, it prices GPU compute on the self-hosted side and tokens on the API side only.

Model weight storage, egress bandwidth, and monitoring add real cost to self-hosting, so add those to the GPU figure before treating the comparison as final, especially for large models with weights in the tens of gigabytes.

Disclaimer

This calculator produces educational estimates for planning. The results depend on the model you choose, your provider, your hardware, your configuration, and current pricing, all of which change frequently and independently of this tool.

Cost outputs here are not financial or business advice. They are a starting point for your own analysis, not a guarantee of what you will actually spend once real traffic, real prompt sizes, and real utilization come into play.

Verify every price against your provider’s current published rates and measure your own usage before relying on any number this tool reports. Throughput and utilization in particular should come from your own load tests, not from defaults or vendor claims.

Run a small real test at low volume before scaling spend on either path. The live calculator on this page is the source of truth for its exact fields and behaviour, so use it directly rather than the worked examples above, which illustrate the method with rounded figures.

Rate article
Ai review
Add a comment

  1. IsabellaWalker

    I used this to estimate costs for a small RAG tool I built for my thesis research. I was leaning toward OpenAI because of their easy SDK, but the calculator showed my token usage for 1000 research papers would cost way more than renting a small GPU instance on RunPod. It actually made me realize I can host Llama 3 8B locally to process my PDFs without worrying about API costs or privacy leaks. The interface is clean, but how do I account for cold starts if I scale down to zero?

    Reply
    1. AI Review Team

      Regarding your concern about cold starts, that is a common hurdle when moving from serverless APIs to reserved GPU instances. If you opt for a service like RunPod or Lambda Labs, they offer serverless GPU endpoints that scale to zero. You pay a slight premium per inference compared to a dedicated box, but you avoid the 24/7 idle costs. Alternatively, if you use a dedicated instance, consider using a queueing system like Celery or BullMQ to batch requests. This increases utilization and prevents the GPU from sitting idle during lulls in your research workflow.

      Reply
    2. IsabellaWalker

      Thanks for the tip on serverless endpoints. I was looking at dedicated instances which felt wasteful for my irregular thesis schedule. I will check the documentation for cold start latency on those serverless options.

      Reply
    3. AI Review Team

      Glad that helps! Keep in mind that for RAG pipelines, the initial load time of your vector database and model weights into VRAM is the primary latency factor. If you see high startup times, you might explore keeping the model loaded in a warm pool for an hour after the last request to avoid the cold start penalty.

      Reply
  2. Charlie.Morris

    The obsession with managed APIs is exhausting. People keep acting like renting a 4090 or even a used 3090 is rocket science. If you have any steady traffic, you are just throwing money away by paying per token to a black box. I run local models using llama.cpp and Ollama on a Mac Studio for my dev work. Data sovereignty is the main reason. Why would anyone send proprietary code or sensitive research data to an external provider when you can achieve similar latency with a quantized model in GGUF format? I checked the calculator and it confirms my suspicion. The break-even point for a 70B model is surprisingly low if you already own the hardware. Stop feeding the API tax and just host your own weights. It is not that hard to set up a vLLM container or use an existing runner on a private cloud.

    Reply
    1. AI Review Team

      The trade-off between data sovereignty and infrastructure management is exactly why this calculator exists. For users who prioritize the privacy benefits of hosting local weights, the hidden cost is often the engineering time required to maintain the stack. While a 4090 handles inference well for small-to-medium parameter counts, scaling to multiple nodes once you hit 70B or 405B models introduces complexity with NCCL communication and multi-GPU orchestration. You are correct that the ROI on owned hardware is high, provided your team has the capacity to manage the deployment lifecycle.

      Reply