LLM Token Cost Calculator with Caching – cut your API bill

LLM Token Cost Calculator with Caching – cut your API bill Calculators

Prompt caching changes the math on API spend, and this calculator shows the change in dollars instead of hand-waving. It takes the shape of your traffic (how many requests, how many input and output tokens each one carries, and how much of that input repeats across calls) and prices it two ways: once at the plain per-token rate, and once with a cached prefix billed at write and read multipliers. The gap between those two numbers is your monthly saving, or in a few cases your monthly loss. It is built for solo developers and small teams who are watching a metered bill climb and want to know whether turning caching on is worth the effort.

Loading calculator...

Most token calculators stop at input times price plus output times price. That misses the part of a modern bill that actually moves: the long, stable chunk at the front of every prompt (a system prompt, a retrieved document, a few-shot block) that you resend on every single call. Cache it once and later calls read it at a fraction of the base rate. The tool splits your input into the repeating part and the changing part, applies the right rate to each, and reports what you keep.

You do not need provider docs open to use it. Type in your traffic, set the two multipliers to match whatever caching scheme your provider uses, and read the result. The defaults describe a mid-tier model with a large reusable prefix, which is the setup where caching pays off most.

How to use the LLM Token Cost Calculator with Caching

Start with volume. Put your real monthly request count in the first field. If you only know a daily figure, multiply by 30 and round; the tool scales linearly, so a rough number still gives a usable estimate you can refine later.

Next, describe one typical request. Enter the total input tokens it sends, then the slice of that input that stays identical from call to call in the cacheable field. A chatbot with a 1,800-token system prompt and a 1,200-token user turn has 3,000 input tokens, of which roughly 1,800 are cacheable. The changing part (the user message, the fresh retrieval) is never cached and always bills at the base input rate.

Set the output tokens per request from your logs or a sample of recent completions. Output is where a lot of the bill hides, because output rates run several times higher than input rates and caching does nothing for output at all.

If you are unsure how much of your prompt is reusable, look for the text that is byte-for-byte identical across two consecutive requests. That identical block is your cacheable prefix, and its length is the single number that decides whether caching helps you.

Now set the cache hit rate. This is the share of requests that arrive while a cached copy of the prefix is still live. High-traffic endpoints with a short, stable system prompt hit often; a quiet tool that fires once an hour may let every cache entry expire before the next call and hit almost never.

Fill in the input and output prices for your model in dollars per million tokens, copied from your provider’s pricing page. Finally, set the two multipliers to match the caching scheme: a cache write multiplier above 1 means writing to cache costs a premium over normal input, and a cache read multiplier below 1 is the discount you get on later reads. The defaults of 1.25 and 0.10 match one common scheme, but they vary by provider, so change them if yours differs.

Read the results top to bottom. The hero number is your monthly cost with caching, the line beneath it is the same traffic priced with no caching, and the difference tells you whether the feature earns its keep at your volume.

Calculator fields explained

Monthly requests – The number of API calls you make in a month. Default is 100,000. The whole bill scales directly with this, so a 10x traffic change is a 10x cost change.

Input tokens per request – Total tokens sent to the model in one call, including the system prompt, any retrieved context, and the user message. Default is 3,000. This is the full input before any split into cached and uncached parts.

Cacheable input tokens per request – The portion of the input that repeats unchanged across calls and can therefore be cached, such as a fixed system prompt or a stable document. Default is 2,000. It must be less than or equal to the total input; the remainder is treated as always-uncached.

Output tokens per request – Tokens the model generates in one response. Default is 500. Caching never touches output, so this cost is identical in both the cached and uncached totals.

Cache hit rate (%) – The share of requests that find a live cache entry for the prefix and pay the read rate instead of the write rate. Default is 80. Requests that miss pay the write premium instead.

Input price (USD per million tokens) – Your model’s base rate for input tokens, in dollars per million. Default is 3.00. This rate applies in full to the uncached part of every request.

Output price (USD per million tokens) – Your model’s rate for generated tokens, in dollars per million. Default is 15.00. Output rates typically run three to five times the input rate.

Cache write multiplier – How much writing the prefix into cache costs relative to the base input rate. Default is 1.25, meaning a 25 percent premium the first time a prefix is cached or refreshed. Set it to 1.00 for schemes with no write premium.

Cache read multiplier – How much reading a cached prefix costs relative to the base input rate. Default is 0.10, a 90 percent discount. Some providers charge 0.50 instead, a 50 percent discount, which shifts the whole result.

Understanding the results

ResultWhat it meansHow to act on it
Monthly cost with cachingYour full bill for the month with the cached prefix priced at write and read ratesCompare against your provider’s current invoice to sanity-check the model
Monthly cost without cachingThe same traffic with every input token at the base rateThis is your baseline if you do nothing
Monthly savings (USD and %)The difference between the two totals, positive or negativeA negative figure means caching costs you money at this hit rate
Cost breakdownUncached input, cache writes, cache reads, and output as separate linesFind which line dominates and attack that one first
Cost per requestThe cached total divided by request countMultiply by projected growth to forecast next quarter
Annual projectionThe monthly cached total times twelveUse for budget planning, not for locking in a number

The hero number, monthly cost with caching, is the one you take to a budget meeting. It already blends the cheap reads and the pricier writes according to your hit rate, so you are not looking at a best case or a worst case but at the weighted average of your actual traffic.

The savings line is where the decision lives. On the defaults it lands around 402 dollars a month against a 1,650 dollar baseline, a saving near 24 percent. That figure grows fast when the cacheable prefix is large and the hit rate is high, and it shrinks toward zero, or goes negative, when the prefix is small or the hit rate is low.

The breakdown is the most useful output for tuning. If the output line dwarfs everything else, caching will barely move your total no matter how you set it, because caching only touches input. In that situation your lever is a shorter completion or a cheaper model, not the cache.

A cache write costs more than an ordinary input token, so a prefix that almost never gets reused is worse than no caching at all. Watch the savings line for a negative value; it is the signal that your hit rate has fallen below the break-even point for your multipliers.

Caching only pays off when the cache read discount outweighs the write premium. Below a certain hit rate the writes pile up faster than the reads save, and the tool will show it plainly by printing a loss instead of a gain.

Very high inputs with a huge stable prefix produce dramatic savings, sometimes over half the bill, which is why long-context RAG and agent setups benefit most. Very low inputs, or inputs where almost nothing repeats, produce a rounding-error saving that is not worth the engineering time to wire up. The cost-per-request line helps here: if it barely changes between the two scenarios, leave caching off.

Calculation formulas

The tool splits each request into three billable input pieces plus output. Let I be total input tokens, C the cacheable tokens, and U the uncached remainder, so U equals I minus C. Let O be output tokens, h the hit rate as a fraction, and the prices be per million tokens. The two multipliers are mw for writes and mr for reads.

Cost per request with caching, in dollars, is:

(U × Pin + h × C × Pin × mr + (1 - h) × C × Pin × mw + O × Pout) ÷ 1,000,000

The uncached slice U always pays the full input rate. The cacheable slice C pays the read rate on the fraction h of requests that hit, and the write rate on the fraction that miss. Output pays its own rate untouched by caching. Multiply the per-request cost by monthly requests R for the monthly total.

The no-caching baseline is simpler, since all input bills at the base rate:

monthly baseline = R × (I × Pin + O × Pout) ÷ 1,000,000

Walk the defaults through it. With R = 100,000, I = 3,000, C = 2,000, U = 1,000, O = 500, h = 0.80, Pin = 3.00, Pout = 15.00, mw = 1.25, mr = 0.10: the uncached input is 1,000 × 3.00 = 3,000; the cache reads are 0.80 × 2,000 × 3.00 × 0.10 = 480; the cache writes are 0.20 × 2,000 × 3.00 × 1.25 = 1,500; the output is 500 × 15.00 = 7,500. Summing gives 12,480, divided by a million is 0.01248 dollars per request, times 100,000 is 1,248 dollars a month. The baseline is 16,500 per million, or 1,650 a month, so caching saves 402 dollars.

The break-even hit rate follows from the multipliers alone. Caching helps the cacheable slice whenever h × mr + (1 – h) × mw is below 1, which rearranges to h greater than (mw – 1) ÷ (mw – mr). For the 1.25 and 0.10 defaults that threshold is about 21.7 percent.

The table below shows the break-even hit rate for several common caching schemes, computed from that same formula. Set your two multipliers to match your provider, then check whether your real hit rate clears the threshold in the right-hand column.

Caching schemeWrite multiplierRead multiplierBreak-even hit rate
Short-TTL premium write1.250.10~21.7%
Long-TTL premium write2.000.10~52.6%
Automatic, no write premium1.000.500% (always helps)
Deep read discount, flat write1.000.100% (always helps)

Practical examples

Each scenario below runs the full formula. Inputs are listed, the arithmetic is shown, and the result tells you what to do next.

Example 1, solo chatbot on a budget model. Inputs: 5,000 requests, 1,200 input tokens (800 cacheable), 300 output, 70 percent hit rate, prices 0.50 and 1.50, multipliers 1.25 and 0.10. Per request: 400 × 0.50 = 200 uncached, 0.70 × 800 × 0.50 × 0.10 = 28 reads, 0.30 × 800 × 0.50 × 1.25 = 150 writes, 300 × 1.50 = 450 output, total 828 per million or 0.000828 dollars. Monthly with caching is 4.14 dollars against a 5.25 baseline. You save 1.11 dollars a month. At this scale, caching is not worth wiring up.

Example 2, personal RAG assistant. Inputs: 20,000 requests, 5,000 input (4,000 cacheable), 600 output, 85 percent hit, prices 3.00 and 15.00. Per request works out to 15,270 per million, or 305.40 dollars a month, versus a 480 baseline. That is 174.60 saved, about 36 percent. The large cacheable document is doing the work here.

Example 3, growing SaaS feature. Inputs: 200,000 requests, 8,000 input (7,000 cacheable), 400 output, 90 percent hit, prices 3.00 and 15.00. Monthly cached total is 2,703 dollars against a 6,000 baseline, a saving of 3,297 dollars a month, roughly 55 percent. Over a year that is 39,564 dollars kept in the business.

The pattern across these three is consistent: caching earns the most when a large block of the prompt repeats and traffic is dense enough to keep the cache warm. Small prefixes and thin traffic get you almost nothing.

Example 4, frontier model with a cold cache. Inputs: 50,000 requests, 4,000 input (3,000 cacheable), 800 output, 30 percent hit, prices 15.00 and 75.00. The write premium on 70 percent of calls nearly cancels the read discount on the other 30. Monthly cached is 5,786.25 dollars against a 6,000 baseline, saving only 213.75, about 3.6 percent. High prices did not help because the hit rate was too low.

Example 5, caching that backfires. Inputs: 30,000 requests, 3,000 input (2,500 cacheable), 400 output, 15 percent hit, prices 3.00 and 15.00. Per request: 1,500 uncached, 112.50 reads, 7,968.75 writes, 6,000 output, total 15,581.25 per million or 0.01558 dollars. Monthly with caching is 467.44 dollars, but the baseline is only 450. Caching here costs 17.44 dollars more per month, not less. The 15 percent hit rate sits below the 21.7 percent break-even, so every write premium is a pure loss.

Example 6, production agent with a fat context. Inputs: 500,000 requests, 12,000 input (10,000 cacheable), 1,000 output, 88 percent hit, prices 3.00 and 15.00. Monthly cached is 14,070 dollars against a 25,500 baseline, saving 11,430 a month or about 45 percent. Annualized, caching removes 137,160 dollars from the bill.

Example 7, high-volume cheap model. Inputs: 2,000,000 requests, 2,000 input (1,500 cacheable), 200 output, 75 percent hit, prices 0.15 and 0.60. Monthly cached is 564.38 dollars against 840, saving 275.62, near 33 percent. Even at fractions of a cent per call, two million calls make the saving real.

Example 8, provider with no write premium. Same idea, different scheme: 100,000 requests, 6,000 input (5,000 cacheable), 500 output, 80 percent hit, prices 2.50 and 10.00, but multipliers 1.00 and 0.50. Monthly cached is 1,500 dollars against a 2,000 baseline, a flat 25 percent saving. With no write penalty, every hit rate above zero helps, so the only question is how deep the read discount runs.

Tips and best practices

Put your longest stable content at the very front of the prompt. Caching works on prefixes, so a fixed system prompt followed by a fixed document followed by the changing user turn caches cleanly, while the same pieces shuffled around cache poorly or not at all. Order is a free optimization.

Measure your real hit rate before trusting a projection. Teams routinely assume 90 percent and discover 40 percent once they log it, because their traffic is bursty and cache entries expire between bursts. Pull a day of logs, count how often two consecutive calls share a prefix inside the cache window, and use that number.

Keep the cacheable prefix genuinely identical. A single changed token near the front, a timestamp, a per-user greeting, a rotating instruction, breaks the prefix and forces a fresh write. Move anything variable to the end of the prompt.

Attack output cost separately. Caching leaves output untouched, so if your breakdown shows output as the largest line, the cache is the wrong tool. Trim verbose completions, cap max tokens, or drop to a cheaper model for the generation step.

A reliable pattern for small teams: cache the system prompt and any retrieved documents, keep the user message uncached at the end, and target a hit rate above 60 percent. That combination captures most of the available saving without fragile prompt engineering.

Recheck the numbers whenever your provider changes prices or caching terms, because both the base rates and the multipliers move, and a scheme that saved you money last quarter can flip. This tool makes the recheck a two-minute job.

For a long-TTL cache with a steep write premium, raise your hit-rate bar. A 2.00 write multiplier pushes break-even past 50 percent, so a long cache only pays when the same prefix is genuinely reused dozens of times before it expires.

Batch similar requests in time when you can. Firing ten related calls in a burst while one cache entry stays live turns nine potential writes into nine reads. Spreading the same ten calls across an hour can let each one miss.

Do not cache one-off prompts. A prefix used once and never again pays the write premium for zero reads, which is strictly worse than the base rate. Reserve caching for prompts you know repeat.

Common mistakes to avoid

Assuming caching always saves money

Caching is a bet that reads will outnumber expensive writes. When the hit rate is low, that bet loses. Example 5 above showed a 15 percent hit rate turning a 450 dollar baseline into a 467 dollar bill.

Enabling caching on an endpoint with a sub-20-percent hit rate and a 1.25 write multiplier actively raises your bill. Check the break-even threshold for your multipliers before flipping the switch, not after the invoice arrives.

Turning caching on blind can add 5 to 10 percent to a low-hit-rate bill. The fix is to measure the hit rate first and compare it against the break-even column in the reference table.

Counting the whole prompt as cacheable

The cacheable field is only the part that repeats byte-for-byte. People enter their full input token count and see an inflated saving that never materializes.

The fix is to diff two real consecutive prompts and cache only the identical leading run. Everything after the first difference is uncached input at the base rate.

Ignoring output cost entirely

Because caching is exciting, teams pour effort into it while output quietly runs three to five times the input rate. In the defaults, output is 7,500 of the 12,480 per-request total, more than half.

The fix is to read the breakdown and size your effort to the biggest line. If output dominates, cap completion length before touching the cache.

Forgetting the cache window expires

A cache entry has a lifetime, and once it lapses the next call pays a write again. Low-frequency endpoints can miss on nearly every call even though the prompt is stable.

Do not model a quiet, once-an-hour endpoint at a 90 percent hit rate just because the prefix never changes. If calls arrive slower than the cache expires, your real hit rate can be close to zero regardless of how stable the text is.

The fix is to align your traffic pattern with the cache lifetime, or to accept a low hit rate in the model and see whether caching still clears break-even.

Putting variable text before stable text

A per-user name or a current date at the top of the prompt invalidates the cache for the entire block behind it. The stable document you meant to cache never gets a hit.

The fix is strict ordering: fixed content first, variable content last. This one change often turns a 20 percent hit rate into an 80 percent one.

Trusting a projection you never measured

Every input here is an estimate until you feed in logged numbers. A guessed hit rate and a guessed cacheable length can be off by a factor that flips the decision.

The fix is to replace each default with a measured value from a real sample before you commit engineering time or budget to caching.

When to use this calculator

Open it when a metered API bill is growing and a big part of every prompt repeats. That describes most RAG systems, most agents with a fixed toolset, and most chatbots with a substantial system prompt. In those cases the cacheable slice is large, the traffic keeps entries warm, and the saving is worth measuring precisely rather than guessing.

Reach for it also when you are choosing between providers whose caching schemes differ. One vendor’s 0.50 read discount with no write premium can beat another’s 0.10 discount with a 1.25 premium, or lose to it, depending entirely on your hit rate. Running both multiplier sets through the tool turns a marketing comparison into a dollar comparison.

The number only changes your decision when the cacheable prefix is a real fraction of the prompt and the traffic is dense enough to keep the cache alive. When either condition fails, the calculator will tell you to move on.

It is less useful for tools where almost nothing repeats, such as a translator that gets a fresh document every call with no shared system prompt. There the cacheable slice is near zero and the result barely differs from a plain token cost. It also adds little when output cost swamps input, since caching cannot touch output.

Use it before you build, not after. A five-minute estimate can tell you the saving is 3 dollars a month and not worth a sprint, or 3,000 dollars a month and worth doing first.

  • Multi-Model API Cost Comparator
  • Monthly AI API Budget Calculator
  • Prompt Length vs Cost Calculator
  • Context Caching Cost Savings
  • RAG Cost Calculator
  • System Prompt Overhead Calculator
  • Context Window Cost Optimizer

Glossary

Token – The unit models bill on, roughly three-quarters of an English word on average. All pricing here is per million tokens.

Input tokens – Tokens sent to the model, covering the system prompt, context, and user message.

Output tokens – Tokens the model generates in its response, billed at a separate and usually higher rate.

Prompt caching – Storing a repeated prefix so later requests read it cheaply instead of paying the full input rate each time.

Cacheable prefix – The leading run of a prompt that stays identical across calls and can therefore be cached.

Cache write – The act of storing a prefix in cache, billed at the write multiplier, often above the base input rate.

Cache read – Retrieving a stored prefix on a later call, billed at the read multiplier, usually well below the base rate.

Cache hit rate – The fraction of requests that find a live cached prefix and pay the read rate.

Is a cache hit rate a fixed property of a model? No. It depends entirely on your traffic and prompt design, not on the model. The same model can show a 90 percent hit rate on a busy endpoint and near zero on a quiet one.

Uncached input – The changing part of the input, always billed at the base rate no matter what caching you enable.

Cost per million tokens – The pricing unit for both input and output rates, quoted in dollars.

Cache multiplier – A factor applied to the base input rate to get the write or read price, such as 1.25 for writes or 0.10 for reads.

Break-even hit rate – The hit rate at which caching neither saves nor costs money, equal to (write multiplier minus 1) divided by (write multiplier minus read multiplier).

TTL – Time to live, the lifetime of a cache entry before it expires and the next call must write again.

Blended cost – The weighted average of read-priced hits and write-priced misses that produces your real per-request cost.

Annual projection – The monthly cached total multiplied by twelve, useful for budgeting but not a guarantee.

Frequently asked questions

Does caching reduce output cost?

No. Caching applies only to input tokens, specifically the repeating prefix. Output tokens bill at the same rate whether or not caching is on.

This matters when output dominates your bill. In the default scenario, output is 7,500 of the 12,480 per-request cost, so even perfect input caching leaves more than half the bill untouched.

What hit rate do I need for caching to pay off?

It depends on your two multipliers. With a 1.25 write premium and a 0.10 read rate, break-even is about 21.7 percent, so anything above that saves money and anything below it loses money.

Schemes with no write premium, such as a 1.00 write and 0.50 read, have a break-even of zero, meaning any hit rate above nothing helps. Always compute the threshold for your own numbers before deciding.

How do I estimate my cacheable token count?

Take two real prompts from consecutive requests and compare them from the start. The identical leading run, before the first character that differs, is your cacheable prefix.

For a typical assistant, that run is the system prompt plus any fixed retrieved context, often 60 to 80 percent of the input. Put that token count in the cacheable field and the changing remainder becomes uncached input.

Why did the tool show a negative saving?

Your hit rate is below the break-even point for your multipliers, so the write premiums on cache misses cost more than the read discounts save.

A negative saving is a valid and useful result. It is the calculator telling you to leave caching off for this endpoint until the hit rate rises.

Raise the hit rate by batching requests or shortening the cache-eligible content, or simply disable caching where traffic is too sparse to keep entries warm.

Can I model different providers with this?

Yes, by setting the two multipliers and the two prices to match each provider. That is exactly how you compare a 0.10-read scheme against a 0.50-read scheme on your own traffic.

Run your real request shape through each provider’s numbers and read the two cached totals side by side. The cheaper base rate does not always win once caching is applied.

Does the order of my prompt matter?

It matters a great deal. Caching keys on the prefix, so stable content must come first and variable content last, or the cache breaks on every call.

Moving a per-user greeting from the top of the prompt to the bottom can lift a hit rate from 20 percent to 80 percent with no other change, because the long document behind it now stays cacheable.

How accurate is the annual projection?

It is the monthly cached total times twelve, so it is only as accurate as your inputs and it assumes steady traffic. Real usage grows, prices change, and caching terms shift.

Treat it as a planning figure with a comfortable margin, not a committed number. Recompute it whenever your volume or your provider’s pricing moves.

Should I cache if my prompts are all unique?

No. If nothing repeats across calls, the cacheable slice is zero, there are no reads to earn the discount, and any write is pure overhead.

In that case the cached and uncached totals are identical, and the calculator will show a zero saving. Spend your effort on output length or model choice instead.

What if only part of my traffic can use caching?

Model each traffic type separately. Run the cacheable, high-repeat endpoints with a real hit rate, and run the one-off endpoints with a zero cacheable count, then add the two monthly totals.

This split gives a truer picture than averaging everything together, because a single blended hit rate can hide an endpoint that is quietly losing money on write premiums.

Disclaimer

This calculator produces educational estimates for planning purposes. The numbers it returns depend on the model you choose, your provider, your hardware and configuration, and the pricing in effect at the moment you run it, all of which change over time and without notice.

Cost outputs are not financial or business advice. They are arithmetic on the inputs you supply, and they carry all the uncertainty of those inputs, especially the hit rate and cacheable token count, which are estimates until you measure them from real logs.

Verify every rate and multiplier against your provider’s current pricing page, and check the results against your own measured usage before you rely on them for a budget or a provider decision. A caching scheme that saves money at one hit rate can cost money at another.

Run a small real test before scaling spend. Enable caching on a slice of traffic, measure the actual bill for a few days, and confirm the saving is real at your volume. The live calculator on this page is the source of truth for its exact fields, so use it directly for your own figures.

Rate article
Ai review
Add a comment

  1. Mia_Allen

    This calculator feels like another attempt to make expensive API usage look palatable. You can tweak the multipliers all you want, but at the end of the day, you are still shipping massive context windows to a black box. Where are the local benchmarks? I have yet to see a real-world test that proves this is not just a glorified spreadsheet for justifying recurring cloud costs. My own tests on a 32GB MacBook Pro show that running quantized models locally often beats the latency of these cached API calls anyway. It looks like a distraction from the fundamental unreliability of these models.

    Reply
    1. AI Review Team

      Regarding the focus on local versus cloud compute, the trade-off is often about deployment scale rather than raw performance. While running quantized models like Llama-3-8B locally offers zero-latency inference on your hardware, the API approach is designed for production applications where high-concurrency throughput exceeds the capacity of consumer-grade VRAM. The caching mechanism described here effectively reduces the redundant processing of static system instructions, which historically accounts for 60 percent of input token charges in RAG pipelines. Benchmarks from various providers suggest that cache hits reduce time-to-first-token by roughly 40 percent in high-traffic scenarios. For developers managing thousands of requests, this is an optimization for operational overhead, not a substitute for local model experimentation.

      Reply
    2. Mia_Allen

      Thanks for the context, though I still worry about the lock-in. If I optimize my prompts for a specific provider’s caching schema, moving to a different provider later becomes a massive technical debt headache.

      Reply
    3. AI Review Team

      That concern is valid and hits on a major friction point in current LLM development. Caching implementations currently lack a unified standard, leading to vendor-specific prefix structures. Using abstraction layers like LangChain or custom middleware to decouple prompt templates from the specific provider’s API implementation can mitigate some of this lock-in risk. By versioning your prompt sets independently, you ensure that migration only requires an update to the serialization logic rather than a total refactor of the caching strategy.

      Reply