This calculator turns a handful of usage numbers into a monthly dollar figure for your LLM API spend. You give it how many requests you send per day, how many input and output tokens an average request uses, and the per-million-token prices your provider charges. It returns the monthly cost, the daily cost, what a single request costs, an annual projection, and a padded budget number you can actually put in a spreadsheet. It exists for the person who signed up for an API, ran a few calls, and now has no idea whether next month reads $12 or $1,200.
Loading calculator...
Most people building on top of language models discover their real cost the same way: an email from the billing system. Token pricing is small enough per call that it feels free during testing, then compounds once a feature ships to actual users. A chatbot answering 400 conversations a day looks harmless until you multiply four numbers together and see the yearly total.
The math here is not complicated, but it is easy to get wrong by an order of magnitude if you forget that output tokens usually cost three to five times what input tokens cost, or that a “small” 40,000-token context window gets billed on every single call. This tool keeps those pieces separate so the estimate holds up.
How to use the Monthly AI API Budget Calculator
Start with the request volume, because everything scales off it. Enter how many API calls you expect to send on a typical day. If you already have a running feature, pull the real number from your provider dashboard or your own logs. If you are still planning, estimate from your user count: a support bot might average two or three calls per active user per day, while a background summarization job could fire thousands of times without a human present.
Next, set the token counts. Input tokens cover everything you send: the system prompt, any retrieved context, conversation history, and the user’s message. Output tokens are what the model writes back. These two fields are separate on purpose, since they are priced differently and often differ by a factor of ten in size. A long RAG prompt might carry 6,000 input tokens and produce a 300-token answer.
Enter the two prices exactly as your provider lists them, per million tokens, in USD. Providers publish these on their pricing pages, and the input and output rates are always shown as two separate figures. Copy both. Guessing a single blended rate is the most common way these estimates drift.
Token counts do not need to be perfect to be useful. Getting them within 20 percent of reality already tells you whether you are looking at a $30 hobby project or a $3,000 line item that needs a budget conversation.
Set days per month to match your traffic pattern. A public consumer app runs seven days a week, so 30 is right. An internal tool used only on workdays runs closer to 22, and using 30 there would overstate your bill by more than a third.
The safety buffer adds a percentage on top of the raw estimate. Real usage is lumpy: some users write essays, some requests retry, and prompts tend to grow over time as you add instructions. A buffer of 15 to 25 percent gives you a budget number you are unlikely to blow past. Set it to zero if you want the bare arithmetic with no padding.
Once every field is filled, the results update immediately. Read the monthly cost first, then check the annual projection to see what the same behavior costs sustained across a year. That yearly number is usually what changes minds about whether to optimize prompts or switch models.
Calculator fields explained
Requests per day – The number of API calls you send in a typical 24-hour period. Default is 500. This is the multiplier that drives the whole estimate, so anchor it to logs if you have them rather than to a hopeful guess.
Input tokens per request – The average size of everything you send to the model per call, measured in tokens. Default is 1,500. Include your system prompt, retrieved documents, chat history, and the user message. As a rough guide, one token is about four characters of English, so 1,500 tokens is roughly 1,100 words.
Output tokens per request – The average length of the model’s reply, in tokens. Default is 500. A one-paragraph answer runs 100 to 200 tokens; a detailed multi-section response can exceed 1,000.
Input price (USD per million tokens) – What your provider charges per million input tokens. Default is 0.50. Cheap models sit near 0.15, mid-tier models around 3.00, and frontier models can reach 15.00 or higher.
Output price (USD per million tokens) – What your provider charges per million output tokens, entered separately because it is almost always higher than the input rate. Default is 1.50. Frontier output pricing can run 60.00 or beyond.
Days per month – How many days the traffic actually runs. Default is 30. Use 30 for a seven-day consumer product, around 22 for a weekday-only internal tool.
Safety buffer (%) – A percentage added on top of the raw estimate to absorb usage spikes and prompt growth. Default is 20. Set it to 0 for the unpadded figure.
Understanding the results
| Result | What it means | How to act on it |
|---|---|---|
| Estimated monthly cost | Raw spend for one month at the volume and prices you entered, before the buffer. | Compare against what you have budgeted for infrastructure this month. |
| Daily cost | Monthly cost divided by your days-per-month value. | Sanity-check it against a single day of real billing once you launch. |
| Cost per request | What one average API call costs, input plus output combined. | Multiply by any campaign or feature volume to price a specific launch. |
| Annual projection | Monthly cost times twelve, assuming steady usage. | Use this for yearly planning and for deciding if optimization is worth engineering time. |
| Buffered budget | Monthly cost with your safety buffer added. | Put this number in the actual budget line, not the raw estimate. |
The estimated monthly cost is the hero number, and it answers the direct question: at this behavior, what does one month cost. It is the sum of every request’s input and output charges across the whole month. If it surprises you, the surprise almost always lives in one of two places, output token size or request volume, and both are visible in the fields above.
Cost per request is the number that scales your thinking. Once you know a single call costs, say, $0.0018, you can price anything: a marketing push that triggers 50,000 calls costs $90, a free tier that lets each user make 100 calls a month costs $0.18 per user before you charge them a cent. This is the figure to keep in your head.
Output tokens usually drive more than half the bill despite being fewer in number. That is because output pricing runs several times the input rate. A request with 4,000 input tokens and 600 output tokens can still spend most of its money on those 600 tokens. When a result looks too high, check the output price and output length before anything else.
The annual projection assumes today’s usage stays flat for twelve months. Real products grow. If your user base is climbing 10 percent a month, the true yearly figure is well above twelve times the current month, so treat this projection as a floor, not a ceiling.
Reading edge cases matters. If you enter a very high request count with tiny token values, the per-request cost can round to something that looks like zero while the monthly total is real money. The opposite happens with a low-volume, huge-context workload: a research tool making 30 calls a day with 100,000-token prompts can outspend a chatbot handling 5,000 daily calls. Volume and size trade off against each other, and the calculator shows both effects at once.
The buffered budget is the number you commit to. If your raw estimate is $800 and your buffer is 20 percent, you plan for $960 and you are rarely wrong on the high side. Underestimating an API bill causes worse problems than overestimating it, so the padding earns its place.
Calculation formulas
The core formula separates input and output, then scales by volume and days:
monthly cost = requests_per_day × days_per_month × [(input_tokens ÷ 1,000,000 × input_price) + (output_tokens ÷ 1,000,000 × output_price)]
The derived outputs follow from that:
cost_per_request = (input_tokens ÷ 1,000,000 × input_price) + (output_tokens ÷ 1,000,000 × output_price)
daily_cost = cost_per_request × requests_per_day
annual_projection = monthly_cost × 12
buffered_budget = monthly_cost × (1 + buffer ÷ 100)
Walk through the defaults. With 1,500 input tokens at $0.50 per million, the input charge per request is 1,500 ÷ 1,000,000 × 0.50 = $0.00075. With 500 output tokens at $1.50 per million, the output charge is 500 ÷ 1,000,000 × 1.50 = $0.00075. Cost per request is $0.0015. At 500 requests a day, daily cost is $0.75. Across 30 days, monthly cost is $22.50. Annual projection is $270. With a 20 percent buffer, the budget line reads $27.00.
Notice that in the default example, input and output happen to cost the same per request despite output being one-third the token count. That is the 3x price ratio doing its work. Change the output price to $6.00 and the same request jumps to $0.00375, more than doubling the whole bill.
The reference table below shows typical per-million pricing tiers so you can slot your provider into the right range without hunting for a decimal.
| Tier | Input price (USD / 1M) | Output price (USD / 1M) | Typical use |
|---|---|---|---|
| Budget | 0.10 – 0.30 | 0.30 – 1.00 | Classification, routing, bulk tagging |
| Standard | 0.40 – 1.00 | 1.20 – 4.00 | Chatbots, summaries, everyday features |
| Mid | 2.50 – 4.00 | 10.00 – 15.00 | Reasoning, code, harder tasks |
| Frontier | 10.00 – 20.00 | 40.00 – 75.00 | Complex agents, high-stakes output |
Token-to-text conversion helps when you only know word counts. English runs about 0.75 words per token, so a 1,000-word document is roughly 1,333 tokens. Code and non-English text pack fewer words per token, so lean higher when either is involved.
Practical examples
Each example below states the scenario, the inputs, the arithmetic, the result, and what you would do with the number.
Example 1, weekend side project chatbot. Inputs: 200 requests/day, 800 input tokens, 250 output tokens, $0.15 input, $0.60 output, 30 days, 15% buffer. Per request: 800÷1M×0.15 + 250÷1M×0.60 = $0.00012 + $0.00015 = $0.00027. Monthly: 0.00027 × 200 × 30 = $1.62. Buffered: $1.86. Result: this costs less than a coffee, so ship it and stop worrying about the API line.
Example 2, indie SaaS support assistant. Inputs: 1,200 requests/day, 2,000 input tokens, 400 output, $0.50 input, $1.50 output, 30 days, 20% buffer. Per request: $0.001 + $0.0006 = $0.0016. Monthly: 0.0016 × 1,200 × 30 = $57.60. Buffered: $69.12. Result: comfortably inside a small tool’s margin, no optimization needed yet.
Example 3, internal weekday summarizer. Inputs: 3,000 requests/day, 4,000 input tokens, 600 output, $0.50 input, $1.50 output, 22 days, 10% buffer. Per request: $0.002 + $0.0009 = $0.0029. Monthly: 0.0029 × 3,000 × 22 = $191.40. Buffered: $210.54. Result: using 22 days instead of 30 saved you from budgeting $261, a real difference on an internal tool.
The same workload in Example 3, run seven days a week instead of weekday-only, costs $261 a month raw. The days-per-month field alone moved the estimate by $70. Match it to how your traffic actually behaves before trusting the total.
Example 4, RAG research tool. Inputs: 300 requests/day, 12,000 input tokens, 800 output, $0.50 input, $1.50 output, 30 days, 20% buffer. Per request: $0.006 + $0.0012 = $0.0072. Monthly: 0.0072 × 300 × 30 = $64.80. Buffered: $77.76. Result: only 300 calls a day, yet the heavy context pushes it above the 1,200-call support bot in Example 2.
Example 5, frontier-model agent. Inputs: 500 requests/day, 5,000 input tokens, 1,500 output, $15 input, $60 output, 30 days, 25% buffer. Per request: 5,000÷1M×15 + 1,500÷1M×60 = $0.075 + $0.09 = $0.165. Monthly: 0.165 × 500 × 30 = $2,475. Buffered: $3,093.75. The same call volume as Example 2 costs 43 times more on a frontier model. Result: this is the number that justifies routing easy calls to a cheaper model.
Example 6, high-volume classification job. Inputs: 50,000 requests/day, 300 input tokens, 20 output, $0.15 input, $0.60 output, 30 days, 15% buffer. Per request: $0.000045 + $0.000012 = $0.000057. Monthly: 0.000057 × 50,000 × 30 = $85.50. Buffered: $98.33. Result: per-request cost rounds toward nothing, but 1.5 million calls a month still adds up to real money.
Example 7, growing consumer app. Inputs: 8,000 requests/day, 1,800 input tokens, 450 output, $0.50 input, $1.50 output, 30 days, 20% buffer. Per request: $0.0009 + $0.000675 = $0.001575. Monthly: 0.001575 × 8,000 × 30 = $378. Annual: $4,536. Result: the yearly figure, not the monthly one, is what tells you a prompt-trimming project would pay for itself.
Example 8, output-heavy content generator. Inputs: 600 requests/day, 500 input tokens, 3,000 output, $0.50 input, $1.50 output, 30 days, 20% buffer. Per request: $0.00025 + $0.0045 = $0.00475. Monthly: 0.00475 × 600 × 30 = $85.50. Buffered: $102.60. Result: output tokens are 95 percent of the cost here, so shortening responses is the only lever that matters.
Tips and best practices
Pull your token counts from a real tokenizer rather than counting characters by hand. Every provider offers one, and the difference between your estimate and reality usually comes down to underguessing how many tokens a system prompt and retrieved context actually consume. Measure a dozen real requests and average them.
Separate your traffic into buckets before you estimate. A single app often mixes a cheap classification step, a mid-tier chat step, and an occasional expensive reasoning step. Running the calculator once per bucket and summing gives a far better number than forcing an average across three very different call types.
Watch the output price above all else. Because output rates run three to five times input rates, and because response length is the thing that quietly grows as you improve your product, that field moves your bill more than any other. When you tune, tune output length first.
Run the calculator twice, once with your current prompt and once with a trimmed version, then compare the annual projections. Seeing a $4,000 gap makes the case for a two-hour prompt cleanup far better than any argument about tokens.
Set your buffer based on how predictable your users are. A background job that always sends the same-sized payload needs almost no buffer. A public chat where users can paste an entire document into the box needs 25 percent or more, because a handful of huge requests can double a day’s spend.
Re-run the estimate whenever you change the system prompt. Adding three sentences of instructions to a prompt that fires 10,000 times a day is not free, and the added tokens ride along on every single call for the rest of the prompt’s life.
Use the cost-per-request number to design your pricing and your free tier. If a call costs you $0.002 and you give away 500 free calls, that free user costs you $1 a month before support and hosting. Knowing this early stops you from building a free tier you cannot afford.
Check the estimate against your first real invoice within the first week of launch, not at the end of the month. If the daily cost is running double your estimate, you want to know on day two, not day thirty.
When comparing two models, hold every field constant except the two prices. That isolates the pricing difference and gives you a clean multiplier, the kind of number in Example 5 where the same volume cost 43 times more on a frontier model than a standard one.
Round your final budget up, never down. The buffered budget already pads the estimate, and rounding $960 to $1,000 rather than $900 costs you nothing in planning while protecting you from a bad surprise.
Common mistakes to avoid
Using one blended price instead of two
The most frequent error is collapsing input and output into a single average rate, usually picked to make the math feel simpler. Because output costs several times more than input, a blended rate either overstates cheap input-heavy workloads or badly understates output-heavy ones.
A content generator sending 500 input tokens and receiving 3,000 output tokens, priced with a single averaged rate, can come out 60 percent below its real cost. That is the gap between a budget you meet and one you blow through in week three.
Always enter the two prices separately, exactly as the provider lists them. The calculator keeps them apart for this reason.
Forgetting the system prompt and context in input tokens
People count the user’s message and stop there. The system prompt, the retrieved documents, and the conversation history all count as input tokens on every call, and together they often dwarf the user’s actual question.
Measure a full real request, not just the visible user text. A chat that feels like a 20-token question can carry 3,000 tokens of hidden context behind it.
Assuming test-phase volume equals production volume
During testing you make a few dozen calls a day and the bill is invisible. Production traffic can be a hundred times that, and the cost scales linearly with it. Underestimating request volume is the single costliest mistake in this whole tool.
Estimate production volume from your user count and expected calls per user, then apply a buffer on top. Do not extrapolate a monthly budget from a testing week.
Setting days per month to 30 for a weekday tool
An internal tool used only Monday through Friday runs about 22 days a month, not 30. Using 30 overstates the bill by roughly 36 percent, which sounds harmless until it makes a viable project look too expensive to approve.
Match the days field to real traffic. Consumer apps get 30, weekday tools get 22, and anything seasonal gets its own honest number.
Ignoring the annual projection
A $200 monthly bill feels fine and gets waved through. The same bill is $2,400 a year, and that framing is what actually decides whether a prompt-optimization sprint is worth an engineer’s week. Skipping the yearly view hides the decision.
Read the annual projection on every serious estimate. Monthly numbers feel small in a way that yearly numbers do not, and yearly is the scale at which optimization pays off.
Treating the buffer as optional padding to delete
Some people zero out the buffer to see a “cleaner” number and then budget from it. That unpadded figure has no room for the spikes, retries, and prompt growth that every real product experiences.
Keep a buffer of at least 15 percent unless your workload is genuinely fixed-size. The padding is the difference between an estimate and a budget.
Confusing tokens with words or characters
Entering a word count where the field asks for tokens undercounts by about a third, since English averages 0.75 words per token. Entering a character count overcounts by roughly four times.
Convert properly or use a tokenizer. When in doubt, multiply your word count by 1.33 to approximate tokens.
When to use this calculator
Reach for this before you launch any feature that calls an LLM API for real users. The moment usage moves from you clicking a test button to thousands of automated calls a day, the cost stops being noise and starts being a line item someone will ask about. A five-minute estimate here prevents the awkward conversation that follows an unexpected invoice.
Use it when you are choosing between models. The pricing gap between a budget model and a frontier one is not 2x, it can be 40x or more, and seeing that spread on your actual volume is what makes the tradeoff concrete instead of theoretical. Run the same usage through both price sets and let the annual figures decide.
It also earns its keep during pricing design. If you are building a product on top of the API, your cost per request sets the floor under everything you charge. Knowing that a call costs $0.0016 tells you what a usage-based plan or a free tier can sustainably offer.
The best time to learn what an AI feature costs is before you build it. The second best time is the day you launch, from a real invoice. Everything after that is cleanup.
When it is not worth opening: a one-off script you will run a handful of times, or an experiment with fewer than a hundred total calls, does not need a budget model. The whole point of the tool is projecting sustained, repeated usage, so for genuinely throwaway work the arithmetic is not worth your time.
Related calculators
- Multi-Model API Cost Comparator
- LLM Token Cost Calculator with Caching
- Prompt Length vs Cost Calculator
- Inference Cost Comparison
- RAG Cost Calculator
- AI Agent Cost Simulator
- Context Window Cost Optimizer
- AI SaaS Pricing Calculator
Glossary
Token – The unit language models bill on, roughly four characters or 0.75 words of English text. Both what you send and what the model returns are counted in tokens.
Input tokens – Everything sent to the model in a request: system prompt, context, history, and the user’s message. Billed at the input rate.
Output tokens – The tokens the model generates in its reply. Billed at the output rate, which is almost always higher than the input rate.
Cost per million tokens – The standard pricing unit providers publish, quoted separately for input and output.
Cost per request – The combined input and output charge for one average API call, the building block for scaling any volume estimate.
Blended cost – A single averaged rate covering both input and output. Convenient but imprecise, and a frequent source of wrong estimates.
Safety buffer – A percentage added to a raw estimate to absorb spikes, retries, and prompt growth, producing a budget you are unlikely to exceed.
Why is output priced higher than input if it is usually fewer tokens? Generation is more computationally expensive than reading, so providers charge more for each token the model produces. That is why output length dominates so many bills despite the smaller token count.
Annual projection – The monthly cost multiplied by twelve, used for yearly planning and optimization decisions. Assumes flat usage.
Context window – The maximum number of tokens a model can process in one request. Large contexts raise input token counts and therefore cost.
RAG – Retrieval-augmented generation, where documents are fetched and added to the prompt. It inflates input tokens, sometimes heavily.
Unit economics – The per-request or per-user cost that determines whether a product built on the API can be profitable.
Frontier model – A provider’s most capable and most expensive tier, where output pricing can exceed $60 per million tokens.
Frequently asked questions
How accurate is this estimate?
It is exactly as accurate as your token counts and volume. The arithmetic is precise; the uncertainty lives entirely in your inputs. If you measure real requests with a tokenizer and pull volume from logs, the estimate typically lands within 10 percent of the invoice.
The largest source of error is usually output length, which varies more than people expect. Averaging 20 real responses instead of guessing one number closes most of that gap.
Why is my output cost higher than my input cost when I send more input?
Because output is priced several times higher per token. A request with 4,000 input tokens and 500 output tokens can still spend more on output when the output rate is 3x the input rate.
At the default prices of $0.50 input and $1.50 output, output tokens cost three times as much each, so 500 output tokens match 1,500 input tokens in dollars. Check the price fields whenever the split surprises you.
Should I include the system prompt in input tokens?
Yes, always. The system prompt is sent on every request and billed as input every time. A 400-token system prompt on a feature making 10,000 daily calls adds 4 million input tokens a day.
Anything that travels with the request counts: system prompt, few-shot examples, retrieved context, and prior conversation turns. If it goes over the wire, it is billed.
Leaving it out is one of the most common reasons an estimate comes in low.
What buffer percentage should I use?
For predictable, fixed-size workloads, 10 to 15 percent is plenty. For public-facing features where users control input length, use 20 to 30 percent because a few oversized requests can swing a day’s total.
If you have no data yet, start at 20 percent. Once you have a month of real billing, tighten it to match the variance you actually see.
Does this account for caching or batch discounts?
No, this calculator estimates standard on-demand pricing. Prompt caching and batch APIs can cut costs substantially, sometimes by 50 percent or more on the cacheable portion, but they apply only to specific usage patterns.
If you use caching or batching, estimate the standard cost here first, then apply your provider’s discount to the relevant slice. Dedicated caching and batch calculators handle those cases with the right math.
How do I convert words to tokens for the input fields?
Multiply your word count by roughly 1.33 for English, since text averages about 0.75 words per token. A 750-word document is close to 1,000 tokens.
Code and non-English languages pack fewer words per token, so for those, lean 10 to 20 percent higher. A tokenizer beats any word-count shortcut every time.
Why does a low-volume tool sometimes cost more than a high-volume one?
Because token size per request can outweigh raw call count. A research tool making 300 calls a day with 12,000-token prompts can outspend a chatbot making 1,200 calls a day with 2,000-token prompts.
Volume and per-request size multiply together. The calculator shows both, which is why plugging in real numbers beats intuition about which workload is “bigger.”
Can I use this for multiple models at once?
Run it once per model and sum the results. If your app routes easy calls to a cheap model and hard calls to an expensive one, estimating each stream separately is far more accurate than averaging.
For a side-by-side view of two or three models on the same workload, the Multi-Model API Cost Comparator is built for exactly that comparison.
Does the annual projection account for growth?
No, it assumes today’s usage holds flat for twelve months. Real products grow, so treat the projection as a floor. If you are adding 10 percent of usage a month, your real annual cost is well above twelve times the current month.
For a growing product, re-run the estimate each month with updated volume rather than trusting a single early projection.
Disclaimer
This calculator produces educational estimates for planning purposes. The numbers depend on the specific model, provider, hardware, configuration, and current pricing you use, all of which change frequently and without notice. Treat every figure as a starting point for your own analysis, not a guaranteed cost. The live calculator on this page is the source of truth for its exact fields, so use it directly rather than relying on the descriptions here.
Provider prices move, models are deprecated and replaced, and token accounting differs slightly between providers. An estimate that was accurate last quarter may not hold today. Verify the current input and output rates on your provider’s official pricing page before committing to any budget.
Cost outputs here are not financial or business advice. They do not account for taxes, minimum commitments, volume discounts, caching, batch pricing, or the many other line items a real bill can carry. Your actual invoice is the only authoritative number.
Before scaling spend on any AI feature, run a small real test at low volume, measure the actual token usage and cost from your provider’s dashboard, and reconcile it against this estimate. Adjust your inputs to match reality, then plan from the corrected figure.








The variability in tokenization overhead is often overlooked when calculating budget. Using tiktoken or a custom BPE tokenizer is standard for counting, but static estimation usually fails because of how system prompts and few-shot examples impact the KV cache context across different models. I have seen production RAG pipelines where the context window utilization sits at 90 percent of the 128k limit, causing huge spikes in input token costs that simple calculators miss. When you factor in FP8 quantization or even AWQ for inference, the latency improves, but the cost per request remains pegged to the provider input-output pricing. Has your math accounted for how prompt template bloat scales with dynamic retrieval from vector databases like Pinecone or Qdrant? It is common for devs to forget that each retrieval call adds significant overhead.
Regarding your point on RAG pipeline overhead, you are correct that dynamic context injection is the primary culprit for budget variance. Many developers calculate based on static prompt lengths and fail to account for the token count of retrieved chunks which often fluctuates based on the Top-K parameters used in the embedding search. We are currently testing an update to the calculator that allows for a variable context input, which should help model those dynamic RAG scenarios more accurately than a fixed token count.
That would be helpful. If you implement variable context, make sure to include a field for calculating the cost of cache hits if the provider exposes that, as that is becoming a major factor for models supporting long-context caching.
Integrating context caching rates is definitely the next logical step for this tool. Many providers are now charging for the initial prefill and then a reduced rate for subsequent calls that use the cached context, which changes the ROI calculation significantly. We will look into adding a cache-hit ratio field in the next release.
I used the alpha of this tool weeks ago. It is much more granular than the basic calculators OpenAI or Anthropic provide on their dashboards. I am currently running this against my internal bot usage and it is pretty close, though I wish it had an export to CSV feature.
Glad to hear the granular approach is proving useful for your internal testing. Export functionality is high on our roadmap, specifically the ability to download a monthly projection report that breaks down daily versus peak-load costs. We are prioritizing the CSV export to make it easier to plug these numbers into existing financial tracking tools.