Every extra sentence you add to a prompt has a price, and that price repeats on every single request you send. This calculator takes the length of your prompt, the length of the reply you expect back, and the per-million-token rates your provider charges, then tells you what one request costs and what a month of them adds up to. It is built for solo developers, indie founders, and small teams who are watching an API bill grow and want to know which half of the equation, the prompt or the response, is driving it.
Loading calculator...
Most people size their prompt by feel. They paste in a few examples, a system message, some retrieved context, and ship it. The bill that arrives three weeks later rarely explains itself. What this tool does is separate input cost from output cost so you can see, in dollars, whether trimming a 6,000-token context down to 4,000 is worth the effort or a rounding error.
The numbers stay small per request and large per month, which is exactly the trap. A request that costs two cents feels free. Multiply it by 200,000 calls and you are looking at a real line item. Working backward from monthly spend to the token count that caused it is the whole point of the exercise.
How to use the Prompt Length vs Cost Calculator
Start with the prompt length field. This is the total number of input tokens you send on a typical request, counting the system prompt, any few-shot examples, retrieved documents, and the user’s actual message. A rough English conversion is about 0.75 tokens per word, so a 2,000-word context lands near 1,500 tokens. If you have a token counter handy, use its real number instead of the estimate.
Next, set the expected output length. This is how many tokens the model writes back on an average call. A short classification reply might be 20 tokens. A full drafted email could be 500. A long structured report can run past 1,500. Output tokens almost always cost more per million than input tokens, so this field carries more weight than its size suggests.
Then enter the two prices. Providers publish these as a cost per million tokens, split into an input rate and an output rate. Copy them from your provider’s pricing page exactly as listed. A budget model might charge 0.15 for input and 0.60 for output. A frontier model might charge 3.00 and 15.00, or higher. The label shows USD; convert if your provider bills in another currency.
Enter the prices in dollars per million tokens exactly as your provider lists them. Do not divide by anything first. The calculator handles the per-million math for you, so a rate of 3.00 goes in as 3.00, not 0.000003.
Last, put in your requests per month. If you only know a daily figure, multiply by 30. If traffic is spiky, use an average rather than a peak day, because the monthly total smooths those spikes out anyway.
Once all five fields are filled, the results update. Read the total cost per request first to sanity-check the scale, then look at the prompt share percentage to see where your money is actually going. If that share is above 70 percent, the prompt is your cost problem, not the response.
Calculator fields explained
Prompt length (input tokens) – The count of tokens you send to the model on one request, including system prompt, examples, context, and user message. Default is 1500. Measured in tokens. Use a real token count when you can, since word-based estimates drift on code, JSON, and non-English text.
Expected output length (output tokens) – The number of tokens the model returns on a typical call. Default is 500. Measured in tokens. Set this from observed replies, not from your maximum token limit, because the limit is a ceiling the model usually does not reach.
Input price (USD per million tokens) – What your provider charges per million input tokens. Default is 3.00. Shown in USD. This is the cheaper of the two rates on almost every model.
Output price (USD per million tokens) – What your provider charges per million output tokens. Default is 15.00. Shown in USD. Output usually costs three to five times the input rate, which is why a chatty model gets expensive fast.
Requests per month – How many calls you make in a month. Default is 10000. A plain count. Use an average across the month rather than your busiest day, or the monthly projection will overstate your bill.
Understanding the results
| Result | What it means | How to act on it |
|---|---|---|
| Input cost per request | Dollars spent sending your prompt on one call | Compare against output cost to find the bigger half |
| Output cost per request | Dollars spent on the model’s reply on one call | If high, cap the response length or ask for shorter answers |
| Total cost per request | Input plus output for a single call | Multiply mentally by your call volume to feel the scale |
| Monthly cost | Total per request times requests per month | Put this straight into your budget line |
| Prompt share of cost (%) | Fraction of each request driven by the prompt | Above 70 percent means trim the prompt first |
| Cost per 1,000 requests | What a batch of 1,000 calls costs | Use it to price a feature or a customer tier |
The hero number here is monthly cost, because that is the figure you defend in a budget meeting or measure against revenue. Everything else exists to explain it. At the default settings, a 1,500-token prompt and a 500-token reply on a 3.00 and 15.00 model across 10,000 requests, the monthly cost lands at 120 dollars. The input side of that is 45 dollars and the output side is 75 dollars.
The prompt share percentage is the number that changes your behavior. At those defaults it reads 37.5 percent, which tells you the response is doing most of the spending. Trimming the prompt would help a little, but capping the output would help more. Flip the inputs to a huge retrieved context and a one-line answer and the share climbs past 90 percent, which sends you in the opposite direction.
A low cost per request hides a high monthly total. Two cents per call sounds harmless until you remember you are making a quarter of a million calls, at which point that harmless number is 5,000 dollars. Always read the monthly figure alongside the per-request one.
Cost per 1,000 requests is the unit you want when you are pricing a product. If a paying customer triggers roughly 1,000 calls a month and that batch costs you 12 dollars, you now know your raw model cost per user before margin. That single number turns an abstract API bill into a per-seat figure you can build a price around.
Prompt share above 70 percent means cut the prompt before anything else. Below that, your effort is better spent shortening replies or moving to a cheaper model for the easy calls.
Edge cases show up at the extremes. A tiny prompt with a long generation, like a one-line instruction that produces a 2,000-token essay, will show an output-dominated split near 90 percent. A giant system prompt with a yes-or-no answer flips it the other way. Neither is wrong, they just point your optimization at different fields.
When the monthly figure looks impossibly small, check your requests-per-month entry. A common slip is entering a daily count and reading it as monthly, which understates the bill by a factor of thirty.
Calculation formulas
The math is straightforward once you keep the per-million division in the right place. For one request:
input cost = input_tokens / 1,000,000 x input_price
output cost = output_tokens / 1,000,000 x output_price
total per request = input cost + output cost
monthly cost = total per request x requests_per_month
prompt share = input cost / total per request x 100
cost per 1,000 requests = total per request x 1,000
Walk the default through it. Input tokens are 1,500, so 1,500 divided by 1,000,000 is 0.0015, times the input price of 3.00 gives 0.0045 dollars. Output tokens are 500, so 0.0005 times 15.00 gives 0.0075 dollars. Add them for 0.012 per request. Times 10,000 requests, the month is 120 dollars. The prompt share is 0.0045 divided by 0.012, which is 37.5 percent. Cost per 1,000 is 0.012 times 1,000, or 12 dollars.
Output tokens usually cost three to five times what input tokens cost on the same model. That ratio is why a 500-token reply can outspend a 1,500-token prompt even though it uses a third of the tokens. Watch the rate, not just the count.
The reference table below shows common per-million rates across tiers so you can see how the same token counts swing in cost. These are representative bands, not any one provider’s live prices.
| Tier | Input (USD per 1M) | Output (USD per 1M) | Output to input ratio |
|---|---|---|---|
| Budget | 0.15 | 0.60 | 4.0x |
| Standard | 0.50 | 1.50 | 3.0x |
| Mid | 3.00 | 15.00 | 5.0x |
| Frontier | 5.00 | 15.00 | 3.0x |
Notice that the output-to-input ratio is not constant across tiers. Budget models can punish long replies more aggressively in relative terms than frontier ones, so the cheapest model is not always cheapest for a generation-heavy workload.
Practical examples
Each example below states the scenario, the inputs, the calculation, the result, and what you would do next. The numbers come straight from the formulas above.
Example 1, a solo classifier. A hobby project tags support messages. Inputs: 300 prompt tokens, 200 output tokens, 0.15 input, 0.60 output, 500 requests. Input is 300/1M x 0.15 = 0.000045. Output is 200/1M x 0.60 = 0.00012. Total per request 0.000165, monthly 0.08 dollars. The cost is a rounding error, so spend your time on accuracy, not tokens.
Example 2, a small chatbot. A side product answers FAQs. Inputs: 800 prompt, 400 output, 0.50 input, 1.50 output, 2,000 requests. Input 0.0004, output 0.0006, total 0.001 per request, monthly 2 dollars. Prompt share 40 percent. Still cheap enough to ignore optimization entirely.
Example 3, a summarizer with long input. A tool condenses articles. Inputs: 6,000 prompt, 400 output, 3.00 input, 15.00 output, 1,500 requests. Input 0.018, output 0.006, total 0.024, monthly 36 dollars. Prompt share 75 percent. The context is the cost, so shortening the fed-in article pays off directly.
The moment your prompt share crosses 70 percent, you stop tuning the model’s output and start editing your own input. That is the single most reliable lever a small team has on an LLM bill.
Example 4, a RAG feature. Retrieval stuffs documents into context. Inputs: 12,000 prompt, 600 output, 3.00 input, 15.00 output, 5,000 requests. Input 0.036, output 0.009, total 0.045, monthly 225 dollars. Prompt share 80 percent. The retrieved chunks are the bill.
Example 5, the same RAG feature trimmed. You cut retrieval from six chunks to two. Inputs: 4,000 prompt, 600 output, same prices, 5,000 requests. Input 0.012, output 0.009, total 0.021, monthly 105 dollars. That single change saved 120 dollars a month against Example 4, and the answers barely moved.
Example 6, production on a frontier model. A live app at scale. Inputs: 2,000 prompt, 800 output, 5.00 input, 15.00 output, 200,000 requests. Input 0.010, output 0.012, total 0.022, monthly 4,400 dollars. Prompt share 45.5 percent. At this volume, both halves matter and a cheaper model deserves a test.
Example 7, the same load on a budget model. You route the easy calls elsewhere. Inputs: 2,000 prompt, 800 output, 0.15 input, 0.60 output, 200,000 requests. Input 0.0003, output 0.00048, total 0.00078, monthly 156 dollars. The model swap took the same traffic from 4,400 to 156 dollars, though quality must be checked before you trust it.
Example 8, a runaway system prompt. A giant instruction block returns a one-word answer. Inputs: 20,000 prompt, 50 output, 3.00 input, 15.00 output, 3,000 requests. Input 0.060, output 0.00075, total 0.06075, monthly 182.25 dollars. The prompt drives 98.8 percent of the cost here. Caching that static system prompt would gut the bill.
Tips and best practices
Measure your real token counts before you trust any estimate. Word-based conversions are fine for prose, but code, JSON, tables, and languages other than English tokenize very differently, sometimes at double the tokens per word. A quick pass through your provider’s tokenizer removes the guesswork and stops you from optimizing the wrong field.
Read the prompt share number before you touch anything. It decides your whole strategy. High share means edit the input, low share means cap the output or change the model. Acting without that percentage is how people spend a week shaving system prompts that were never the problem.
Separate the static part of your prompt from the dynamic part. A system message and few-shot examples that never change are perfect candidates for context caching, which many providers bill at a fraction of the normal input rate. The variable user text stays full price, but the fixed scaffolding can drop sharply.
Set your output token limit to just above your longest reasonable answer rather than leaving it at the model default. A generous cap does not raise cost by itself, but it removes the guardrail that stops a runaway generation from billing you for 4,000 tokens you never wanted.
Test cheaper models on a slice of live traffic instead of committing blind. Example 6 versus Example 7 showed a swing from 4,400 to 156 dollars a month on identical volume. That gap is worth a controlled experiment even if only 60 percent of your calls can safely move down a tier.
Price your feature in cost per 1,000 requests, not cost per token. Per-token figures are too small to reason about. If a customer triggers a thousand calls and that batch costs you 12 dollars, you can build a subscription price on top of a number you can actually hold in your head.
Recheck the calculator whenever your provider updates prices, which happens more often than most people track. A rate change of even a dollar per million on the output side reshapes every projection you have made. The formula stays the same, the coefficients do not.
Watch for prompts that grow silently. A retrieval system that started fetching three chunks and now fetches eight has quietly tripled its input cost while nobody changed a line of application logic. Re-running this calculator monthly catches that drift before the bill does.
Do not over-optimize a cheap workload. If the monthly cost sits under a few dollars, as in Examples 1 and 2, the engineering time to shave tokens costs more than the tokens ever will. Spend that attention where the monthly figure has three digits.
Common mistakes to avoid
Entering a daily request count as monthly
The single most frequent error is putting a day’s traffic into the requests-per-month field. If you make 10,000 calls a day and enter 10,000, the calculator shows a month that is really a day, understating your true bill by thirty times.
A 300-dollar monthly estimate that is actually a daily one hides a 9,000-dollar bill. This is the costliest slip in the whole tool, and it comes from a single mislabeled field.
Confirm the units before you read the result. Multiply your daily average by 30, enter that, and the monthly figure will match what your provider actually charges. A daily count read as monthly undercounts the bill thirtyfold.
Using the max output limit as the expected output
Setting expected output to your token ceiling, say 4,096, assumes every reply hits the maximum. Real replies almost never do. A model told it can write up to 4,096 tokens might average 300, so budgeting for the ceiling overstates your output cost by an order of magnitude.
Pull the real average from a sample of production responses instead. If your logs show replies clustering around 350 tokens, use 350, and let the occasional long answer average out.
Ignoring the output-to-input price gap
Treating input and output as if they cost the same leads you to fixate on prompt length while the reply quietly does the damage. On a mid-tier model the output rate can be five times the input rate, so a short prompt with a long answer costs more than the token counts suggest.
Counting only tokens and not their prices is how a team spends a sprint compressing prompts while the output side, billed at five times the rate, goes untouched and keeps the total flat.
Always look at input cost and output cost as separate dollar figures, not token counts. The calculator splits them for exactly this reason.
Estimating tokens from words on non-prose content
The 0.75-tokens-per-word rule works for English sentences and falls apart on structured data. JSON payloads, code, and markup can run well past one token per character in places, so a word-based estimate on a code-heavy prompt can undercount by half.
Feed representative samples through the real tokenizer. The difference between an estimated 1,500 tokens and a measured 2,800 tokens is the difference between an accurate budget and a fantasy one.
Forgetting that context resets every request
A stateless API charges you for the full prompt on every single call, including the system message and examples you have sent a thousand times before. People assume the model remembers, then wonder why a fixed system prompt keeps costing money.
Count the repeated scaffolding as part of every request’s input, because that is how billing sees it. Caching is the fix, but only after you have measured what the repetition costs.
Optimizing a workload that is already free
Spending hours trimming a prompt on a project that costs two dollars a month is negative return on your time. The monthly figure tells you whether the effort is worth it before you start.
Set a threshold, maybe 50 dollars a month, below which you simply do not optimize. Direct the saved attention at the features whose monthly cost actually shows up in the total.
When to use this calculator
Reach for this tool when an API bill has started to matter and you cannot tell whether the prompt or the response is responsible. The split between input cost and output cost is the specific thing it gives you that a raw invoice does not, and that split changes what you do next.
It earns its place during a pricing decision. Before you set a subscription tier or a usage cap, the cost per 1,000 requests turns your model spend into a per-customer number you can build margin on top of. Guessing that figure instead of computing it is how a plan ends up losing money on its heaviest users.
The best time to run these numbers is before you ship the feature, not after the invoice explains, in hindsight, what your prompt design decided for you.
It also helps when you are weighing a model swap or a caching change. Plug the same token counts into two price pairs and the monthly difference is immediate, as Examples 6 and 7 showed with their swing from 4,400 to 156 dollars. That comparison justifies or kills the migration in one screen.
There are times it is not worth opening. If your monthly cost is a handful of dollars, the tool will confirm that and you should move on. And if you have not measured your real token counts yet, do that first, because feeding it guesses produces confident numbers built on sand.
Related calculators
- Multi-Model API Cost Comparator
- LLM Token Cost Calculator with Caching
- Monthly AI API Budget Calculator
- System Prompt Overhead Calculator
- Few-shot Example Cost Amplifier
- Prompt Compression Savings
- Context Window Cost Optimizer
- Paste-Text Token Counter
Glossary
Input tokens – The tokens you send to the model, covering system prompt, examples, retrieved context, and the user message. Billed at the input rate.
Output tokens – The tokens the model generates in its reply. Billed at the output rate, which is usually higher than the input rate.
Token – The unit a model reads and writes in, roughly three-quarters of an English word on average, though this varies widely by content type.
Cost per million tokens – The standard way providers publish pricing, giving a dollar figure for every 1,000,000 tokens of input or output.
Prompt share – The percentage of a request’s cost that comes from the input side rather than the output side.
Why do output tokens cost more than input tokens when both are just text? Generation is more computationally expensive than reading. The model produces each output token one at a time through a full forward pass, while it can process the input prompt in a more parallel fashion, and pricing reflects that difference.
System prompt – The fixed instruction block that sets the model’s behavior, sent on every request and billed every time unless cached.
Few-shot examples – Sample input-output pairs included in the prompt to steer the model, which add input tokens on every call.
Context caching – A provider feature that stores a repeated prompt segment and bills it at a reduced rate on later requests.
Monthly cost – The total cost per request multiplied by requests per month, the headline budget figure.
Cost per 1,000 requests – The cost of a batch of a thousand calls, a handy unit for pricing features and customer tiers.
RAG – Retrieval-augmented generation, a pattern that fetches documents and inserts them into the prompt, which inflates input token counts.
Tiered pricing – The grouping of models into bands like budget, standard, mid, and frontier, each with its own input and output rates.
Frequently asked questions
How do I count the tokens in my prompt?
Use your provider’s tokenizer or a token-counting endpoint, which gives an exact number for the text you paste in. That is more reliable than any word-based rule.
If you need a quick estimate without a tool, multiply your word count by roughly 1.3 for English prose. A 1,000-word context lands near 1,300 tokens, though code and JSON will run higher, sometimes well past 2,000 for the same word count.
Why is my output cost higher than my input cost when the reply is shorter?
Because output tokens are priced higher, often three to five times the input rate on the same model. A shorter reply can still outspend a longer prompt.
On a mid-tier model at 3.00 input and 15.00 output, a 500-token reply costs 0.0075 dollars while a 1,500-token prompt costs 0.0045. The reply is a third of the length and 67 percent more expensive.
Does the system prompt count toward every request?
Yes. A stateless API bills the full input on every call, so a system message you send unchanged a thousand times is charged a thousand times.
A 500-token system prompt at a 3.00 input rate adds 0.0015 dollars to every request. Across 100,000 requests that fixed block alone costs 150 dollars a month, which is why caching it matters at volume.
Should I switch to a cheaper model to cut costs?
Often yes for straightforward tasks, but test quality before you commit. The savings can be large, as shown when identical traffic dropped from 4,400 to 156 dollars a month between a frontier and a budget model.
Route the easy calls to the cheap model and keep the hard ones on the expensive one. A hybrid split usually beats an all-or-nothing choice on both cost and quality.
The right move depends on how many of your calls genuinely need the stronger model. If only 30 percent do, the other 70 percent can move down a tier and take most of the bill with them.
Is caching worth setting up?
It pays off when a large chunk of your prompt is identical across requests, like a long system message or a fixed set of examples. The bigger and more repeated that chunk, the more caching saves.
For the runaway system prompt in Example 8, where the prompt drove 98.8 percent of a 182-dollar monthly cost, caching the static portion could cut the bill by most of that share. Caching a repeated system prompt can remove most of the input cost.
What if my traffic is very spiky?
Use an average across the month rather than a peak day. The monthly projection smooths spikes out, so an average gives you a truer figure than your busiest hour would.
If a normal day is 5,000 calls but launch days hit 40,000, base the estimate on the monthly total divided by 30, not the launch spike. Otherwise you will budget for a ceiling you rarely touch.
Why does my word-based token estimate keep coming out low?
Because the 0.75-tokens-per-word rule assumes plain English prose. Structured content breaks it.
Code, JSON, tables, and non-English text tokenize far more densely, sometimes at two tokens per word or more. A code-heavy prompt you estimated at 1,500 tokens can measure 2,800 in reality, which nearly doubles the input cost you budgeted for.
Can I use this for providers that bill in a currency other than USD?
Yes, as long as you convert the per-million rates into the same currency before entering them and read the outputs in that currency. The math does not care which currency the numbers represent.
Convert both the input and output rates at the same exchange figure so the two stay consistent. The prompt share percentage is currency-neutral either way, since it is a ratio of two costs in the same units.
Disclaimer
This calculator produces educational estimates for planning purposes. The figures it returns depend on the model you choose, the provider you use, your exact token counts, and the prices in effect on the day you run it, all of which change over time and vary between vendors.
Cost outputs are not financial or business advice. They are a way to reason about token spend before you commit, not a guarantee of what any provider will charge you. Treat the monthly projection as a starting point for your own budgeting rather than a fixed bill.
Verify every rate against your provider’s current pricing page and check your token counts against a real tokenizer before relying on the numbers. Published rates move, and a small change on the output side reshapes every projection here. The live calculator on this page is the source of truth for its exact fields, so use it directly rather than working from these examples.
Run a small real test at low volume before scaling spend. A few hundred actual requests will tell you more about your true cost than any estimate, and they will surface the token drift and output-length surprises that a projection cannot see.








This calculator is cute for someone just starting out, but it ignores the reality of API overhead and model drift. Calling this a tool for decision making is a stretch when most enterprise bills are driven by RAG implementation errors, not just token count. You are modeling a static environment while LLM output length fluctuates wildly based on system prompt temperature settings. I have seen developers waste weeks optimizing for token cost while their actual latency issues were caused by inefficient vector database lookups. Show me a tool that tracks end-to-end request lifecycle and I will listen.
Regarding the focus on static modeling, you raise a fair point about the complexity of production environments. This calculator is specifically designed for API-based workflows where developers often lack visibility into how system prompt bloat impacts monthly invoices. You are correct that RAG retrieval latency and vector database performance are separate, often larger, bottlenecks. We focus on token billing because providers like OpenAI and Anthropic charge per-token, and those costs scale linearly. For tracking request lifecycles, we recommend looking into observability tools like LangSmith or Arize Phoenix, which provide the granularity needed to identify latency spikes caused by retrieval layers.
That makes sense. I have used LangSmith for tracing, but I find the UI often masks the underlying token math. I will look into your suggestion for integrating those metrics with my current cost tracker.
Glad to hear that helps. Integrating tracing data with your financial models usually provides the clarity needed to distinguish between architectural inefficiencies and pure operational costs. Let us know if you find a specific setup that bridges that gap well.
Does this handle quantization overhead? When I run local inference on an A100 or a 4090, the VRAM footprint is the primary constraint, not just the token price. I spend my time managing KV cache sizes and testing 4-bit vs 8-bit configs to keep latency under 50ms. Most folks dont realize that packing a massive context window into a 24GB card leads to severe thermal throttling if you are pushing constant throughput. I need to know if this factors in the hardware utilization metrics or if it is strictly for cloud API pricing.
About your question on hardware metrics, this tool is strictly for cloud API pricing models where the provider manages the compute layer. When you are hosting models locally on an A100 or 4090, the cost structure changes from per-token billing to power consumption, depreciation, and infrastructure overhead. To manage VRAM effectively during inference, you should look into vLLM for optimized throughput and PagedAttention, which handles KV cache memory management much better than standard Hugging Face implementations. Monitoring hardware thermal throttling and token-per-second throughput is essential for local setups, but that logic falls outside the scope of this particular financial calculator.