This calculator shows what the context you carry into every call actually costs, and what you save by trimming it. Context is re-billed as input on every single call, so a large system prompt or long history quietly multiplies across your volume. It is built for developers deciding whether summarizing history or retrieving fewer chunks is worth the effort against just sending everything.
Loading calculator...
The habit of stuffing the window feels harmless because each call looks cheap. Across 100,000 calls, though, every extra 1,000 context tokens becomes real money, and this tool puts a monthly figure on the part you could cut.
How to use it
Enter your context tokens per call first, meaning the system prompt, conversation history, and any documents you attach to every request. Then set the output tokens, the monthly call count, and your input and output prices.
The last field, keep this percentage of context, is where you model a change. If you plan to summarize history or retrieve only the top chunks and expect to carry half the tokens, set it to 50. The tiles then show the full-context cost, the trimmed cost, and the saving.
Set the keep percentage from a real plan, not a wish. If summarizing your chat history typically compresses it to a third of its length, keep 33, and remember the summary itself adds a few tokens back.
Below the tiles, a table shows four fixed keep levels, 100%, 75%, 50%, and 25%, with the monthly cost and saving at each. It lets you see the whole trimming curve without changing the field repeatedly.
Fields explained
Context tokens / call – system prompt, history, and documents sent every call. Default 8000, step 1.
Output tokens / call – reply length, unaffected by trimming. Default 500, step 1.
Calls / month – monthly call volume. Default 100000, step 1.
Input price / 1M – input token rate. Default 3, step 0.01.
Output price / 1M – output token rate. Default 15, step 0.01.
Keep this % of context – the share of context you retain after trimming. Default 50, step 1.
Reading the results
| Result | What it means | How to act |
|---|---|---|
| Full context / mo | Monthly cost keeping 100% of context | Your current baseline |
| Trimmed / mo | Monthly cost at your keep percentage | The bill after trimming |
| Saved / mo | Dollar and percentage saving | Weigh against quality risk and effort |
| Levels table | Cost and saving at 100, 75, 50, 25 percent | See the whole trimming curve |
At the defaults, full context costs $3,150 a month, keeping 50% costs $1,950, and the saving is $1,200, or 38.1%. The saving is not 50% because output cost is fixed and does not shrink when you trim the prompt.
That fixed output floor is the key to reading the result. Only input shrinks when you trim, never output. At the defaults, $750 of the monthly cost is output and stays put no matter how aggressively you cut context.
Trimming saves money only if the model still has what it needs. Cutting relevant history or dropping the chunk that held the answer lowers quality and can trigger retries or follow-up questions that cost more than the tokens you saved. Verify answers after trimming.
When output dominates, trimming barely helps. A call with 1,000 context tokens and 2,000 output tokens saves almost nothing from halving context, because the input was already small.
The formula
The cost at any keep level is the retained context priced as input, plus the fixed output cost:
outputCost = calls x outputTokens x outPrice / 1e6
costAt(keep) = calls x (contextTokens x keep) x inPrice / 1e6 + outputCost
saved = costAt(1) - costAt(keep)
Walk the defaults. Output cost is 100,000 x 500 x 15 / 1e6 = $750, present at every level. Full context is 100,000 x 8,000 x 3 / 1e6 plus $750, which is $2,400 plus $750, or $3,150. Keeping 50% is 100,000 x 4,000 x 3 / 1e6 plus $750, which is $1,200 plus $750, or $1,950. The saving is $1,200.
| Context kept | Monthly cost (default) | Saved vs full |
|---|---|---|
| 100% | $3,150.00 | $0.00 |
| 75% | $2,550.00 | $600.00 |
| 50% | $1,950.00 | $1,200.00 |
| 25% | $1,350.00 | $1,800.00 |
The saving is linear in the tokens you cut, but the percentage is diluted by the fixed output floor. At the defaults, cutting to 25% removes $1,800 of a $2,400 input cost, yet the total only falls to $1,350 because $750 of output remains.
Change the output tokens and the whole table shifts by the same fixed amount, since output is added at every level.
Worked examples
Small context, cheap model. 3,000 context, 300 output, 10,000 calls, at 0.5 / 1.5, keeping 50%. Output cost is $4.50. Full context is 10,000 x 3,000 x 0.5 / 1e6 plus $4.50, or $19.50. Trimming to half gives $12.00, saving $7.50, a 38.5% cut. Small numbers, same shape.
Default mid-scale. The shipped inputs give $3,150 full, $1,950 at 50%, and $1,200 saved at 38.1%. The levels table shows a $1,800 saving if you can reach 25%.
Large context at scale. 32,000 context, 800 output, 500,000 calls, at 3 / 15, keeping 25%. Output cost is $6,000. Full context is $54,000, and keeping a quarter is $18,000. Trimming to 25% saves $36,000 a month here. Aggressive retrieval pays off hardest when context is huge.
Output-dominated call. 1,000 context, 2,000 output, 100,000 calls, at 3 / 15, keeping 50%. Output cost is $3,000 against just $300 of input. Full is $3,300 and trimmed is $3,150, saving only $150, a 4.5% cut. Here the lever is shorter output, not less context.
Common mistakes
Expecting the saving to match the trim. Halving context does not halve the bill, because output is fixed. Read the percentage from the tile rather than assuming it equals the cut.
Trimming past what the model needs. Cutting relevant context lowers quality. The cheapest call is worthless if the answer is wrong and the user asks again.
Do not trim blind to accuracy. Measure answer quality before and after a cut on a real sample. A context reduction that saves $1,200 a month but raises the wrong-answer rate can cost far more in support load and lost trust than it saves in tokens.
Ignoring output when it dominates. If output is the larger cost, context trimming is the wrong lever. Cap response length instead.
FAQ
Why does halving context not halve the cost?
Output cost is fixed and added at every level, so only the input portion shrinks. At the defaults, $750 of output stays while input falls from $2,400 to $1,200.
That is why the 50% level saves 38.1%, not 50%.
How do I actually trim context?
Summarize older conversation history into a short recap, retrieve only the most relevant chunks instead of whole documents, and drop system prompt boilerplate you can move to a fine-tune or shorter instruction.
Each of these lowers the tokens re-sent on every call, which is what the keep percentage models.
Does trimming hurt answer quality?
It can, if you remove material the model needed. The saving is only real when the trimmed context still contains the relevant information.
Test on a sample of real queries and compare answers before trimming in production.
What keep percentage should I use?
Base it on your trimming method. Summarization often reaches 25% to 40% of the original, while dropping to top-k retrieval depends on how many chunks you were sending.
Use the levels table to see the saving at each round number before committing to one.
Is this the same as prompt caching?
No. Caching keeps the full context but discounts a repeated prefix, while trimming sends fewer tokens. They can combine: trim what you do not need, then cache the stable part of what remains.
For the caching side, use the LLM token cost calculator with caching.
Disclaimer
This is an educational estimate for planning. Results depend on the token counts, prices, and keep percentage you enter, and on your provider’s current rates, all of which change. It prices tokens only and cannot tell you whether a trimmed context still answers correctly.
The live calculator on this page is the source of truth for its exact fields. Confirm current prices, measure answer quality on a real sample before and after trimming, and run a small test before rolling a context change into production.







