This calculator compares an all-real-time API bill against a realistic mix where only some jobs can tolerate batch turnaround. Batch endpoints trade latency for a discount, so the saving depends on both the discount size and the share of work you can actually route to batch. It is aimed at teams with a large volume of jobs, some of which run in the background where a few hours of delay costs nothing.
Loading calculator...
The number people miss is the batchable share. A 50% batch discount sounds like a 50% saving, but if only a fifth of your jobs can wait, the real cut is far smaller. This tool multiplies the two together so the headline discount does not fool you.
How to use it
Start with your monthly job count and the average tokens per job, counting input and output together since batch pricing usually discounts both. Then enter your blended price per million tokens, which is your normal real-time rate.
The last two fields set the discount and the batchable share. Batch discount is the percentage off you get for accepting delayed turnaround. Jobs that can be batched is the share of your work that tolerates that delay, meaning offline or background tasks rather than anything a user waits on.
Decide the batchable share by task type, not by hope. Evals, bulk data enrichment, and overnight generation can batch. Anything in a live request path cannot, so exclude it from the percentage.
The three tiles show the all-real-time monthly cost, the cost with batching applied to the eligible share, and the monthly saving with its percentage. The saving grows only when both the discount and the batchable share are high.
Fields explained
Jobs per month – total jobs processed monthly. Default 1000000, step 1.
Tokens per job (in + out) – combined tokens per job. Default 2000, step 1.
Blended price per 1M tokens – your real-time rate across input and output. Default 5, step 0.01.
Batch discount (%) – the percentage off for batch turnaround. Default 50, step 1.
Jobs that can be batched (%) – share of jobs that tolerate delay. Default 80, step 1.
Reading the results
| Result | What it means | How to act |
|---|---|---|
| All real-time | Monthly cost with every job at full price | Your baseline |
| With batching | Cost with the eligible share discounted | Your bill after routing batchable work |
| Saved / month | Dollar gap and percentage cut | Weigh against the effort of a batch pipeline |
At the defaults, all real-time costs $10,000 a month, batching brings it to $6,000, and the saving is $4,000, or 40.0%. That 40% comes from an 80% batchable share at a 50% discount, since the two multiply.
The relationship is worth internalizing. Savings equal the batchable share times the discount. An 80% share at a 50% discount gives 40%, and no combination of the two beats the discount itself.
A generous discount is wasted if little of your work can batch. At a 50% discount but only a 20% batchable share, the saving drops to 10%, since the other 80% still pays full price. Grow the batchable share before chasing a bigger discount.
Batch turnaround is typically measured in hours, not seconds. Route only latency-tolerant work to it, and keep anything a person waits on in real time regardless of the saving.
The formula
The mixed cost applies the discount to the batchable share and full price to the rest:
perJobReal = tokensPerJob x price / 1e6
perJobBatch = perJobReal x (1 - discount)
allRealTime = jobs x perJobReal
mixed = jobs x batchable x perJobBatch + jobs x (1 - batchable) x perJobReal
Walk the defaults. Per job real-time is 2,000 x 5 / 1e6 = $0.01, and per job batched is $0.005. All real-time is 1,000,000 x 0.01 = $10,000. The mixed cost is 1,000,000 x 0.8 x 0.005 plus 1,000,000 x 0.2 x 0.01, which is $4,000 plus $2,000, or $6,000. The saving is $4,000.
| Batchable share | Saving at a 50% discount |
|---|---|
| 20% | 10% |
| 50% | 25% |
| 80% | 40% |
| 100% | 50% |
Because the saving is batchable share times discount, the discount is a hard ceiling. Even if every job batches, you save exactly the discount, never more. This is why the batchable share is the field to grow.
Raising the discount lifts the ceiling, but the batchable share decides how close you get to it.
Worked examples
Half the work batchable. 100,000 jobs, 1,500 tokens, $5 per million, 50% discount, 50% batchable. Per job is $0.0075, so all real-time is $750. Mixed is 100,000 x 0.5 x 0.00375 plus 100,000 x 0.5 x 0.0075, which is $562.50, saving $187.50, a 25.0% cut.
Default high-volume mix. The shipped inputs give $10,000 real-time, $6,000 with batching, and $4,000 saved at 40.0%. A strong result driven by the 80% batchable share.
Large offline workload. 5,000,000 jobs, 3,000 tokens, $8 per million, 50% discount, 90% batchable. Per job is $0.024, so all real-time is $120,000. Mixed is 5,000,000 x 0.9 x 0.012 plus 5,000,000 x 0.1 x 0.024, which is $66,000, saving $54,000. Batching saves $54,000 a month here.
Mostly latency-bound. 1,000,000 jobs, 2,000 tokens, $5 per million, 50% discount, but only 20% batchable. All real-time is $10,000 and mixed is $9,000, saving just $1,000, a 10.0% cut. The discount is fine; the eligible share is the limit.
Common mistakes
Reading the discount as the saving. A 50% discount is not a 50% saving unless every job can batch. Multiply by the batchable share to get the real figure.
Overstating what can wait. Anything a user watches load cannot batch. Being honest about the batchable share keeps the estimate from promising savings you cannot capture.
Do not route latency-sensitive work to batch to chase the discount. Batch jobs can take hours to return, so pushing a live feature onto batch to save money breaks the user experience and often costs more in lost usage than it saves in tokens.
Ignoring the pipeline cost. Batch needs job submission, polling, and result handling. For small volumes the saving may not cover the engineering, so check the dollar figure, not just the percentage.
FAQ
Why is my saving less than the discount?
Because only the batchable share gets the discount; the rest pays full price. The saving is the share times the discount, so an 80% share at 50% gives 40%.
The reference table shows how the saving tracks the batchable share at a fixed discount.
Does batch pricing discount input and output equally?
This tool uses one blended price and one discount across combined tokens. If your provider discounts input and output differently, blend them into the single price and discount fields.
Enter a weighted average so the per-job figure matches your real mix.
What kind of work can I batch?
Offline and background tasks: evaluations, bulk enrichment, overnight report generation, and anything where a few hours of delay is fine.
Live chat, autocomplete, and any user-facing response must stay real-time and should be excluded from the batchable share.
How long does batch turnaround take?
Typically hours rather than seconds, and providers often quote a completion window rather than a guarantee. Plan around the upper bound of that window.
If your task needs results within minutes, batch is the wrong tool regardless of the discount.
Is a small percentage saving worth building for?
Look at the dollar figure, not the percentage. A 10% saving on a $100,000 bill is $10,000 and likely worth a pipeline; the same percentage on a $200 bill is not.
The saved-per-month tile gives you that absolute number directly.
Disclaimer
This is an educational estimate for planning. Results depend on the volumes, prices, discount, and batchable share you enter, and on your provider’s current batch rates and turnaround, all of which change. It models token cost only and omits the engineering cost of a batch pipeline.
The live calculator on this page is the source of truth for its exact fields. Confirm your provider’s current batch discount and completion window, and run a small batch job before committing latency-tolerant work at scale.








This calculator is actually super helpful for my thesis work. I’ve been processing research datasets through an API and didn’t realize how much the batchable share matters. I was looking at a 50% discount and thinking I’d save half my budget, but most of my jobs are real-time queries that users submit. Only my overnight data enrichment can wait. Plugging in 15% batchable share at 50% discount gives me 7.5% savings, which is way less exciting but at least now I know what to expect before I build out the batch pipeline. The formula explanation helped me understand why the discount alone doesn’t mean much if you can’t actually use it.
Regarding your thesis workflow, you’ve identified exactly the decision point that matters most. That 7.5% savings calculation is honest, and it actually informs a bigger architectural question: is the engineering effort to implement batch processing worth the ROI? For research datasets, the answer often depends on your infrastructure maturity and how tightly your budget is constrained. One thing worth considering as you scale: if you can shift even a few percentage points of work toward batch without sacrificing user experience, the compounding effect becomes more interesting. For instance, if you moved from 15% to 25% batchable share at that same 50% discount, you’d hit 12.5% savings. Some teams find they can batch things like background feature generation or async report compilation that they initially thought required real-time processing. The calculator gives you a clear frame to revisit that assumption as your pipeline matures. Feel free to test different batchable percentages to find your actual threshold where implementation effort justifies the cost reduction.