Fine-tuning a model looks cheap until the second invoice arrives. The training run is a one-time charge you can see coming, but the fine-tuned model you deploy usually bills at a higher per-token rate than the base model it came from, and that rate applies to every request for as long as the model stays in production. This calculator separates those two costs so you can see the one-time training charge and the recurring inference charge side by side, then judge whether a custom model earns its keep at your actual request volume.
Loading calculator...
It is built for solo developers and small teams deciding between prompt engineering on a base model and paying to train a specialized one. You enter the size of your training set, how many passes the trainer makes over it, the model tier you are fine-tuning, and your expected monthly traffic. The tool returns the training bill, the monthly inference bill, a first-year total, and a per-thousand-request unit cost you can compare against a base-model baseline.
The numbers are estimates for planning. Real providers publish their own training and inference prices, and those prices move, so treat the output as a way to compare scenarios rather than a quote. The point is to catch the expensive surprises before you commit a dataset to a training run.
How to use the Fine-Tuning Cost Estimator
Start with your training set. Enter the number of training examples you have prepared, then the average tokens per example. If you have a JSONL file ready, the example count is just the line count, and the average token length is something you can measure on a sample of a few hundred rows rather than guessing. A support-ticket dataset might average 300 tokens per example, while a long-document summarization set could run past 1,500.
Set the training epochs next. An epoch is one full pass over your dataset. Most fine-tuning jobs use 3 by default, and this is the single input most likely to quietly multiply your training bill, because the token count you pay for scales directly with it. Two epochs on the same data costs two thirds of what three epochs costs.
Pick the base model tier you intend to fine-tune. The tier sets three prices at once: the training rate per million tokens, and the input and output rates the fine-tuned model will bill at afterward. A Small tier trains cheaply and serves cheaply but has less headroom for hard tasks. A Frontier tier trains at twenty times the Small rate and serves at a steep multiple, so you only reach for it when a smaller model has already failed the task on your evaluation set.
Measure your average example length on a real sample before you fill in the field. Estimating 400 tokens when your examples actually average 900 will understate the training bill by more than half, and the error compounds with every epoch.
Now describe your traffic. Enter the monthly inference requests you expect the fine-tuned model to serve, then the average input and output tokens per request. Input tokens are everything you send, including the system prompt and any retrieved context. Output tokens are what the model generates. These two fields drive the recurring cost, and for any high-traffic deployment they matter far more than the training charge.
Read the results as a pair. The one-time training cost is what you pay to create the model. The monthly inference cost is what you pay to run it. The first-year total combines them, and the per-thousand-request figure lets you line the fine-tuned model up against a base model serving the same traffic. If the base model would answer as well, the delta is the price of customization.
Calculator fields explained
Number of training examples – the count of labeled examples in your training set, one per line in a typical JSONL file. Default 5,000. Training cost scales linearly with this, so doubling the set doubles the training charge at fixed epochs and example length.
Average tokens per example – the mean token length of a single training example, counting both the prompt and the target completion. Default 500 tokens. Measure it on a sample rather than estimating, since prompt templates and retrieved context inflate this more than people expect.
Training epochs – how many complete passes the trainer makes over the dataset. Default 3. Fewer epochs cost less and risk underfitting; more epochs cost proportionally more and risk overfitting to the training data.
Base model tier – the model family you are fine-tuning, which fixes the training rate and the fine-tuned inference rates. Options are Small (cheapest, best for narrow classification and extraction), Base (the default, a general workhorse), Large (for reasoning-heavy tasks), and Frontier (highest quality and highest cost, reserved for tasks smaller tiers fail). Default Base.
Monthly inference requests – the number of calls you expect to make to the fine-tuned model each month. Default 100,000. This multiplies the per-request token cost into your recurring bill.
Average input tokens per request – the mean number of tokens you send per call, including the system prompt and any context. Default 400 tokens. Fine-tuning often lets you shorten prompts, which lowers this number and is one of the real ways a custom model pays for itself.
Average output tokens per request – the mean number of tokens generated per call. Default 200 tokens. Output tokens usually bill at roughly twice the input rate, so a chatty model costs more than the request count alone suggests.
Understanding the results
| Result | What it means | How to act on it |
|---|---|---|
| Total training tokens | Examples multiplied by average length multiplied by epochs | The billable unit for training; shrink it by cutting epochs or trimming examples |
| One-time training cost | Training tokens priced at the tier’s training rate | Compare against the value of a better model; it is paid once |
| Monthly inference cost | Requests priced at the fine-tuned input and output rates | The number that recurs forever; the main lever on total spend |
| First-year total cost | Training plus twelve months of inference | Use it to compare fine-tuning against staying on the base model |
| Cost per 1,000 requests | Monthly inference divided by requests, scaled to a thousand | Line it up against the base model’s per-thousand cost |
The hero number for most people is the first-year total, because it puts the visible one-time charge and the easy-to-forget recurring charge on the same footing. A $60 training run feels trivial, but if it locks you into $240 a month of inference, the honest cost of the decision is closer to $2,940 across the year.
The monthly inference cost is where the leverage lives. Training happens once. Inference happens on every request, every day, and a fine-tuned model at a higher token rate keeps charging that premium whether it answers well or poorly. When you are weighing a custom model against a base model, the per-thousand-request figure is the comparison that actually decides it.
At high volume, inference dwarfs training by a factor of hundreds. A dataset of 500 short examples might cost under two dollars to train, and the same model serving half a million requests a month can run past eleven thousand dollars a year. The training charge is a rounding error against that; the tier you picked, which set your inference rate, is the real decision.
The edges of the input range tell you which cost dominates. A very large dataset with low traffic pushes almost all the cost into training, which means a single run and then near-free serving. A tiny dataset with heavy traffic inverts it, and the training charge becomes noise against the inference bill. Read where your scenario sits before you optimize, because trimming epochs on a training-dominated job saves real money while trimming them on an inference-dominated job saves nothing worth the risk to model quality.
The fine-tuned inference rate is usually higher than the base model’s rate for the same tier. If your fine-tuned model answers no better than the base model with a good prompt, you are paying that premium on every single request for nothing.
One more reading habit worth building: check whether fine-tuning lets you shorten your prompts. A base model often needs a long system prompt and several few-shot examples to behave. A fine-tuned model has that behavior baked in, so your average input tokens can drop sharply. That reduction shows up directly in the monthly inference cost and is one of the few ways a custom model genuinely lowers your bill instead of raising it.
Calculation formulas
The tool runs two calculations that meet in the first-year total. The first prices the training run.
training tokens = examples × avg_tokens_per_example × epochs
training cost = (training tokens ÷ 1,000,000) × training_rate
The second prices a month of serving the fine-tuned model.
monthly inference = requests × [(input_tokens ÷ 1,000,000 × ft_input_rate) + (output_tokens ÷ 1,000,000 × ft_output_rate)]
And the two combine into the annual view.
first-year total = training cost + (monthly inference × 12)
The tier you select supplies the three rates. These are the region-neutral planning rates the calculator uses, expressed in dollars per million tokens.
| Tier | Training ($/1M tokens) | Fine-tuned input ($/1M) | Fine-tuned output ($/1M) |
|---|---|---|---|
| Small | 3.00 | 0.40 | 1.60 |
| Base | 8.00 | 3.00 | 6.00 |
| Large | 25.00 | 9.00 | 18.00 |
| Frontier | 60.00 | 18.00 | 36.00 |
Walk one case through end to end. Take 5,000 examples averaging 500 tokens, trained for 3 epochs on the Base tier, then served to 100,000 monthly requests at 400 input and 200 output tokens each. Training tokens come to 5,000 × 500 × 3 = 7,500,000. At the Base training rate that is 7.5 × 8.00 = $60.00. For inference, each request costs (400 ÷ 1,000,000 × 3.00) + (200 ÷ 1,000,000 × 6.00) = 0.0012 + 0.0012 = $0.0024, and 100,000 of them make $240.00 a month. The first-year total is 60 + 240 × 12 = $2,940.00, and the per-thousand-request cost is 240 ÷ 100 = $2.40.
Epochs multiply only the training side of the ledger. Going from 3 epochs to 6 doubles the training charge but leaves the monthly inference cost untouched, because inference bills per request regardless of how the model was trained.
Notice that the inference rates, not the training rate, decide the shape of the year. The $60 training charge is fixed the moment the run finishes. The $2,880 of annual inference is what you signed up for by choosing the Base tier, and it is the number that changes if your traffic grows.
Practical examples
Each scenario below states the inputs, runs the formula, and ends with what the number tells you to do.
Example 1, a solo classifier. 1,000 examples at 300 tokens, 3 epochs, Small tier, serving 10,000 requests at 300 input and 150 output tokens. Training tokens are 1,000 × 300 × 3 = 900,000, costing 0.9 × 3.00 = $2.70. Monthly inference is 10,000 × [(300 ÷ 1M × 0.40) + (150 ÷ 1M × 1.60)] = 10,000 × 0.00036 = $3.60. First-year total: 2.70 + 43.20 = $45.90. A whole custom model for under fifty dollars a year, because both the data and the traffic are small.
Example 2, a small support bot. 2,000 examples at 400 tokens, 3 epochs, Base tier, 20,000 requests at 400 input and 200 output. Training tokens 2,000 × 400 × 3 = 2,400,000, costing 2.4 × 8.00 = $19.20. Monthly inference 20,000 × 0.0024 = $48.00. First-year total 19.20 + 576 = $595.20, at $2.40 per thousand requests. The recurring side already outweighs training twelve to one.
Example 3, the default configuration. 5,000 examples at 500 tokens, 3 epochs, Base tier, 100,000 requests at 400/200. This is the walkthrough above: $60.00 training, $240.00 monthly, $2,940.00 for the year. Useful as the anchor everyone can adjust from.
The first fine-tuning bill I ever approved was the training run. The one that actually hurt was the twelfth inference invoice, because nobody had multiplied the monthly number by twelve before we shipped.
Example 4, a growing product. 8,000 examples at 600 tokens, 4 epochs, Large tier, 200,000 requests at 500 input and 300 output. Training tokens 8,000 × 600 × 4 = 19,200,000, costing 19.2 × 25.00 = $480.00. Monthly inference 200,000 × [(500 ÷ 1M × 9.00) + (300 ÷ 1M × 18.00)] = 200,000 × 0.0099 = $1,980.00. First-year total 480 + 23,760 = $24,240.00, or $9.90 per thousand. The Large tier’s inference rate is doing the damage here, not the training.
Example 5, epoch sensitivity. Hold 5,000 examples at 500 tokens on the Base tier and change only the epochs. At 1 epoch, training tokens are 2,500,000 and the charge is $20.00. At 6 epochs, tokens rise to 15,000,000 and the charge is $120.00. Same data, same model, a six-fold spread in training cost driven by one field. Pick epochs from an evaluation curve, not out of habit.
Example 6, production scale on Frontier. 20,000 examples at 800 tokens, 3 epochs, Frontier tier, 1,000,000 requests at 600 input and 400 output. Training tokens 20,000 × 800 × 3 = 48,000,000, costing 48 × 60.00 = $2,880.00. Monthly inference 1,000,000 × [(600 ÷ 1M × 18.00) + (400 ÷ 1M × 36.00)] = 1,000,000 × 0.0252 = $25,200.00. First-year total 2,880 + 302,400 = $305,280.00. At this scale the training charge is under one percent of the year.
Example 7, tiny data and heavy traffic. 500 examples at 200 tokens, 2 epochs, Base tier, 500,000 requests at 350 input and 150 output. Training tokens 500 × 200 × 2 = 200,000, costing 0.2 × 8.00 = $1.60. Monthly inference 500,000 × [(350 ÷ 1M × 3.00) + (150 ÷ 1M × 6.00)] = 500,000 × 0.00195 = $975.00. First-year total is $11,701.60 on a $1.60 training run. When traffic is this heavy, obsessing over training cost is a waste of attention; the tier and the token counts are everything.
Example 8, big data and light traffic. 50,000 examples at 700 tokens, 3 epochs, Large tier, only 5,000 requests at 400 input and 200 output. Training tokens 50,000 × 700 × 3 = 105,000,000, costing 105 × 25.00 = $2,625.00. Monthly inference 5,000 × [(400 ÷ 1M × 9.00) + (200 ÷ 1M × 18.00)] = 5,000 × 0.0072 = $36.00. First-year total 2,625 + 432 = $3,057.00. Here training is 86 percent of the year, the exact mirror of Example 7, and trimming an epoch would save real money.
Tips and best practices
Run the estimate before you clean the dataset, not after. The two numbers that move the training bill most are epochs and average example length, and both are decisions you make while preparing data. Knowing that six epochs cost twice what three cost changes how carefully you validate before committing the run.
Multiply the monthly inference figure by twelve every single time. The training charge is loud and one-time; the inference charge is quiet and forever. A model that looks like a $60 decision is really a $2,940 decision once you annualize it, and that framing is the one that should drive the tier choice.
Fine-tune to shorten prompts, then measure the saving. If your base-model prompt carries a 600-token system message and four few-shot examples, a fine-tuned model that needs only 150 input tokens cuts your input bill by three quarters. That reduction often pays back the training run faster than any quality improvement does.
Start one tier lower than you think you need. Small and Base tiers train and serve at a fraction of Large and Frontier rates. Prove the smaller model fails your evaluation set before you pay for a bigger one, because the tier decision follows you into every future invoice through the inference rate.
Watch epochs like a cost lever, not a quality dial. More passes do not reliably mean a better model, and past a point they overfit. Treat the epoch count as something you tune against a validation curve, and let the calculator show you what each additional pass adds to the training bill.
Build the base-model baseline first. Estimate what the same traffic would cost on the base model with a good prompt, then compare. If the fine-tuned per-thousand cost is not lower or the quality gain is not measurable, keep the prompt and skip the training run.
Keep your token measurements honest by sampling real production traffic rather than idealized test prompts. Retrieved context, conversation history, and tool definitions all land in the input count, and a request you imagined at 200 tokens can arrive at 900. The inference estimate is only as good as those two token averages.
Re-run the estimate whenever traffic changes by more than a factor of two. The whole cost picture shifts when you move from a training-dominated regime to an inference-dominated one, and the optimization that made sense at 5,000 requests a month is the wrong one at 500,000. A quick Base tier recheck takes a minute and can redirect where you spend engineering effort.
Common mistakes to avoid
Pricing only the training run
The most common error is treating the training charge as the cost of fine-tuning and stopping there. Training is the entry fee. The recurring inference bill, charged at the fine-tuned rate on every request, is where nearly all the money goes at any real volume.
A $19 training run on the Base tier that serves 100,000 requests a month is a $2,900-a-year commitment, not a $19 one. Approving fine-tuning on the training number alone is how teams walk into recurring costs they never sized.
Fix it by reading the first-year total, not the training line. That figure already annualizes the inference charge and is the honest number to weigh against staying on the base model.
Ignoring the higher fine-tuned inference rate
Fine-tuned models typically bill at a higher per-token rate than the base model of the same tier. Teams that assume the rate is unchanged budget the training charge, ship the model, and then find every request costs more than it did before.
The inference premium, not the training run, is the costliest mistake to miss. It applies forever and scales with traffic, so a small rate difference becomes a large annual number. Check the fine-tuned input and output rates for your tier and price a month of real traffic against them before you commit.
Setting epochs by default and forgetting them
Leaving epochs at the default and never revisiting it wastes training budget on data-heavy jobs. Each epoch multiplies the token count you pay for, and past the point of diminishing returns you are paying to overfit.
Cranking epochs to 6 or 8 hoping for a better model usually just doubles the training bill without a matching quality gain. Tune the count against a validation curve instead of assuming more is better.
Set epochs deliberately, watch validation loss flatten, and stop there. The calculator makes the cost of each extra pass visible, which is exactly the discipline this mistake needs.
Estimating token lengths instead of measuring them
Guessing average example length or per-request tokens produces an estimate that can be off by a factor of two or three. Prompt templates, retrieved context, and system messages all add tokens people forget to count.
Measure both averages on a real sample. A few hundred training rows and a day of production logs give you numbers you can trust, and the whole estimate rests on them.
Comparing against nothing
Running the fine-tuning estimate without a base-model baseline leaves you with a number and no way to judge it. Two thousand dollars a year is cheap or expensive only relative to the alternative.
Price the same traffic on the base model with a strong prompt first. If the base model matches the quality at a lower per-thousand cost, the fine-tuning premium buys you nothing.
Forgetting that data grows
People size a fine-tuning job for today’s dataset and today’s traffic, then retrain quarterly as data accumulates. Each retrain is another training charge, and a model whose traffic keeps climbing keeps raising the inference bill underneath it.
Estimate the retraining cadence and the traffic trajectory, not just the first run. A model you retrain four times a year at $480 a run carries $1,920 of annual training cost on top of inference.
When to use this calculator
Reach for it whenever you are deciding between prompt engineering on a base model and paying to train a custom one. That decision hinges on two numbers this tool produces: the first-year total and the per-thousand-request cost. If the fine-tuned model is not measurably better or measurably cheaper per request, the base model with a careful prompt is the rational choice, and the calculator makes that comparison concrete.
It earns its place again at budget time. Before you approve a training run, the annualized inference figure tells finance what they are actually signing up for, which is usually many times the visible training charge. A team that brings the first-year total to that conversation avoids the awkward moment three months later when the recurring bill lands.
Use it when you are choosing a tier. The tier decision sets your inference rate for the life of the deployment, so the difference between Base and Large is not a one-time choice, it is a rate you pay on every request. Running both tiers through the calculator at your real traffic shows the true spread.
Fine-tuning is worth it when a smaller, cheaper model does the job a bigger model was doing, and the training charge buys you a permanently lower inference rate. It is worth it far less often than the pitch decks suggest.
Skip the calculator when your traffic is trivial and your data is tiny, because at that scale nearly any configuration costs a few dollars and the estimate will not change your decision. Skip it too when you have not yet proven a base model fails the task; fine-tuning to fix a problem a better prompt would have solved is spending you can avoid entirely.
Related calculators
- Multi-Model API Cost Comparator
- Monthly AI API Budget Calculator
- Inference Cost Comparison
- Distributed Training Total Cost Estimator
- Training Time Estimator
- Self-Hosting Break-Even Calculator
- Quality vs Cost Tradeoff Advisor
Glossary
Fine-tuning – continuing to train a pre-trained model on your own labeled examples so it specializes in your task and format.
Base model – the general-purpose model you start from before any custom training.
Training example – one labeled item in your dataset, usually a prompt paired with the target completion, stored as a single line in a JSONL file.
Epoch – one complete pass of the trainer over the entire dataset. Training cost scales directly with the number of epochs.
Training tokens – the total tokens processed during training, equal to examples times average length times epochs. This is the billable quantity for the training run.
Inference – running the finished model to answer requests in production, billed per token rather than per training run.
Why does a fine-tuned model cost more per token than the base model it came from? Providers charge a premium for serving a custom model because it occupies dedicated capacity and cannot be batched with other customers’ traffic as efficiently as a shared base model.
Input tokens – the tokens you send in a request, including the system prompt, any retrieved context, and the user message.
Output tokens – the tokens the model generates in response, usually billed at a higher rate than input tokens.
Model tier – the size class of the model, which fixes both the training rate and the fine-tuned inference rates. Larger tiers cost more at every stage.
Overfitting – when too many epochs make a model memorize the training set instead of generalizing, so it performs worse on real inputs.
Per-thousand-request cost – monthly inference cost divided by requests and scaled to a thousand, the unit that lets you compare a fine-tuned model against a base model directly.
First-year total – the one-time training cost plus twelve months of inference, the figure that reflects the true cost of the decision.
Baseline – the cost and quality of solving the same task on the base model with a good prompt, the reference every fine-tuning decision should be measured against.
Frequently asked questions
Is fine-tuning cheaper than using a base model?
Not usually on cost alone. Fine-tuned models bill at a higher per-token rate than their base model, so you pay a training charge upfront and then a premium on every request, which raises your bill rather than lowering it in most cases.
The exception is when fine-tuning lets you drop to a smaller tier or drastically shorten your prompts. If a fine-tuned Small model does the work a prompted Base model was doing, the per-thousand cost can fall by more than half, and that is when the training charge pays for itself.
How many epochs should I use?
Three is the common default and a reasonable starting point. Fewer risks underfitting, and more risks overfitting while multiplying the training charge.
The right number comes from watching validation loss flatten, not from a rule. Since each epoch scales the training cost linearly, going from 3 to 6 on a 15-million-token job adds $60 on the Base tier for a gain you should confirm before paying for it.
Why is inference cost so much larger than training?
Training is a one-time charge on a fixed number of tokens. Inference recurs on every request for as long as the model is deployed, so at any meaningful volume it accumulates far past the training figure.
A 500-example dataset can cost $1.60 to train and over $11,000 a year to serve at half a million monthly requests. The ratio between the two is a signal about where to spend your optimization effort.
Read the first-year total to see both on the same scale, and let the monthly inference figure, annualized, drive your tier and token decisions.
Does a bigger training set always mean a better model?
No. Quality depends more on how clean and representative your examples are than on raw count, and past a few thousand good examples the returns flatten for many tasks.
A curated set of 2,000 examples often beats a noisy set of 20,000, and it costs a tenth as much to train. Measure quality on a held-out evaluation set rather than assuming more data buys more performance.
What token length should I enter for my examples?
Measure it rather than guess. Take a sample of a few hundred rows from your training file and compute the average token count across both the prompt and the completion.
An estimate off by a factor of two doubles or halves your training bill. Prompt templates and system instructions inflate example length more than people expect, so the measured number is almost always higher than the guessed one.
Can I lower the recurring cost after training?
Yes, mainly by shrinking your per-request token counts. Fine-tuning often lets you cut the system prompt and few-shot examples that a base model needed, which lowers input tokens directly.
Dropping average input from 600 to 150 tokens on 100,000 monthly requests at the Base rate saves about $135 a month, or $1,620 a year, which frequently exceeds the entire training charge.
How often will I need to retrain?
It depends on how fast your data and task drift. Teams with stable tasks retrain rarely, while those in fast-moving domains retrain quarterly or monthly as fresh examples accumulate.
Each retrain is another training charge, so a quarterly cadence at $480 a run adds $1,920 of annual training cost on top of inference. Factor the cadence into your first-year planning rather than treating training as a single event.
Which tier should I fine-tune?
Start with the smallest tier that could plausibly handle the task, because the tier sets your inference rate for the life of the deployment. Small and Base train and serve for a fraction of Large and Frontier.
Only move up when a smaller tier fails your evaluation set. The rate difference is not a one-time cost; on the Large tier you pay roughly three times the Base input rate on every request forever, so the tier choice deserves more scrutiny than the training charge does.
Disclaimer
This calculator produces educational estimates for planning. The training and inference rates it uses are region-neutral placeholders chosen to illustrate how the costs relate, and they are not quotes from any provider. Your actual costs depend on the model, the provider, the hardware, your configuration, and current pricing, all of which change frequently.
The cost outputs are not financial or business advice. They are meant to help you compare scenarios and spot which cost dominates yours, not to tell you what to spend or which model to buy.
Before you rely on any of these numbers, verify them against your provider’s current published training and inference prices, and against your own measured token usage from real traffic. A token average taken from production logs will always beat one taken from an idealized test prompt.
Run a small real fine-tuning job and serve a modest slice of traffic through it before scaling spend. The live calculator on this page is the source of truth for its exact fields and current behavior, so use it directly, and let a measured pilot confirm the estimate before you commit a full budget.







