A retrieval-augmented generation pipeline hides its bill in a place most people never look: the retrieved chunks you paste into every prompt. You pay for those tokens on each query, forever, and they usually dwarf the question itself. This calculator takes the numbers you already control, retrieved chunks, chunk size, output length, and the price per million tokens for your model, and returns the cost of a single query, the recurring monthly total, and the one-time cost of building your index.
Loading calculator...
It is built for solo developers wiring a chatbot into their docs, small teams shipping a support assistant, and anyone deciding whether RAG belongs in a product before the first invoice lands. You enter usage and pricing; it separates the one-time indexing spend from the per-query spend that scales with traffic.
The reason this matters is simple arithmetic. A question of 40 tokens can carry 2,000 tokens of retrieved context into the model. Your generation call is billed on all 2,040, not on the 40 you typed. Multiply that across 50,000 monthly queries and the retrieved context, not the answers, becomes the line item you optimize.
How to use the RAG Cost Calculator
Start with monthly queries. This is the number of times a user asks a question and your system runs one full retrieval and one generation call. If you are pre-launch, estimate from expected daily active users multiplied by questions per user per day, then multiply by 30.
Next set the retrieval shape. Retrieved chunks, the top-k value, is how many passages you pull from your index for each query. Tokens per chunk is the size of each passage after your chunking step. These two multiply into the context you inject, so a top-k of 5 with 400-token chunks adds 2,000 tokens to every prompt before the model reads a single word of the actual question.
Then enter the request shape. Base prompt tokens covers your system instructions plus the user question. Output tokens per query is the length of the model’s answer. Query embedding tokens is the small cost of turning the question into a vector so it can be matched against your index.
Now the prices. Input price and output price are what your model charges per million tokens, taken straight from your provider’s pricing page. Embedding price is the per-million rate of your embedding model, which is far cheaper than the generation model and applies both to queries and to indexing.
Pull the three prices from your provider on the day you plan, not from memory. Input and output rates for the same model often differ by a factor of three or more, and mixing them up is the single most common source of a wrong estimate.
Finally, knowledge base chunks to index is the total number of passages in your corpus. This drives the one-time indexing cost, the price of embedding every document once so it can be retrieved later. Set it to the size of your full corpus, not the slice you retrieve per query.
Read the results top to bottom. Cost per query tells you the marginal price of one more question. Monthly query cost is that figure at your traffic. One-time indexing cost is paid once per full re-index. The token outputs show you where the money actually goes.
Calculator fields explained
Monthly queries – the count of full RAG requests per month, each being one retrieval plus one generation. Default 50,000. Unit: queries.
Retrieved chunks (top-k) – how many passages you fetch from the index for each query. Higher values improve recall but add tokens linearly. Default 5. Unit: chunks.
Tokens per chunk – the token length of each retrieved passage. Set by your chunking strategy, commonly 200 to 600. Default 400. Unit: tokens.
Base prompt tokens – system instructions plus the user’s question, before any retrieved context is added. Default 200. Unit: tokens.
Output tokens per query – the expected length of the model’s answer. Default 300. Unit: tokens.
Query embedding tokens – the token length of the question sent to the embedding model for retrieval. Default 40. Unit: tokens.
Input price (USD per million tokens) – your generation model’s charge for input tokens. Default 0.50. Unit: USD per million tokens.
Output price (USD per million tokens) – your generation model’s charge for output tokens, usually higher than input. Default 1.50. Unit: USD per million tokens.
Embedding price (USD per million tokens) – your embedding model’s rate, applied to both queries and one-time indexing. Default 0.02. Unit: USD per million tokens.
Knowledge base chunks to index – total passages in your full corpus, used for the one-time indexing cost. Default 20,000. Unit: chunks.
Understanding the results
| Result | What it means | How to act on it |
|---|---|---|
| Cost per query | The marginal price of one full RAG request | Multiply by expected traffic to sanity-check any pricing tier you charge users |
| Monthly query cost | Per-query cost across your monthly volume | Compare against your budget; this is the number that grows with users |
| One-time indexing cost | Price to embed the whole corpus once | Add it when you first build or fully rebuild the index, not every month |
| Retrieved context tokens per query | top-k multiplied by tokens per chunk | Cut this first when the per-query cost is too high |
| Input tokens per query | Base prompt plus retrieved context | Watch the ratio of context to base; context should not swamp the question |
| Annual cost projection | Monthly recurring cost across twelve months | Use for runway planning and provider commitment discounts |
The hero number is monthly query cost. It answers the only question that keeps a project alive: what does this cost me every month at my current traffic. Everything else exists to explain that figure or to show you which lever moves it.
Look at retrieved context tokens next to input tokens per query. In the default setup, 2,000 of the 2,200 input tokens are retrieved context. The question you actually asked is a rounding error. This is normal for RAG and it is exactly why the retrieval shape, not the prompt wording, controls your bill.
Retrieved context is usually 80 to 95 percent of your input tokens.
Output cost behaves differently. Output tokens are fewer than input tokens in almost every RAG setup, but the output price per million is often two to five times higher. A 300-token answer at 1.50 per million costs more than you might expect relative to its length, so trimming verbose answers is a real saving, not a rounding exercise.
One-time indexing cost is deceptively small until your corpus is large or you re-index often. Embedding 20,000 chunks costs cents, but a pipeline that re-embeds the whole corpus on every deploy turns a one-time line into a recurring one. Re-index only changed documents.
Read the edge cases carefully. Very high top-k with a small query volume produces a high cost per query but a low monthly total, so the pipeline feels cheap while each request is wasteful. Very high query volume with a lean context produces the opposite: a tiny per-query cost that still adds up to a serious monthly figure. The two numbers tell different stories and you need both.
Annual cost projection multiplies the monthly recurring figure by twelve and leaves indexing out of the repeat, since indexing is not a monthly event. Use it when you are deciding whether an annual provider commitment or a self-hosted embedding model would pay for itself.
Calculation formulas
The calculator runs one chain of arithmetic. Retrieved context is the product of your two retrieval inputs:
context_tokens = top_k × tokens_per_chunk
Input tokens for the generation call are the base prompt plus that context:
input_tokens = base_prompt_tokens + context_tokens
Each query then has three billed components, generation input, generation output, and the query embedding:
cost_per_query = (input_tokens ÷ 1,000,000 × input_price) + (output_tokens ÷ 1,000,000 × output_price) + (query_embedding_tokens ÷ 1,000,000 × embedding_price)
Monthly recurring cost scales that by traffic, and the one-time indexing cost embeds every chunk in the corpus once:
monthly_cost = cost_per_query × monthly_queries
indexing_cost = knowledge_base_chunks × tokens_per_chunk ÷ 1,000,000 × embedding_price
Walk the default numbers through it. Context is 5 × 400 = 2,000 tokens. Input is 200 + 2,000 = 2,200 tokens. Generation input is 2,200 ÷ 1,000,000 × 0.50 = 0.0011 dollars. Generation output is 300 ÷ 1,000,000 × 1.50 = 0.00045 dollars. Query embedding is 40 ÷ 1,000,000 × 0.02 = 0.0000008 dollars. Per query that totals 0.0015508 dollars, and across 50,000 queries the monthly cost is 77.54 dollars. Indexing 20,000 chunks of 400 tokens costs 20,000 × 400 ÷ 1,000,000 × 0.02 = 0.16 dollars, paid once.
The reference table below shows rough per-million rates across common model tiers so you can see how much the input and output prices swing your result. Rates change constantly, so treat these as illustrative brackets and confirm against your provider.
| Tier | Input (USD / 1M) | Output (USD / 1M) | Embedding (USD / 1M) |
|---|---|---|---|
| Budget | 0.10 to 0.20 | 0.40 to 0.60 | 0.01 to 0.02 |
| Standard | 0.30 to 0.60 | 1.00 to 1.60 | 0.02 to 0.03 |
| Mid | 0.80 to 1.50 | 3.00 to 5.00 | 0.05 to 0.10 |
| Frontier | 2.50 to 5.00 | 10.00 to 20.00 | 0.10 to 0.13 |
The embedding price barely moves your per-query cost, because a query embeds only its own short text. It matters far more for one-time indexing, where you embed millions of tokens across the whole corpus at once. Optimize embedding spend at index time, not at query time.
Notice which term dominates. The context token count multiplied by the input price is almost always the largest slice of the per-query cost. That is the lever with the most leverage, which is why the token outputs are placed right next to the money outputs in the results.
| top-k | Tokens per chunk | Context tokens | Input tokens (base 200) |
|---|---|---|---|
| 3 | 300 | 900 | 1,100 |
| 5 | 400 | 2,000 | 2,200 |
| 8 | 500 | 4,000 | 4,200 |
| 20 | 600 | 12,000 | 12,200 |
Practical examples
Each scenario below states the inputs, runs the formula, and ends with what the reader does next. Prices reflect the tiers in the reference table.
Example 1, solo docs bot. Inputs: 2,000 queries, top-k 3, 300-token chunks, base 150, output 200, input 0.15, output 0.60, embedding 0.02, query embed 30, corpus 2,000 chunks. Context is 900, input is 1,050. Per query: (1,050 ÷ 1M × 0.15) + (200 ÷ 1M × 0.60) + (30 ÷ 1M × 0.02) = 0.0001575 + 0.00012 + 0.0000006 = 0.0002781 dollars. Monthly is 0.56 dollars; indexing is 0.012 dollars. Action: ship it, the cost is not worth optimizing yet.
Example 2, blog Q&A widget. Inputs: 5,000 queries, top-k 4, 350-token chunks, base 180, output 250, input 0.25, output 1.00, embedding 0.02, query embed 35, corpus 5,000. Context 1,400, input 1,580. Per query: 0.000395 + 0.00025 + 0.0000007 = 0.0006457 dollars. Monthly is 3.23 dollars; indexing is 0.035 dollars. Action: comfortably inside a hobby budget, revisit only if traffic grows tenfold.
Example 3, small-team support bot. Inputs: 30,000 queries, top-k 5, 400-token chunks, base 200, output 300, input 0.50, output 1.50, embedding 0.02, query embed 40, corpus 15,000. Context 2,000, input 2,200. Per query 0.0015508 dollars. Monthly is 46.52 dollars; indexing is 0.12 dollars. Action: fine to run, but the retrieved context is 90 percent of input, so trimming top-k is the first optimization.
The jump from Example 2 to Example 3 is not the model getting fancier. It is six times the traffic on a standard-tier model instead of a budget one. Volume and price tier move your bill far more than any single prompt tweak.
Example 4, growing knowledge base. Inputs: 60,000 queries, top-k 6, 450-token chunks, base 220, output 350, input 0.50, output 1.50, embedding 0.02, query embed 45, corpus 40,000. Context 2,700, input 2,920. Per query: 0.00146 + 0.000525 + 0.0000009 = 0.0019859 dollars. Monthly is 119.15 dollars; indexing is 0.36 dollars. Action: at this level, a reranker that lets you drop top-k from 6 to 3 pays for itself.
Example 5, production on a frontier model. Inputs: 200,000 queries, top-k 8, 500-token chunks, base 250, output 400, input 3.00, output 15.00, embedding 0.13, query embed 50, corpus 100,000. Context 4,000, input 4,250. Per query: 0.01275 + 0.006 + 0.0000065 = 0.0187565 dollars. Monthly is 3,751.30 dollars; indexing is 6.50 dollars. Action: this is real money, reserve the frontier model for hard queries only.
$3,751 per month for the same retrieval on a frontier model.
Example 6, same volume on a budget model. Inputs identical to Example 5 but input 0.15, output 0.60, embedding 0.02. Input tokens are still 4,250. Per query: 0.0006375 + 0.00024 + 0.000001 = 0.0008785 dollars. Monthly is 175.70 dollars. Action: the model choice alone cut the bill from 3,751 to 176 dollars, so route easy queries here and escalate only when quality demands it.
Example 7, edge case with heavy context and light traffic. Inputs: 1,000 queries, top-k 20, 600-token chunks, base 300, output 500, input 0.50, output 1.50, embedding 0.02, query embed 60, corpus 200,000. Context 12,000, input 12,300. Per query: 0.00615 + 0.00075 + 0.0000012 = 0.0069012 dollars. Monthly is 6.90 dollars; indexing is 2.40 dollars. Action: here the one-time indexing cost is a third of a month’s recurring cost, so re-indexing carelessly is your biggest waste.
Example 8, cutting top-k as an alternative config. Inputs: 100,000 queries, base 200, output 300, input 0.50, output 1.50, corpus 50,000. At top-k 8 with 400-token chunks, input is 3,400 and monthly is 215.08 dollars. At top-k 3, input is 1,400 and monthly is 115.08 dollars. Action: dropping five chunks per query saves 100 dollars a month, so test whether recall actually suffers before paying for the extra context.
Tips and best practices
Treat top-k as a budget dial, not a fixed setting. Every extra chunk multiplies across every query for the life of the pipeline. Start low, measure answer quality, and only add chunks when you can point to a question that failed without them. Most teams find that top-k 3 to 5 answers as well as top-k 10 while costing half as much.
Right-size your chunks. Small chunks retrieve precisely but need higher top-k to cover an answer; large chunks cover more per chunk but carry filler tokens you pay for. Somewhere near 300 to 500 tokens per chunk balances recall against waste for most document types.
Separate index cost from query cost in your own head. Indexing is a capital expense you pay once; queries are an operating expense you pay forever. A pipeline that looks expensive to build can be cheap to run, and the reverse is far more dangerous.
Re-index incrementally. When a document changes, embed only that document. Re-embedding the entire corpus on every content update converts your one-time cost into a recurring one, which the calculator’s indexing figure will not warn you about because it assumes a single build.
Route by difficulty. Send routine questions to a budget model and reserve a reranking pass or a frontier model for the queries that genuinely need them. Example 6 showed a 21-fold gap between tiers on identical retrieval, so a router that catches even half your traffic pays off quickly.
Cap output length. Output tokens carry the highest per-million price in most model families. A hard limit on answer length trims the most expensive tokens in your pipeline without touching retrieval quality.
Log the actual token counts your pipeline sends in production for a week, then feed the measured averages back into this calculator. Estimated context sizes are almost always lower than real ones, because retrieved chunks vary and prompt templates grow over time.
Compress your system prompt. Base prompt tokens ride on every single query alongside the context. A 400-token system prompt trimmed to 150 saves 250 tokens per query, which at 200,000 queries is 50 million tokens a month.
Store your vector database costs separately. This calculator prices the tokens flowing through the model, not the monthly cost of hosting the vectors. Add that hosting line from your database provider before you call the total complete.
Common mistakes to avoid
Ignoring retrieved context in the token count
The most frequent error is estimating cost from the question length alone. A user types 40 tokens, so the pipeline is assumed to cost 40 tokens’ worth of input. In reality the retrieved chunks add thousands of tokens to every prompt.
If your estimate ignores retrieved context, it can be off by a factor of fifty. The context tokens, not the question, are what the model bills you for on input, and they scale with top-k on every request.
Fix this by always computing input tokens as base prompt plus context, exactly as the calculator does. The retrieved context tokens output exists precisely to make this impossible to overlook.
Confusing input and output prices
Providers publish two rates for the same model, and swapping them silently wrecks an estimate. Output rates are commonly three to ten times input rates, so using the output rate for input inflates the bill wildly, while the reverse hides real cost.
Indexing your whole corpus twice doubles a one-time bill.
Read the pricing page carefully and enter each rate in its own field. When a provider lists a single blended number, ask which direction it applies to before trusting it.
Treating indexing as a monthly cost
Some planners add the one-time indexing cost to the monthly total, which overstates recurring spend and can kill a viable project on paper. Indexing happens once per full build, not every month.
Do not fold indexing into your monthly run rate. Keep it as a separate build cost, and only count it again when you fully rebuild the index or migrate embedding models. Mixing the two makes cheap pipelines look unaffordable.
Keep the two figures in separate columns of your budget. The calculator returns them separately for exactly this reason.
Setting top-k too high by default
Copying a top-k of 10 or 20 from a tutorial and never revisiting it is a quiet money leak. Each extra chunk multiplies across your entire query volume with no quality benefit past the point where the answer is already covered.
Test lower values against real questions. If top-k 4 answers as well as top-k 12, you have halved your largest cost component for free.
Forgetting vector database hosting
The tokens are only part of the bill. Hosting the vectors themselves carries a monthly cost that this calculator does not model, since it depends entirely on your database provider and index size.
Add the hosting line before you commit to a total. For a small corpus it is often a few dollars a month, but at scale it can rival the query cost.
Estimating output length too optimistically
Assuming short answers when your model tends to be verbose understates the most expensive tokens you buy. Output price per million is the highest rate in most model families.
Measure real answer lengths in production rather than guessing, and set a maximum output length to keep the number honest.
When to use this calculator
Reach for it before you commit to a model. The gap between a budget and a frontier model on identical retrieval can be more than twentyfold, as Examples 5 and 6 showed, and that decision is far cheaper to change on paper than after you have shipped.
Use it when you are pricing a feature. If you charge users per seat or per query, the cost-per-query output tells you the floor beneath your price. Charging below it means every active user loses you money, which no growth rate fixes.
Open it whenever you change the retrieval shape. A new chunking strategy, a higher top-k, or a longer system prompt all move the input token count, and the calculator turns that change into a dollar figure before it hits your invoice.
The cheapest RAG optimization is the chunk you never retrieve. Every token you leave out of the prompt is a token you never pay for, on every query, for as long as the pipeline runs.
Skip it for a throwaway prototype serving a handful of test queries. At a few hundred queries a month on a budget model, the entire bill is under a dollar, and the time spent planning costs more than the pipeline ever will. Come back when traffic or the model tier starts to matter.
Related calculators
- LLM Token Cost Calculator with Caching
- RAG vs Long Context Cost Analyzer
- Embedding Generation Cost
- Vector Database Storage Calculator
- Context Window Cost Optimizer
- Multi-Model API Cost Comparator
- Monthly AI API Budget Calculator
Glossary
RAG – retrieval-augmented generation, a pattern where relevant passages are fetched from a knowledge base and inserted into the prompt so the model answers from your data.
Top-k – the number of passages retrieved per query. A top-k of 5 means the five closest matches are pulled and added to the prompt.
Chunk – a slice of a source document, sized during ingestion, that is embedded and stored as a single retrievable unit.
Chunking – the process of splitting documents into chunks, which sets the tokens-per-chunk figure that drives context cost.
Embedding – a numeric vector representing the meaning of text, used to match a query against stored chunks.
Embedding model – the model that produces embeddings, priced far below generation models and used for both indexing and query lookup.
Vector database – the store that holds embeddings and returns the nearest matches for a query vector.
Context tokens – the retrieved passages injected into the prompt, equal to top-k multiplied by tokens per chunk.
Why is retrieved context counted as input rather than output? Because the model reads it as part of the prompt. Everything you place in front of the model, including retrieved chunks, is billed at the input rate; only the model’s generated answer is billed at the output rate.
Input tokens – every token the model reads on a request, here the base prompt plus retrieved context.
Output tokens – the tokens the model generates in its answer, usually billed at a higher per-million rate than input.
Indexing – the one-time act of embedding an entire corpus so its chunks can later be retrieved.
Reranking – a second-pass scoring step that reorders retrieved chunks so you can keep top-k low without losing the best passage.
Cost per query – the marginal price of a single full RAG request, combining generation input, generation output, and the query embedding.
Annual projection – the recurring monthly cost multiplied by twelve, used for runway and commitment planning.
Frequently asked questions
Does this calculator include vector database hosting costs?
No. It prices the tokens that flow through the embedding and generation models, which is where most RAG spend hides, but it does not model the monthly fee for hosting your vectors.
Add that line separately from your database provider. For a corpus of 20,000 chunks it is often a few dollars a month, but a corpus in the millions can push hosting into the same range as your query cost.
Why is the one-time indexing cost so low compared to the monthly cost?
Indexing embeds each chunk once at the cheap embedding rate, while queries run through the far more expensive generation model on every request. In the default setup, indexing 20,000 chunks costs 16 cents, while a month of queries costs over 77 dollars.
The imbalance flips only for tiny query volumes with a huge corpus, as in Example 7, where indexing 200,000 chunks cost 2.40 dollars against a monthly recurring bill of just 6.90 dollars.
How do I lower my cost per query the fastest?
Cut retrieved context, because it dominates the input token count. Lowering top-k or shrinking chunk size reduces the largest slice of the bill on every single query.
Cutting top-k from 8 to 3 nearly halves generation input cost. Test the lower value against real questions first to confirm answer quality holds.
Should I use a cheaper model to save money?
Often yes, and the saving is larger than people expect. Example 6 ran the same retrieval as Example 5 on a budget model and cut the monthly bill from 3,751 dollars to 176 dollars.
The practical answer is a router: send routine queries to the budget model and escalate only the hard ones. Even catching half your traffic on the cheap tier can halve your bill.
What top-k should I start with?
Begin at 3 to 5 for most document sets. This covers the answer for the majority of questions while keeping context tokens manageable.
Raise it only when you can name a question that failed for lack of context. A reranking pass lets you keep top-k low without sacrificing recall, since the best chunk is more likely to reach the top of a short list.
How accurate are the token estimates?
They are as accurate as your inputs. The formula itself is exact, but real chunk sizes and answer lengths vary from query to query.
Measure your production averages for a week and feed them back in. Real context sizes usually run higher than initial estimates because prompt templates and retrieved chunks both tend to grow over time.
Does caching change these numbers?
It can, significantly, when your provider supports prompt caching for a stable system prompt or repeated context. Cached input tokens are billed at a fraction of the standard input rate.
This calculator assumes no caching, so its figures are a conservative ceiling. If you cache a large fixed system prompt across many queries, your real input cost can fall well below the estimate here.
To model caching, use a dedicated caching calculator or reduce the base prompt tokens to reflect the discounted cached portion.
Why do output tokens cost more per million than input tokens?
Generating text is more computationally demanding than reading it, so providers price output higher, commonly two to five times the input rate.
This is why a 300-token answer can cost as much as a 900-token chunk of input. Capping output length trims the most expensive tokens in the pipeline.
Can I use this for a self-hosted model?
Indirectly. Enter your effective cost per million tokens based on your hardware and throughput, rather than a provider’s published rate.
The token arithmetic is identical whether you rent an API or run your own GPU. Only the price inputs change, so a self-hosting break-even calculator pairs well with this one.
Disclaimer
This calculator produces educational estimates for planning. The numbers it returns depend on the model you choose, your provider, your hardware, your configuration, and current pricing, all of which change frequently and without notice.
Cost outputs are not financial or business advice. They are a tool for reasoning about token economics, not a quote, and they exclude costs the tool does not model, including vector database hosting, reranking services, and infrastructure.
Verify every rate against your provider’s current pricing page and against your own measured token usage before you rely on any figure here. The live calculator on this page is the source of truth for its exact fields and defaults, so use it directly rather than reproducing its inputs from memory.
Run a small real test at low volume before scaling spend. A week of production logs will tell you more about your true cost per query than any estimate, and it will surface the gap between planned and actual token counts while that gap is still cheap to fix.








Running local RAG pipelines is a different beast regarding hardware. When you push 5 chunks of 400 tokens through a 70B parameter model at FP16, you hit a memory wall fast. My dual RTX 3090 setup handles the context window okay for now, but latency jumps to 2 tokens/sec once I hit the KV cache limit. Using 4-bit quantization helps, but retrieval speed usually bottlenecks at the embedding step if I am not caching vectors in RAM. How does this calculator account for the overhead of local quantization versus API-based inference pricing?
Regarding your question on quantization and API overhead: this calculator focuses on the per-token billing model used by providers like OpenAI or Anthropic, where you pay for input context regardless of your local compression. If you are running 70B models locally on your dual 3090s, your primary cost is electricity and hardware depreciation rather than per-query pricing. For local setups, I suggest tracking your VRAM usage via nvidia-smi and calculating your cost based on the power draw and the lifespan of your GPUs. The embedding overhead you mentioned is indeed significant, especially if you move beyond simple FAISS indexing to something like Qdrant or Milvus.
Thanks for the clarification. I actually found that offloading the embedding model to a smaller dedicated card keeps the main GPUs free for the generation logic, which helped my throughput significantly.
That is a smart architectural choice. By decoupling the embedding retrieval from the primary inference engine, you essentially parallelize the pipeline and prevent the vector search from causing a stutter in token generation. It is a common pattern for production-grade RAG systems looking to keep latency under 500ms.
I am tired of hidden costs everywhere. Is there a way to calculate this without giving my email to another SaaS tool? I just want a simple spreadsheet template. Paying for every token added to my prompt feels like a trap for small projects.
Pricing transparency is the main reason we built this. You do not need to provide any personal information to use the tool, and we do not store your inputs. It is designed to run entirely in your browser. If you prefer a manual approach, you can take the formula: (Total Tokens * Price per Million / 1,000,000) and drop it into a Google Sheet. The tool just automates the math so you can see how increasing your top-k parameter quickly inflates your monthly bill.