Understanding What an AI Model Actually Costs You
AI Productivity

Understanding What an AI Model Actually Costs You

6 September 20268 min read1523 words
Tags#token-costs#prompt-caching#ai-budgeting#cost-control

Most people meet AI pricing as a line on a rate card and a bill that does not match their expectations. The gap is almost always in how the work is counted, not in the price itself. By the end of this post you will know what a token is, why input, output and cached tokens cost different amounts, why an agent that works for ten minutes costs far more than a chat message, and how to estimate your own monthly spend from your actual usage instead of from a price list.

What a token actually is

A token is a chunk of text, usually a few characters long. Common words tend to be one token. Longer or unusual words get split into several. Punctuation, spaces and line breaks count too. Code tends to tokenise less efficiently than prose because of all the brackets, indentation and symbol characters.

A rough working rule for English prose is that a token is somewhat shorter than a word, so a page lands in the hundreds of tokens and a long document in the tens of thousands. Do not treat that as precise. Every model family tokenises differently, and most providers expose a token counter or return exact counts with each response. Use the real counts when money is involved.

The important consequence: you are not billed per question or per conversation. You are billed for text moving in and out, measured in these chunks.

Input, output and cached tokens are three different prices

Nearly every provider splits billing into at least two rates, and increasingly three.

Input tokens are everything you send: your question, the system prompt, the documents you attached, the conversation history, the tool definitions. All of it counts, every single time.

Output tokens are what the model generates. These are typically much more expensive per token than input, because generating text is the computationally heavy part. Reasoning models that think before answering can produce a lot of output tokens you never see in the final answer, and those usually bill as output.

Cached tokens are input you have sent before, that the provider kept in a ready state so it does not have to be processed from scratch. Reading from that cache is much cheaper than sending the same text fresh.

For a concrete set of published figures, Anthropic lists Claude Fable 5.1 at $10 per million input tokens, $50 per million output tokens, and $0.25 per million cache-read tokens. Output is five times input. Cache reads are a fortieth of a normal input token. Those ratios, more than the absolute numbers, are what should shape how you work. Note that these are one provider's published rates for one model at one point in time, and rates change. We are not quoting figures for any other provider here, because we do not have published ones to quote.

Why caching matters so much for repeated context

Think about what a typical serious AI task looks like. You load a long system prompt, a style guide, a schema, maybe a set of reference documents. Then you ask twenty questions against that same context. Without caching, that whole block is re-sent and re-billed as fresh input twenty times.

Work it through with the published Fable 5.1 rates. Say your fixed context is 100,000 tokens. At $10 per million, sending it once as fresh input is $1.00, so twenty uncached questions cost $20.00 for the context alone, before a single word of your actual questions or answers. Read those same tokens from cache at $0.25 per million and each read is $0.025. Nineteen cached reads plus one initial load is a different order of magnitude for the same work.

Two honest caveats. Most providers charge something to write to the cache in the first place, and that write rate is not in the figures we are working from, so treat the arithmetic above as the shape of the saving rather than a full quote. And caches expire: one that goes cold between batches saves you nothing, which is why running related work close together matters.

Anthropic reduced that cache-read rate from $1.00 to $0.25 per million, a 75% cut, and states that it measures roughly 25% lower costs on typical workloads and up to roughly 45% on agentic workloads as a result. Those are the maker's own figures on the maker's own workload mix, not a guarantee for yours. Why agentic work benefits most is worth understanding on its own.

Why multi-step and agentic work costs more

A single question is one round trip. You send some input, you get some output, you are done.

An agent doing a multi-step task does something structurally different. It sends the context, gets a step, runs a tool, then sends the context plus the previous step plus the tool result back again for the next decision. Repeat that for fifteen steps and the input grows on every single one. The conversation is not fifteen questions. It is fifteen increasingly long questions, each carrying everything that came before.

That is why a task that feels like one instruction to you can cost fifty times a chat message. It also explains why caching helps agents disproportionately: the growing history has a large stable prefix that is identical on every step, and a stable prefix is exactly what a cache is good at.

Practical implication: measure agent runs separately from chat usage. Averaging the two together produces a number that describes neither.

Estimating your bill from your usage, not from a rate card

A rate card cannot tell you what you will spend, because it does not know your token volumes. Work from the other end.

Start by listing your actual jobs. Not "we use AI for support", but "we summarise about 40 tickets a day" and "we run a report agent twice a week". Then run each job once for real and record the token counts the provider returns with each response. Most dashboards also show usage per key or per project.

Multiply out from the measured single run:

cost per run = (input tokens / 1,000,000 × input rate)
             + (cached tokens / 1,000,000 × cache-read rate)
             + (output tokens / 1,000,000 × output rate)

monthly cost = cost per run × runs per month

Do that per job type and add them up, then add a margin, because real usage includes retries, abandoned runs, people experimenting, and prompts that grow as someone keeps adding instructions to the system message. A measured estimate with a 30% to 50% cushion lands far closer than any guess from a price page.

Then set a budget alert on the account. Cost problems with AI are rarely a gradual drift. They are usually one loop that retried, or one prompt that started attaching a whole database dump.

Habits that actually reduce the bill

Send less context. The most common source of waste is attaching a whole document when three relevant paragraphs would do. Retrieving the right section costs a fraction of sending everything and hoping.

Put the stable part of the prompt first. Caches generally work on a matching prefix. Keep your system prompt, schema and reference material at the front, and whatever changes per request at the end. Reordering that block is often the highest-value change available.

Run related work together. Twenty questions against the same context in one session get the cache benefit. The same twenty spread across a week probably do not.

Cap output length. Output is the expensive side. Asking for a five-bullet summary instead of leaving length open changes the bill directly, and usually improves the answer.

Match the model to the task. Classifying whether an email is a complaint does not need the same model as reasoning through a legal document. Most providers offer smaller, cheaper models in the same family. Routing simple, high-volume tasks to the smaller one and reserving the large model for genuinely hard work is standard practice, and it is usually where the biggest savings sit.

Batch where latency does not matter. If a job runs overnight and nobody is waiting, check whether your provider offers a discounted batch or asynchronous mode. Terms here change, so read the current ones.

Trim conversation history. Long sessions re-send the entire history every turn. Summarising older turns into a short recap keeps input from growing without limit.

What to watch for

Prices and cache terms change, sometimes substantially, as the Fable 5.1 cache-read reduction shows. Re-check your rates every few months rather than trusting a number you wrote down once. Re-measure your usage after any significant prompt change too, because a system message that quietly doubled in length has doubled a cost on every request you make.

Where to go next

Stuck working out which of your jobs is driving the bill? Write to the desk with your usage numbers and we will help you read them.

Comments

No comments yet. Be the first to share your thoughts.

AI Token Costs Explained: Input, Output, Caching