Every system design conversation I have ever sat in eventually reaches the same moment. Someone sketches the boxes, someone else says "roughly what does that cost", and the room goes quiet. For twenty years we had an answer for that moment: Jeff Dean's latency numbers. L1 cache, main memory, disk seek, transatlantic round trip. Memorise six of them and you can size almost anything.
That table still works. It just no longer answers the question anyone is actually asking, because the expensive part of a modern feature is not a disk seek. It is a token.
Back-of-the-envelope estimation
is the practice of sizing a system from a handful of memorised constants before you build it, so the design conversation happens against real orders of magnitude instead of instinct.
So this is the replacement table. Every figure below is external, published, and dated, because I would rather you check my sources than trust me. Nothing here is a measurement I took.
Every number in one place
Here is the whole set. Sources and dates are in the last section, and every derived figure says so.
Text and tokens, as documented in July 2026
| Quantity | Number | Basis |
|---|---|---|
| One token, English prose (OpenAI) | ~4 characters, ~0.75 words | OpenAI API documentation |
| One page of prose | ~800 tokens | OpenAI's pages-per-dollar embedding pricing |
| Claude Opus 5, 1M token window | ~555,000 words, ~2.5M characters | Anthropic models overview |
| Claude Opus 4.6, 1M token window | ~750,000 words, ~3.4M characters | Anthropic models overview |
| Tokenizer change, Opus 4.6 to 4.7 and later | ~30% more tokens for identical text | Anthropic models overview |
Price per million tokens, as published on the vendors' pricing pages and rechecked on 26 July 2026
| Model | Input | Output | Cached input |
|---|---|---|---|
| Claude Fable 5 | $10.00 | $50.00 | $1.00 (derived, 0.1x) |
| Claude Opus 5 | $5.00 | $25.00 | $0.50 (derived, 0.1x) |
| GPT-5.6 Sol | $5.00 | $30.00 | $0.50 |
| Claude Sonnet 5 | $3.00 / $2.00 intro | $15.00 / $10.00 intro | $0.30 / $0.20 intro (0.1x) |
| GPT-5.6 Terra | $2.50 | $15.00 | $0.25 |
| Gemini 3.1 Pro Preview (up to 200k in) | $2.00 | $12.00 | $0.20 cache, plus storage |
| Claude Haiku 4.5 | $1.00 | $5.00 | $0.10 (derived, 0.1x) |
| GPT-5.6 Luna | $1.00 | $6.00 | $0.10 |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | $0.03 cache, plus storage |
| text-embedding-3-small | $0.02 (derived) | n/a | n/a |
| text-embedding-3-large | $0.13 (derived) | n/a | n/a |
One row in that table has a clock on it. I first listed Claude Sonnet 5 at its standard $3.00 and $15.00; rechecking Anthropic's pricing page on 26 July 2026, introductory pricing of $2.00 and $10.00 is in effect through 31 August 2026, with the standard rate taking over on 1 September. Everything I model below uses the standard rate, because that is the number still true in October.
Three multipliers do most of the work and are stable across vendors. Batch processing is 50% off at Anthropic, OpenAI and Google. A cache read is 0.1x a fresh input token. A cache write is 1.25x, which OpenAI states outright as $6.25 against its $5.00 base for GPT-5.6 Sol, and which Anthropic documents as 1.25x for a five minute TTL and 2x for an hour.
The embedding prices are the only derived ones in that table. OpenAI publishes them as pages per dollar — 62,500 for the small model, 9,615 for the large — on an 800 tokens-per-page basis. Do the division and you get two cents and thirteen cents per million.
Limits, as published in July 2026
| Model | Context window | Max output |
|---|---|---|
| Claude Fable 5, Opus 5, Sonnet 5 | 1M tokens | 128k |
| Claude Haiku 4.5 | 200k tokens | 64k |
| GPT-5.6 Sol, Terra, Luna | 1.05M tokens | 128k |
| Gemini 3.1 Pro Preview | 1,048,576 tokens | 65,536 |
| OpenAI embedding models, per call | 8,192 tokens | n/a |
Long context is not always priced flat. Google charges $2.00 input for Gemini 3.1 Pro Preview up to 200,000 tokens and $4.00 above it, with output moving from $12.00 to $18.00. If your design plan is "just put the whole corpus in the prompt", price the tier you will actually land in.
Embeddings, storage and latency, as documented in July 2026
| Quantity | Number | Basis |
|---|---|---|
| text-embedding-3-small dimensions | 1,536 | OpenAI docs |
| text-embedding-3-large dimensions | 3,072 | OpenAI docs |
| Bytes per vector in pgvector | 4 x dimensions + 8 | pgvector README |
| A 1,536 dimension vector on disk | 6,152 bytes | derived |
| RAM to hold vectors in Qdrant | vectors x dims x 4 bytes x 1.5 | Qdrant capacity planning |
| pgvector HNSW dimension ceiling | 2,000 | pgvector README |
| pgvector HNSW defaults | m = 16, ef_construction = 64 | pgvector README |
| Top-ten ANN query, 10M x 384 dims | 13.1 ms p99 | AWS pgvector 0.8.0 benchmark |
| Same query, category-filtered | 85.7 ms p99 | same |
| Same query returning 10,000 rows | 160.3 ms p99 | same |
| Fastest time to first token | 0.35 s | Artificial Analysis, July 2026 |
| Fastest output speed | 901.6 tokens/sec | Artificial Analysis, July 2026 |
| One A10G GPU-hour on AWS g5.xlarge | $1.006, or $0.00028 per GPU-second | published on-demand rate |
Why does the Jeff Dean latency table no longer size an AI feature?
Because every entry in it lives between a nanosecond and a few hundred milliseconds, and the thing you are now sizing has a floor around 350 milliseconds before it does any work. Dean's table puts a main memory reference at 100 ns and a round trip inside the same datacenter at 500,000 ns. The fastest model on Artificial Analysis's leaderboard in July 2026 took 0.35 seconds to produce its first token. That is roughly 700 times a same-datacenter round trip, and that is the fast one.
The table is not wrong. It is answering a question that has stopped being the bottleneck.
Two things changed. First, the dominant cost moved from time to money: an AI feature's constraint is usually the bill, not the clock. Second, the units moved. Dean's numbers are denominated in bytes moved. The new ones are denominated in tokens billed, vectors stored, and seconds a user waits for a stream to start.
That reframing has a practical consequence. A retrieval step you spent a sprint optimising from 40 ms to 12 ms is invisible next to a 350 ms first token. Whereas a system prompt you trimmed by 800 tokens shows up on every request, forever.
How much does a chat feature cost at 100,000 monthly users?
Between $7,900 and $208,000 a month, entirely depending on which model you pick — and the arithmetic that gets you there is four multiplications. Here are the assumptions, stated so you can disagree with them: 100,000 monthly active users, eight conversations each, six turns per conversation, a 1,500 token system prompt, 60 tokens per user message, 350 tokens per reply, and the full history resent every turn.
That last assumption is the one people forget. Input grows on every turn because you resend everything. Turn one sends 1,560 tokens; turn six sends 3,610. Across six turns that is 15,510 input tokens and 2,100 output tokens per conversation.
At 800,000 conversations a month, you are buying 12,408 million input tokens and 1,680 million output tokens.
| Model | Input $/M | Output $/M | Monthly bill | Per active user |
|---|---|---|---|---|
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | $7,922 | $0.08 |
| Claude Haiku 4.5 | $1.00 | $5.00 | $20,808 | $0.21 |
| GPT-5.6 Luna | $1.00 | $6.00 | $22,488 | $0.22 |
| Claude Sonnet 5 | $3.00 | $15.00 | $62,424 | $0.62 |
| Claude Opus 5 | $5.00 | $25.00 | $104,040 | $1.04 |
| Claude Fable 5 | $10.00 | $50.00 | $208,080 | $2.08 |
Those rows use standard list prices. On Claude Sonnet 5's introductory rate of $2.00 and $10.00, in effect through 31 August 2026, the same traffic costs $41,616 a month, or $0.42 per active user, until it reverts on 1 September. Budget against the standard rate and treat the discount as headroom.
Twenty-six times, same feature, same traffic. This is the single strongest argument for routing by request complexity rather than picking one model — which is most of why Codelit routes every generation across a chain of providers instead of standing on one.
Now look at where the money goes at the Sonnet 5 row. Output is 1,680 million of the 14,088 million tokens — 12% of the volume — and $25,200 of the $62,424 bill, which is 40% of the cost. Capping reply length is worth roughly three times as much as trimming the same number of prompt tokens.
Cache the system prompt and the picture moves again. Without caching, that 1,500 token prompt is resent six times per conversation, so it accounts for 9,000 of the 15,510 input tokens — 58% of your input spend on a string that never changes. Cached, it costs one write at 1.25x plus five reads at 0.1x, or 2,625 token-equivalents instead of 9,000. The Sonnet 5 bill drops from $62,424 to $47,124.
For a sanity check against something real rather than modelled: Codelit's actual AI spend runs about $0.30 per active user per month across diagram generation, which sits between the Haiku and Sonnet rows above. Different workload, same order of magnitude. That is what a good estimate is supposed to do.
How much RAM does semantic search over five million documents need?
About 369 GB, which is the number that kills most naive RAG designs. Assume five million documents averaging four pages, chunked at 500 tokens with 20% overlap. That is roughly eight chunks per document, so 40 million vectors — and the vector count, not the document count, is what you have to size.
The embedding bill is the cheap part and almost nobody's problem. Five million documents at 3,200 tokens each, plus overlap, is 19,200 million tokens: $384 on text-embedding-3-small, $2,496 on text-embedding-3-large. Re-embedding your entire corpus after a model change is a long afternoon and a rounding error.
Storage is the expensive part.
| Configuration | Vectors | Dims | pgvector on disk | Qdrant RAM formula |
|---|---|---|---|---|
| One chunk per document | 5M | 1,536 | 31 GB | 46 GB |
| Eight chunks per document | 40M | 1,536 | 246 GB | 369 GB |
| Same, reduced dimensions | 40M | 512 | 82 GB | 123 GB |
| Same, text-embedding-3-large | 40M | 3,072 | 492 GB | 737 GB |
Three decisions fall straight out of that table. Chunking strategy is a capacity decision, because eight-way chunking is an eight-times storage multiplier. Dimension count is a capacity decision, because OpenAI lets you request fewer dimensions from the same model and 512 dimensions costs a third of 1,536. And the model choice is a capacity decision, because the large model is double the storage before you have evaluated whether the extra couple of points on MTEB matter for your corpus.
There is also a hard wall worth knowing about. pgvector's HNSW index accepts up to 2,000 dimensions. A 3,072 dimension embedding cannot be HNSW-indexed in Postgres at all without dropping to half-precision vectors or reducing dimensions first. That is a schema decision you want to make before the backfill, not after.
Index builds have their own memory story: pgvector's docs say the build is significantly faster when the graph fits in maintenance_work_mem, and Postgres emits a notice the moment it stops fitting. At 40 million vectors it will not fit, and you should plan for that rather than discover it at 2 a.m.
What does an agent averaging twelve tool calls actually cost?
About 68 cents a run on Claude Opus 5, which is twenty times what the same task costs as a single request — and the reason is that an agent resends its entire transcript on every step. Assume a 4,000 token system prompt including tool definitions, a 200 token task, twelve tool calls where the model emits 120 tokens and the tool returns 800, and a 500 token final answer.
Thirteen requests. The first sends 4,200 tokens. The last sends 15,240. Total input across the run is 126,360 tokens against just 1,940 output tokens.
| Shape | Input tokens | Output tokens | Cost on Claude Opus 5 |
|---|---|---|---|
| Single request, no tools | 4,200 | 500 | $0.03 |
| Twelve tool calls, no caching | 126,360 | 1,940 | $0.68 |
| Twelve tool calls, prefix cached | 30,162 equivalent | 1,940 | $0.20 |
Only 15,240 of those 126,360 input tokens are ever new. The other 111,120 are the same bytes, resent. Price them as cache reads at 0.1x and pay 1.25x once for each new segment, and the input side collapses to about 30,000 token-equivalents. At ten thousand agent runs a day, that is roughly $204,000 a month becoming roughly $60,000.
Caching is the whole ballgame for agents, and it is fragile in a specific way: it is a prefix match. Anything that changes early in the prompt invalidates everything after it. A timestamp in the system prompt, a tool list assembled from an unordered set, a per-user string interpolated into the preamble — each of those quietly turns a 71% discount into zero, with no error and no warning. Freeze the prefix and put everything volatile at the end.
Which rules of thumb actually change an architecture decision?
Three, and they are all consequences of the tables above rather than restatements of them.
Price the reply, not the prompt. Output runs five to six times input on every frontier model, so a max output cap and an instruction to be concise beat a system prompt diet by roughly three to one. This is the cheapest change on the list and the one teams reach for last.
Design the prompt around the cache boundary, not around readability. A cache read costs a tenth of a fresh token, which makes prompt layout an architecture concern. Stable content first, volatile content last, deterministic serialisation everywhere. It is the same lesson I kept running into across fifty-five real-world architectures, where caching, not orchestration, was doing the actual scaling work.
Treat embedding dimensions as a storage budget. Dimensions multiply through storage, index RAM, shard count and node cost simultaneously. Halving them is a one-line API change; adding 246 GB of RAM to a cluster is a project. Reduce dimensions and measure recall before you shard.
There is a fourth that is less a rule than a warning. Tokenizers are per-vendor constants and they move. Anthropic's own documentation says the tokenizer introduced with Opus 4.7 produces roughly 30% more tokens for identical text than the one before it. A model upgrade with no code change can raise your bill by a third and shrink your effective context at the same time. Re-baseline with the vendor's token counting endpoint after every migration.
What these numbers do not tell you
They do not tell you which model is cheapest for your task. Cheapest per token and cheapest per solved problem are different quantities, and they frequently point in opposite directions. An agent that needs three attempts on a $1 model costs more than one attempt on a $5 model, and it also costs three times the latency. Cost per successful outcome is the metric that matters, and none of these tables can give it to you.
A failed generation is a paid generation, too. The pipeline that turns a prompt into a Codelit diagram validates structure before it treats a response as finished, and that check catches around 3% of responses that arrived with a clean 200 and content that would have rendered as a broken diagram. Every one of those was billed at full rate for output nobody could use. Whatever your equivalent failure rate is, it belongs in the denominator.
They also do not tell you about rate limits, which are often the real constraint long before price is. Nor about the quality cliff at long context: a model advertising a million tokens is telling you what it will accept, not what it will reason over reliably.
And they go stale. Every price above is dated, and the dates are the important part. I built a design-time cost estimator into Codelit precisely because pricing tables rot, and even with a weekly refresh the honest claim is directional rather than exact — accurate enough to compare two designs, not accurate enough to reconcile an invoice. Treat this page the same way.
Where every figure came from
| Figure | Source | Date basis |
|---|---|---|
| Characters and words per token | OpenAI API documentation | Rechecked 26 July 2026 |
| 800 tokens per page, pages per dollar | OpenAI embeddings guide | Rechecked 26 July 2026 |
| Claude model prices, context, max output | Anthropic models overview and pricing page | Rechecked 26 July 2026 |
| Claude Sonnet 5 introductory pricing through 31 August 2026 | Anthropic pricing page | Rechecked 26 July 2026 |
| Words and characters per 1M token window, tokenizer change | Anthropic models overview | Rechecked 26 July 2026 |
| Cache read 0.1x, write 1.25x and 2x, per-model cache minimums | Anthropic prompt caching docs | Rechecked 26 July 2026 |
| Anthropic batch discount of 50% | Anthropic batch processing docs | Rechecked 26 July 2026 |
| GPT-5.6 prices, cached input, cache writes, batch | OpenAI pricing page | Rechecked 26 July 2026 |
| GPT-5.6 context and output limits | OpenAI models page | Rechecked 26 July 2026 |
| Gemini prices, context caching, long-context tiers | Google Gemini API pricing | Page dated 21 July 2026 |
| Embedding dimensions and 8,192 token input limit | OpenAI embeddings guide | Rechecked 26 July 2026 |
| Bytes per vector, HNSW 2,000 dimension ceiling, m and ef_construction defaults | pgvector README | Rechecked 26 July 2026 |
| Vector RAM formula and 1.5x overhead | Qdrant capacity planning docs | Rechecked 26 July 2026 |
| ANN query p99 latencies on 10M x 384 dimensions | AWS Database Blog, pgvector 0.8.0 on Aurora | Published 28 May 2025 |
| Time to first token and output speed leaders, measurement method | Artificial Analysis models leaderboard and methodology | Rechecked 26 July 2026 |
| Jeff Dean latency numbers | Widely republished 2012 compilation of Dean's figures | Original circa 2000s |
| g5.xlarge on-demand hourly rate, A10G GPU | AWS EC2 G5 instance specs and published on-demand pricing | Rechecked 26 July 2026 |
Everything not in that table is arithmetic performed on something that is, and I have shown the arithmetic so you can check it. If a number here is wrong, it is wrong because the source moved or because I divided badly, and both are things you can catch in ten minutes.
That is the whole point of a table like this. Not to be authoritative — to be checkable, fast, before anyone writes code.
Questions people actually ask
- How many tokens is a page of text?
- About 800, which is the figure OpenAI uses when it quotes embedding prices in pages per dollar. The underlying rule of thumb in OpenAI's API documentation is one token per four characters, or roughly three tokens per four words of English. Anthropic publishes a different ratio for its own tokenizer, describing a one million token Claude Opus 5 window as around 555,000 words, which is closer to two tokens per word. Measure with the vendor's own tokenizer before you commit to a number.
- How do I estimate what an AI chat feature will cost per user?
- Multiply four things you can write down in a minute. Conversations per user per month, turns per conversation, tokens sent per turn including the full history you resend, and tokens generated per reply. Price input and output separately, because output tokens cost five to six times more than input tokens on every frontier model. A six turn chat at 100,000 monthly active users lands somewhere between eight cents and two dollars per user per month depending purely on which model tier you pick.
- How much RAM does a vector database need per million embeddings?
- Qdrant's capacity planning documentation gives a formula. Number of vectors times dimensions times four bytes times 1.5 for overhead. At 1,536 dimensions that is about 9.2 kilobytes per vector, so one million embeddings needs roughly 9 gigabytes and five million needs about 46 gigabytes. Postgres with pgvector stores each vector as four bytes per dimension plus eight, which is 6,152 bytes at 1,536 dimensions. Reducing dimensions is the cheapest lever you have.
- Does prompt caching actually save money on an AI feature?
- Yes, and it is usually the largest single lever. Anthropic prices a cache read at one tenth of a normal input token and a five minute cache write at 1.25 times. OpenAI's pricing page shows the same shape, listing cached input at fifty cents and cache writes at six dollars twenty five against a five dollar base for its flagship model. For an agent that resends its whole transcript on every tool call, caching cut a modelled run from about 68 cents to about 20 cents.
- Why is time to first token so much slower than a database query?
- Because they are different orders of magnitude and always will be. The fastest model on Artificial Analysis's leaderboard in July 2026 answered in 0.35 seconds to first token, while a top ten vector search over ten million embeddings measured 13.1 milliseconds at the ninety ninth percentile in Amazon's own pgvector benchmark. The model call is roughly twenty seven times slower than the retrieval step, which means retrieval tuning almost never shows up in what a user feels.
- Should I use 1536 or 3072 dimensional embeddings?
- Start at 1,536 and only go higher with evidence. Doubling dimensions doubles storage and index memory for a few points of benchmark score, and pgvector's HNSW index will not accept vectors above 2,000 dimensions at all, so a 3,072 dimensional model needs half precision vectors or dimension reduction before Postgres can index it. OpenAI's embedding models let you request fewer dimensions directly, which is usually the better trade.