Mo Sharif
← ~/blog

Numbers Every Engineer Should Know, Rewritten for AI

Every system design conversation I have ever sat in eventually reaches the same moment. Someone sketches the boxes, someone else says "roughly what does that cost", and the room goes quiet. For twenty years we had an answer for that moment: Jeff Dean's latency numbers. L1 cache, main memory, disk seek, transatlantic round trip. Memorise six of them and you can size almost anything.

That table still works. It just no longer answers the question anyone is actually asking, because the expensive part of a modern feature is not a disk seek. It is a token.

Back-of-the-envelope estimation

is the practice of sizing a system from a handful of memorised constants before you build it, so the design conversation happens against real orders of magnitude instead of instinct.

So this is the replacement table. Every figure below is external, published, and dated, because I would rather you check my sources than trust me. Nothing here is a measurement I took.

Here is the whole set. Sources and dates are in the last section, and every derived figure says so.

Text and tokens, as documented in July 2026

QuantityNumberBasis
One token, English prose (OpenAI)~4 characters, ~0.75 wordsOpenAI API documentation
One page of prose~800 tokensOpenAI's pages-per-dollar embedding pricing
Claude Opus 5, 1M token window~555,000 words, ~2.5M charactersAnthropic models overview
Claude Opus 4.6, 1M token window~750,000 words, ~3.4M charactersAnthropic models overview
Tokenizer change, Opus 4.6 to 4.7 and later~30% more tokens for identical textAnthropic models overview

Price per million tokens, as published on the vendors' pricing pages and rechecked on 26 July 2026

ModelInputOutputCached input
Claude Fable 5$10.00$50.00$1.00 (derived, 0.1x)
Claude Opus 5$5.00$25.00$0.50 (derived, 0.1x)
GPT-5.6 Sol$5.00$30.00$0.50
Claude Sonnet 5$3.00 / $2.00 intro$15.00 / $10.00 intro$0.30 / $0.20 intro (0.1x)
GPT-5.6 Terra$2.50$15.00$0.25
Gemini 3.1 Pro Preview (up to 200k in)$2.00$12.00$0.20 cache, plus storage
Claude Haiku 4.5$1.00$5.00$0.10 (derived, 0.1x)
GPT-5.6 Luna$1.00$6.00$0.10
Gemini 3.5 Flash-Lite$0.30$2.50$0.03 cache, plus storage
text-embedding-3-small$0.02 (derived)n/an/a
text-embedding-3-large$0.13 (derived)n/an/a

One row in that table has a clock on it. I first listed Claude Sonnet 5 at its standard $3.00 and $15.00; rechecking Anthropic's pricing page on 26 July 2026, introductory pricing of $2.00 and $10.00 is in effect through 31 August 2026, with the standard rate taking over on 1 September. Everything I model below uses the standard rate, because that is the number still true in October.

Three multipliers do most of the work and are stable across vendors. Batch processing is 50% off at Anthropic, OpenAI and Google. A cache read is 0.1x a fresh input token. A cache write is 1.25x, which OpenAI states outright as $6.25 against its $5.00 base for GPT-5.6 Sol, and which Anthropic documents as 1.25x for a five minute TTL and 2x for an hour.

The embedding prices are the only derived ones in that table. OpenAI publishes them as pages per dollar — 62,500 for the small model, 9,615 for the large — on an 800 tokens-per-page basis. Do the division and you get two cents and thirteen cents per million.

Limits, as published in July 2026

ModelContext windowMax output
Claude Fable 5, Opus 5, Sonnet 51M tokens128k
Claude Haiku 4.5200k tokens64k
GPT-5.6 Sol, Terra, Luna1.05M tokens128k
Gemini 3.1 Pro Preview1,048,576 tokens65,536
OpenAI embedding models, per call8,192 tokensn/a

Long context is not always priced flat. Google charges $2.00 input for Gemini 3.1 Pro Preview up to 200,000 tokens and $4.00 above it, with output moving from $12.00 to $18.00. If your design plan is "just put the whole corpus in the prompt", price the tier you will actually land in.

Embeddings, storage and latency, as documented in July 2026

QuantityNumberBasis
text-embedding-3-small dimensions1,536OpenAI docs
text-embedding-3-large dimensions3,072OpenAI docs
Bytes per vector in pgvector4 x dimensions + 8pgvector README
A 1,536 dimension vector on disk6,152 bytesderived
RAM to hold vectors in Qdrantvectors x dims x 4 bytes x 1.5Qdrant capacity planning
pgvector HNSW dimension ceiling2,000pgvector README
pgvector HNSW defaultsm = 16, ef_construction = 64pgvector README
Top-ten ANN query, 10M x 384 dims13.1 ms p99AWS pgvector 0.8.0 benchmark
Same query, category-filtered85.7 ms p99same
Same query returning 10,000 rows160.3 ms p99same
Fastest time to first token0.35 sArtificial Analysis, July 2026
Fastest output speed901.6 tokens/secArtificial Analysis, July 2026
One A10G GPU-hour on AWS g5.xlarge$1.006, or $0.00028 per GPU-secondpublished on-demand rate

Because every entry in it lives between a nanosecond and a few hundred milliseconds, and the thing you are now sizing has a floor around 350 milliseconds before it does any work. Dean's table puts a main memory reference at 100 ns and a round trip inside the same datacenter at 500,000 ns. The fastest model on Artificial Analysis's leaderboard in July 2026 took 0.35 seconds to produce its first token. That is roughly 700 times a same-datacenter round trip, and that is the fast one.

The table is not wrong. It is answering a question that has stopped being the bottleneck.

Two things changed. First, the dominant cost moved from time to money: an AI feature's constraint is usually the bill, not the clock. Second, the units moved. Dean's numbers are denominated in bytes moved. The new ones are denominated in tokens billed, vectors stored, and seconds a user waits for a stream to start.

That reframing has a practical consequence. A retrieval step you spent a sprint optimising from 40 ms to 12 ms is invisible next to a 350 ms first token. Whereas a system prompt you trimmed by 800 tokens shows up on every request, forever.

Between $7,900 and $208,000 a month, entirely depending on which model you pick — and the arithmetic that gets you there is four multiplications. Here are the assumptions, stated so you can disagree with them: 100,000 monthly active users, eight conversations each, six turns per conversation, a 1,500 token system prompt, 60 tokens per user message, 350 tokens per reply, and the full history resent every turn.

That last assumption is the one people forget. Input grows on every turn because you resend everything. Turn one sends 1,560 tokens; turn six sends 3,610. Across six turns that is 15,510 input tokens and 2,100 output tokens per conversation.

At 800,000 conversations a month, you are buying 12,408 million input tokens and 1,680 million output tokens.

ModelInput $/MOutput $/MMonthly billPer active user
Gemini 3.5 Flash-Lite$0.30$2.50$7,922$0.08
Claude Haiku 4.5$1.00$5.00$20,808$0.21
GPT-5.6 Luna$1.00$6.00$22,488$0.22
Claude Sonnet 5$3.00$15.00$62,424$0.62
Claude Opus 5$5.00$25.00$104,040$1.04
Claude Fable 5$10.00$50.00$208,080$2.08

Those rows use standard list prices. On Claude Sonnet 5's introductory rate of $2.00 and $10.00, in effect through 31 August 2026, the same traffic costs $41,616 a month, or $0.42 per active user, until it reverts on 1 September. Budget against the standard rate and treat the discount as headroom.

Twenty-six times, same feature, same traffic. This is the single strongest argument for routing by request complexity rather than picking one model — which is most of why Codelit routes every generation across a chain of providers instead of standing on one.

Now look at where the money goes at the Sonnet 5 row. Output is 1,680 million of the 14,088 million tokens — 12% of the volume — and $25,200 of the $62,424 bill, which is 40% of the cost. Capping reply length is worth roughly three times as much as trimming the same number of prompt tokens.

Cache the system prompt and the picture moves again. Without caching, that 1,500 token prompt is resent six times per conversation, so it accounts for 9,000 of the 15,510 input tokens — 58% of your input spend on a string that never changes. Cached, it costs one write at 1.25x plus five reads at 0.1x, or 2,625 token-equivalents instead of 9,000. The Sonnet 5 bill drops from $62,424 to $47,124.

For a sanity check against something real rather than modelled: Codelit's actual AI spend runs about $0.30 per active user per month across diagram generation, which sits between the Haiku and Sonnet rows above. Different workload, same order of magnitude. That is what a good estimate is supposed to do.

About 369 GB, which is the number that kills most naive RAG designs. Assume five million documents averaging four pages, chunked at 500 tokens with 20% overlap. That is roughly eight chunks per document, so 40 million vectors — and the vector count, not the document count, is what you have to size.

The embedding bill is the cheap part and almost nobody's problem. Five million documents at 3,200 tokens each, plus overlap, is 19,200 million tokens: $384 on text-embedding-3-small, $2,496 on text-embedding-3-large. Re-embedding your entire corpus after a model change is a long afternoon and a rounding error.

Storage is the expensive part.

ConfigurationVectorsDimspgvector on diskQdrant RAM formula
One chunk per document5M1,53631 GB46 GB
Eight chunks per document40M1,536246 GB369 GB
Same, reduced dimensions40M51282 GB123 GB
Same, text-embedding-3-large40M3,072492 GB737 GB

Three decisions fall straight out of that table. Chunking strategy is a capacity decision, because eight-way chunking is an eight-times storage multiplier. Dimension count is a capacity decision, because OpenAI lets you request fewer dimensions from the same model and 512 dimensions costs a third of 1,536. And the model choice is a capacity decision, because the large model is double the storage before you have evaluated whether the extra couple of points on MTEB matter for your corpus.

There is also a hard wall worth knowing about. pgvector's HNSW index accepts up to 2,000 dimensions. A 3,072 dimension embedding cannot be HNSW-indexed in Postgres at all without dropping to half-precision vectors or reducing dimensions first. That is a schema decision you want to make before the backfill, not after.

Index builds have their own memory story: pgvector's docs say the build is significantly faster when the graph fits in maintenance_work_mem, and Postgres emits a notice the moment it stops fitting. At 40 million vectors it will not fit, and you should plan for that rather than discover it at 2 a.m.

About 68 cents a run on Claude Opus 5, which is twenty times what the same task costs as a single request — and the reason is that an agent resends its entire transcript on every step. Assume a 4,000 token system prompt including tool definitions, a 200 token task, twelve tool calls where the model emits 120 tokens and the tool returns 800, and a 500 token final answer.

Thirteen requests. The first sends 4,200 tokens. The last sends 15,240. Total input across the run is 126,360 tokens against just 1,940 output tokens.

ShapeInput tokensOutput tokensCost on Claude Opus 5
Single request, no tools4,200500$0.03
Twelve tool calls, no caching126,3601,940$0.68
Twelve tool calls, prefix cached30,162 equivalent1,940$0.20

Only 15,240 of those 126,360 input tokens are ever new. The other 111,120 are the same bytes, resent. Price them as cache reads at 0.1x and pay 1.25x once for each new segment, and the input side collapses to about 30,000 token-equivalents. At ten thousand agent runs a day, that is roughly $204,000 a month becoming roughly $60,000.

Caching is the whole ballgame for agents, and it is fragile in a specific way: it is a prefix match. Anything that changes early in the prompt invalidates everything after it. A timestamp in the system prompt, a tool list assembled from an unordered set, a per-user string interpolated into the preamble — each of those quietly turns a 71% discount into zero, with no error and no warning. Freeze the prefix and put everything volatile at the end.

Three, and they are all consequences of the tables above rather than restatements of them.

Price the reply, not the prompt. Output runs five to six times input on every frontier model, so a max output cap and an instruction to be concise beat a system prompt diet by roughly three to one. This is the cheapest change on the list and the one teams reach for last.

Design the prompt around the cache boundary, not around readability. A cache read costs a tenth of a fresh token, which makes prompt layout an architecture concern. Stable content first, volatile content last, deterministic serialisation everywhere. It is the same lesson I kept running into across fifty-five real-world architectures, where caching, not orchestration, was doing the actual scaling work.

Treat embedding dimensions as a storage budget. Dimensions multiply through storage, index RAM, shard count and node cost simultaneously. Halving them is a one-line API change; adding 246 GB of RAM to a cluster is a project. Reduce dimensions and measure recall before you shard.

There is a fourth that is less a rule than a warning. Tokenizers are per-vendor constants and they move. Anthropic's own documentation says the tokenizer introduced with Opus 4.7 produces roughly 30% more tokens for identical text than the one before it. A model upgrade with no code change can raise your bill by a third and shrink your effective context at the same time. Re-baseline with the vendor's token counting endpoint after every migration.

They do not tell you which model is cheapest for your task. Cheapest per token and cheapest per solved problem are different quantities, and they frequently point in opposite directions. An agent that needs three attempts on a $1 model costs more than one attempt on a $5 model, and it also costs three times the latency. Cost per successful outcome is the metric that matters, and none of these tables can give it to you.

A failed generation is a paid generation, too. The pipeline that turns a prompt into a Codelit diagram validates structure before it treats a response as finished, and that check catches around 3% of responses that arrived with a clean 200 and content that would have rendered as a broken diagram. Every one of those was billed at full rate for output nobody could use. Whatever your equivalent failure rate is, it belongs in the denominator.

They also do not tell you about rate limits, which are often the real constraint long before price is. Nor about the quality cliff at long context: a model advertising a million tokens is telling you what it will accept, not what it will reason over reliably.

And they go stale. Every price above is dated, and the dates are the important part. I built a design-time cost estimator into Codelit precisely because pricing tables rot, and even with a weekly refresh the honest claim is directional rather than exact — accurate enough to compare two designs, not accurate enough to reconcile an invoice. Treat this page the same way.

FigureSourceDate basis
Characters and words per tokenOpenAI API documentationRechecked 26 July 2026
800 tokens per page, pages per dollarOpenAI embeddings guideRechecked 26 July 2026
Claude model prices, context, max outputAnthropic models overview and pricing pageRechecked 26 July 2026
Claude Sonnet 5 introductory pricing through 31 August 2026Anthropic pricing pageRechecked 26 July 2026
Words and characters per 1M token window, tokenizer changeAnthropic models overviewRechecked 26 July 2026
Cache read 0.1x, write 1.25x and 2x, per-model cache minimumsAnthropic prompt caching docsRechecked 26 July 2026
Anthropic batch discount of 50%Anthropic batch processing docsRechecked 26 July 2026
GPT-5.6 prices, cached input, cache writes, batchOpenAI pricing pageRechecked 26 July 2026
GPT-5.6 context and output limitsOpenAI models pageRechecked 26 July 2026
Gemini prices, context caching, long-context tiersGoogle Gemini API pricingPage dated 21 July 2026
Embedding dimensions and 8,192 token input limitOpenAI embeddings guideRechecked 26 July 2026
Bytes per vector, HNSW 2,000 dimension ceiling, m and ef_construction defaultspgvector READMERechecked 26 July 2026
Vector RAM formula and 1.5x overheadQdrant capacity planning docsRechecked 26 July 2026
ANN query p99 latencies on 10M x 384 dimensionsAWS Database Blog, pgvector 0.8.0 on AuroraPublished 28 May 2025
Time to first token and output speed leaders, measurement methodArtificial Analysis models leaderboard and methodologyRechecked 26 July 2026
Jeff Dean latency numbersWidely republished 2012 compilation of Dean's figuresOriginal circa 2000s
g5.xlarge on-demand hourly rate, A10G GPUAWS EC2 G5 instance specs and published on-demand pricingRechecked 26 July 2026

Everything not in that table is arithmetic performed on something that is, and I have shown the arithmetic so you can check it. If a number here is wrong, it is wrong because the source moved or because I divided badly, and both are things you can catch in ten minutes.

That is the whole point of a table like this. Not to be authoritative — to be checkable, fast, before anyone writes code.

Questions people actually ask

How many tokens is a page of text?
About 800, which is the figure OpenAI uses when it quotes embedding prices in pages per dollar. The underlying rule of thumb in OpenAI's API documentation is one token per four characters, or roughly three tokens per four words of English. Anthropic publishes a different ratio for its own tokenizer, describing a one million token Claude Opus 5 window as around 555,000 words, which is closer to two tokens per word. Measure with the vendor's own tokenizer before you commit to a number.
How do I estimate what an AI chat feature will cost per user?
Multiply four things you can write down in a minute. Conversations per user per month, turns per conversation, tokens sent per turn including the full history you resend, and tokens generated per reply. Price input and output separately, because output tokens cost five to six times more than input tokens on every frontier model. A six turn chat at 100,000 monthly active users lands somewhere between eight cents and two dollars per user per month depending purely on which model tier you pick.
How much RAM does a vector database need per million embeddings?
Qdrant's capacity planning documentation gives a formula. Number of vectors times dimensions times four bytes times 1.5 for overhead. At 1,536 dimensions that is about 9.2 kilobytes per vector, so one million embeddings needs roughly 9 gigabytes and five million needs about 46 gigabytes. Postgres with pgvector stores each vector as four bytes per dimension plus eight, which is 6,152 bytes at 1,536 dimensions. Reducing dimensions is the cheapest lever you have.
Does prompt caching actually save money on an AI feature?
Yes, and it is usually the largest single lever. Anthropic prices a cache read at one tenth of a normal input token and a five minute cache write at 1.25 times. OpenAI's pricing page shows the same shape, listing cached input at fifty cents and cache writes at six dollars twenty five against a five dollar base for its flagship model. For an agent that resends its whole transcript on every tool call, caching cut a modelled run from about 68 cents to about 20 cents.
Why is time to first token so much slower than a database query?
Because they are different orders of magnitude and always will be. The fastest model on Artificial Analysis's leaderboard in July 2026 answered in 0.35 seconds to first token, while a top ten vector search over ten million embeddings measured 13.1 milliseconds at the ninety ninth percentile in Amazon's own pgvector benchmark. The model call is roughly twenty seven times slower than the retrieval step, which means retrieval tuning almost never shows up in what a user feels.
Should I use 1536 or 3072 dimensional embeddings?
Start at 1,536 and only go higher with evidence. Doubling dimensions doubles storage and index memory for a few points of benchmark score, and pgvector's HNSW index will not accept vectors above 2,000 dimensions at all, so a 3,072 dimensional model needs half precision vectors or dimension reduction before Postgres can index it. OpenAI's embedding models let you request fewer dimensions directly, which is usually the better trade.