Engineering
NOT ALL TOKENS ARE CREATED EQUAL
Most RAG pipelines quietly assume every token carries equal expected information value. That frequentist assumption compounds at every layer — and it shows up on the bill.
A typical RAG pipeline makes one statistical commitment that never appears in the architecture diagram: that tokens carry equal expected information value. Chunks are embedded uniformly, documents are loaded by recency or similarity rather than relevance weight, and re-embedding passes touch every record at flat per-token cost. That is the frequentist worldview applied to your infrastructure bill — and it is not what anyone actually believes about their data.
Where the assumption hides
- Ingestion. Fixed-size chunking treats paragraphs uniformly. A boilerplate disclaimer appearing in 200 documents gets 200 embeddings; a novel claim appearing once gets one. Both occupy the same storage and compete for the same retrieval slots.
- Migration. When the embedding model changes, the standard pattern re-embeds the whole corpus and charges by total tokens. A document not queried in eighteen months is migrated with the same urgency as one queried this morning.
- Retrieval. Top-k by similarity, stuffed into context. A reranker, where used, works over surface text without knowing that two chunks describe the same entity or that one document supersedes another.
- Inference. The bill scales with input tokens, all priced identically. The architecture has no way to express which tokens were predictably going to matter.
Each default is defensible alone. The problem is the composition: every layer multiplies the cost of treating the corpus as flat, and the multiplications compound.
The Bayesian alternative, in production
- Storage: deduplication and entity resolution. The same conceptual claim appearing in five documents becomes one entity with five provenance edges, embedded once. Vector storage typically drops 30–50% on enterprise corpora with cross-referenced content.
- Migration: posterior-weighted re-embedding. High query-weight documents migrate promptly; the long tail migrates lazily or not until a query forces it.
- Retrieval: graph traversal. Rather than asking the model to find what matters by surface similarity, hand it the entity, the relationships, the evidence chain and the contradictions, pre-resolved at ingest. Traversal stops being an optimisation on top of retrieval and becomes the retrieval — you retrieve a subgraph whose structure is already the answer to which tokens deserved attention.
This is where a graph-native engine with vector search in the same query matters in a way a pure vector store cannot: a two-hop traversal starting from a vector-similar entity and walking out along typed edges is one query, one transaction, one engine.
The arithmetic
A 50M-token corpus at $0.12 per million tokens costs $6.00 to re-embed flat. Posterior-weighted, with the long tail discounted to 0.03 of full weight, it costs roughly $1.41 — a 76% reduction. At platform scale across 25 billion tokens, that turns $3,000 of recurring cost per migration cycle into about $700.
At retrieval, on a model priced at $3 per million input tokens, a 200k-token bundle costs $0.60 per query before output; a 4k-token subgraph costs $0.012. With prompt caching at a realistic 40% hit rate, the bundle falls to about $0.24 and the subgraph to $0.008. Caching narrows the gap from 50× to 30×; it does not close it.
Why it is harder to copy than it looks
A vendor on a flat-token architecture cannot reach this position by tweaking one layer. Deduplicate at storage and you still pay flat re-embed costs. Add posterior-weighted migration and you still stuff bundles at query time. Swap in graph retrieval and you still maintain a separate vector index with its own migration logic. The advantages compound the same way the costs do: dedup makes traversal more efficient, traversal often makes the reranker unnecessary, skipping the reranker simplifies the audit trail, and the audit trail is what makes the system defensible to a regulator.
Three questions for your own architecture
- When a new embedding model arrives, is your migration cost a function of total tokens or of relevant tokens? The fix is not a cheaper embedding model; it is a posterior over which embeddings are worth keeping current.
- At retrieval time, is your system returning text or returning structure? The fix is not a better reranker; it is a graph that encodes the relationships your queries depend on.
- When a compliance officer asks what the system relied on, do you have a deterministic answer pointing at specific structural elements — or a 200,000-token bundle and the model's claim about what it used?
These are three angles on one fix: refuse the assumption that all data deserves equal treatment, and refuse it at every layer. Flat-token architectures will keep working until the bills make the case.
Adapted from "Not all tokens are created equal." by Jonathan Aiken, first published in Nodes & Edges, 27 April 2026.