Embedding cost
Embeddings are the cheapest per-token models on the market — which is exactly why teams forget to budget them. At search scale, query embeddings run every single day, and the refresh loop never stops.
The formula
corpus_tok = pages × words × 1.33
initial = corpus_tok/1M × E_in (one-time)
refresh = initial × refresh% (monthly)
queries = queries/day × qTok/1M × E_in × 30 (monthly)
total = queries + refresh + vectorDB
Embedding prices as of , pulled daily from OpenRouter's API. Chunking with overlap inflates indexed tokens ~10–15%; embeddings of the same corpus with a new model version are a full re-index — the refresh% input is where you budget for that migration risk.