Costs not included by default
- Vector database storage and query operations.
- Reranking APIs, OCR, parsing and document extraction.
- Metadata filters, network egress and observability.
- Re-embedding caused by model migration.
Model ingestion frequency, chunk overlap, retrieved context and monthly query volume in one transparent workflow.
Use an exact corpus token count for a billing-grade estimate.
Reduce retrieved context only after measuring answer quality. Smaller chunks do not automatically mean lower total cost: overlap can materially increase ingestion tokens.