AI cost stack

The infrastructure categories behind AI cost control.

A cost program needs more than a calculator. Use this map to design observability, routing, evaluation and financial ownership.

O

Observability & tracing

Capture model, tokens, cache usage, tool calls, latency, errors, user and feature attribution for every run.

R

Gateways & routing

Centralize provider access, fallback behavior, budgets, rate limits and cost-aware model selection.

E

Evaluation

Measure quality and task completion so optimizations target cost per successful outcome.

C

Caching

Identify stable prefixes, semantic reuse and provider-specific cache write, hit and storage economics.

D

Data & retrieval

Track parsing, embedding, vector storage, reranking and retrieved-context costs as one system.

F

FinOps ownership

Allocate cost to team, feature, tenant and customer; set budgets; flag anomalies; forecast unit economics.

Minimum event schema

  • Provider, model ID, snapshot and service tier.
  • Input, cached, cache-write, output and reasoning tokens.
  • Tool calls, retries, status, latency and trace ID.
  • Product feature, tenant, customer and environment.
  • Published price version and actual invoice amount.

Build or buy?

Use third-party observability where it speeds implementation, but keep a provider-neutral internal event schema. Your cost history and attribution model should remain portable.