Mistral pricing profile

Mistral Small 4

Efficient hybrid instruct, reasoning and coding model.

activeeconomy256,000 contextVerified 2026-08-02

Pricing modes

USD per one million tokens. A documented long-context band applies to the full request when its threshold is crossed.

ModeInputCached inputCache writeOutputLong-context band
standard$0.15$0.6
batch$0.075$0.3
Worked example

100K input + 2K output

$0.0162

Estimated standard-tier token cost for one request. At 10,000 identical requests per month, the token subtotal would be $162.00.

Open this scenario →
Technical limits

Model constraints

Context window256,000 tokens
Maximum output64,000
Long-context thresholdNo separate band recorded
Tokenizer profilemistral

What can change the invoice

  • No separate cached-input rate is recorded for the active pricing schedule.
  • Taxes, cloud marketplaces, enterprise contracts, retries and quality differences are not included in this example.

Related Mistral models

Compare models from the same provider before evaluating quality and latency.