How to use this table
Shortlist several models, then run the same representative evaluation set against each one. Compare cost per successful task—not cost per raw request.
Normalize input, output and monthly volume, then sort by estimated list-price cost. Cost is not a quality score.
| Model | Per request | Monthly | vs cheapest | Applied pricing |
|---|
Shortlist several models, then run the same representative evaluation set against each one. Compare cost per successful task—not cost per raw request.
Output and reasoning tokens are frequently more expensive than input. A model that completes the task in fewer generated tokens can be cheaper despite a higher list price.