| Provider | $/mo at selected usage ▼ | Model | Type | Tier | Intelligence (AA) | Input $/1M | Output $/1M | Combined (1M+1M) | Notes |
|---|---|---|---|---|---|---|---|---|---|
| Loading… | |||||||||
How to choose?
Cheapest are open model hosts and DeepSeek (V4 Flash $0.14/$0.28) and Mistral Small — a fraction of flagship prices. The same open model (e.g. Llama 70B) costs different amounts on different hosts: Groq is fastest (LPU), Fireworks and Together offer a wide catalog. DeepSeek and Mistral are strong if you want a European or budget option.
Flagships (GPT-5.6, Claude Opus 4.8) cost 5–6× more for output but are strongest in demanding reasoning. Real price depends on input:output ratio: if you send lots of context and get a short answer, input price matters; in chat output weighs more. "Combined" column assumes 1:1 ratio — calculate with your own ratio. On closed models prompt caching significantly lowers repeated context cost.
Intelligence column is the Artificial Analysis Intelligence Index (v4.1, 0–100) — an independent quality metric so cheapest does not automatically look best. Best value (intelligence per combined price) is currently DeepSeek V4 Flash: near-flagship level at a fraction of the price. Note: reasoning and non-reasoning mode scores are not fully comparable (mode noted in comments), and Mistral's exact version was not found in the index. Index changes — check the latest figure via the link.