Pricecomparison.cloud · ai-api

LLM API: price per million tokens

Language models bill per token: separately for input (prompt + context) and output (model response). Output is typically 2–5× more expensive than input. We compare flagship and budget models. Closed (OpenAI, Anthropic, Mistral) vs. open model hosts (Groq, Together, Fireworks, DeepSeek) — where the same open model can cost very different amounts.

Prices verified models
Adjust monthly usage with sliders — $/mo column calculates cost (input Mtok × input $/1M + output Mtok × output $/1M) · click headers to sort · bar = combined price (1M input + 1M output)
5 Mtok/mo
2 Mtok/mo

Provider $/mo at selected usage Model Type Tier Intelligence (AA) Input $/1M Output $/1M Combined (1M+1M) Notes
Loading…

How to choose?

Cheapest are open model hosts and DeepSeek (V4 Flash $0.14/$0.28) and Mistral Small — a fraction of flagship prices. The same open model (e.g. Llama 70B) costs different amounts on different hosts: Groq is fastest (LPU), Fireworks and Together offer a wide catalog. DeepSeek and Mistral are strong if you want a European or budget option.

Flagships (GPT-5.6, Claude Opus 4.8) cost 5–6× more for output but are strongest in demanding reasoning. Real price depends on input:output ratio: if you send lots of context and get a short answer, input price matters; in chat output weighs more. "Combined" column assumes 1:1 ratio — calculate with your own ratio. On closed models prompt caching significantly lowers repeated context cost.

Intelligence column is the Artificial Analysis Intelligence Index (v4.1, 0–100) — an independent quality metric so cheapest does not automatically look best. Best value (intelligence per combined price) is currently DeepSeek V4 Flash: near-flagship level at a fraction of the price. Note: reasoning and non-reasoning mode scores are not fully comparable (mode noted in comments), and Mistral's exact version was not found in the index. Index changes — check the latest figure via the link.