LLM API pricing

What each model actually costs per million tokens.

Rates read off each provider’s own pricing documentation, stamped with the date they were checked, and cross-checked on every deploy against the live model catalog Spendline uses to price real API traffic. If the two ever disagree, this site fails to build rather than showing you a stale number.

Looking for what Spendline itself costs? See our plans.

By model

Claude Sonnet 5 pricing
$2 input, $10 output per million tokens. Anthropic.
Verified August 12, 2026
GPT-5.6 Terra pricing
$2 input, $12 output per million tokens. OpenAI.
Verified September 16, 2026
Claude Opus 5 pricing
$5 input, $25 output per million tokens. Anthropic.
Verified September 16, 2026
Claude Opus 4.8 pricing
$5 input, $25 output per million tokens. Anthropic.
Verified September 16, 2026
Claude Haiku 4.5 pricing
$1 input, $5 output per million tokens. Anthropic.
Verified September 16, 2026
Claude Fable 5.1 pricing
$10 input, $50 output per million tokens. Anthropic.
Verified September 16, 2026
GPT-6 Astra pricing
$10 input, $50 output per million tokens. OpenAI.
Verified September 16, 2026
GPT-5.6 Sol pricing
$4 input, $20 output per million tokens. OpenAI.
Verified September 16, 2026
GPT-5.6 Luna pricing
$0.20 input, $1.20 output per million tokens. OpenAI.
Verified September 16, 2026
Gemini 3.1 Pro Preview pricing
$2 input, $12 output per million tokens. Google Gemini.
Verified September 16, 2026
Gemini 3.8 Flash pricing
$0.75 input, $3.75 output per million tokens. Google Gemini.
Verified September 16, 2026
Gemini 2.5 Pro pricing
$1.25 input, $10 output per million tokens. Google Gemini.
Verified September 16, 2026
Grok 4.6 pricing
$2 input, $6 output per million tokens. xAI.
Verified September 16, 2026
Grok 4.5 pricing
$2 input, $6 output per million tokens. xAI.
Verified September 16, 2026
Mistral Large 3 pricing
$0.50 input, $1.50 output per million tokens. Mistral.
Verified September 16, 2026
Mistral Medium 3.5 pricing
$1.50 input, $7.50 output per million tokens. Mistral.
Verified September 16, 2026
DeepSeek V4 Pro pricing
$0.66 input, $1.98 output per million tokens. DeepSeek.
Verified September 16, 2026
DeepSeek Flash pricing
$0.15 input, $0.60 output per million tokens. DeepSeek.
Verified September 16, 2026
Qwen3.8 Max pricing
$2 input, $6 output per million tokens. Alibaba Qwen.
Verified September 16, 2026
Qwen3.7 Max pricing
$2.50 input, $7.50 output per million tokens. Alibaba Qwen.
Verified September 16, 2026
Kimi K3 (Together AI) pricing
$3 input, $15 output per million tokens. Together AI.
Verified September 16, 2026
DeepSeek V4 Pro 0813 (Together AI) pricing
$1.32 input, $3.96 output per million tokens. Together AI.
Verified September 16, 2026
Qwen3.8-2.4T-A95B (Together AI) pricing
$2 input, $6 output per million tokens. Together AI.
Verified September 16, 2026
GLM-5.3 (Together AI) pricing
$1.40 input, $4.40 output per million tokens. Together AI.
Verified September 16, 2026
Kimi K3 (Fireworks AI) pricing
$3 input, $15 output per million tokens. Fireworks AI.
Verified September 16, 2026
DeepSeek V4 Pro 0813 (Fireworks AI) pricing
$1.32 input, $3.96 output per million tokens. Fireworks AI.
Verified September 16, 2026
Qwen 3.8 Max (Fireworks AI) pricing
$2 input, $6 output per million tokens. Fireworks AI.
Verified September 16, 2026
GLM-5.3 (Fireworks AI) pricing
$1.40 input, $4.40 output per million tokens. Fireworks AI.
Verified September 16, 2026
GPT-OSS 120B (Groq) pricing
$0.15 input, $0.60 output per million tokens. Groq.
Verified September 16, 2026
GPT-OSS 20B (Groq) pricing
$0.075 input, $0.30 output per million tokens. Groq.
Verified September 16, 2026
Qwen3.8-27B (Groq) pricing
$0.80 input, $4 output per million tokens. Groq.
Verified September 16, 2026

Head to head

When these prices last moved

Providers change rates and the old number disappears. We keep a dated log of every price change we have observed, including the ones that have been announced but have not taken effect yet.

Why these numbers are not the whole bill

Every page here shows a rate card, and a rate card is the smallest part of what a team actually pays. Output usually costs several times input, so responses drive the invoice even when prompts look larger. Reasoning tokens bill as output, which means raising an effort setting raises the bill without changing a prompt. Retries and agent fan-out pay full price for calls that produced nothing. And a provider invoice arrives as one number for the entire organization, so none of it is attributable to a customer or a feature by default.

The guides go into that in depth: what Claude actually costs and why the token estimate comes in low, how to track LLM costs per customer, and LLM unit economics.

Rates on this page last verified September 16, 2026.