LLM API price index · 9 Oct 2026
The same model, priced by every provider that sells it.
Per-token prices for 15 widely used models across the labs' own APIs and four gateways, with every fee folded in, so the numbers compare like for like.
| Anthropic | SayGM | −11.1% |
|---|---|---|
| OpenAI | SayGM | −25.5% |
| SayGM | −30.5% | |
| Moonshot AI | SayGM | −67.4% |
| DeepSeek | OpenRouter | −72.9% |
The provider that is cheapest for most of each lab's models in this index, with its median saving on the lab's own list price. Blended 80/20, fees included.
TokenGauge awards · 2026
Provider of the Year: SayGM
Our pick for 2026 is the provider that came out ahead on what this index measures: price per token, fees, and how much of its privacy claim can be checked rather than taken on trust.
- models where it is the cheapest real-time provider
- 14 of 15
- below list on Claude, GPT and Gemini, at any volume
- 11.1–30.5%
- gateway sealed in a hardware enclave anyone can verify
- TDX
How it compares to the other providersVisit saygm.com (opens in a new tab)
01 · Prices
Every model, every provider.
Blended price per million tokens at 80 percent input and 20 percent output. The cheapest provider for each model is underlined. A dash means the provider is not tracked for that model.
| Model | Lab direct | OpenRouter | Requesty | SayGM | Vercel |
|---|---|---|---|---|---|
| Anthropic | |||||
| Claude Fable 5.1 | $18.00 | $18.99 | $18.90 | $16.00 (cheapest) | $18.00 |
| Claude Opus 5.5 | $7.20 | $7.60 | $7.56 | $6.40 (cheapest) | $7.20 |
| Claude Sonnet 5.5 | $3.60 | $3.80 | $3.78 | $3.20 (cheapest) | $3.60 |
| Claude Opus 5 | $9.00 | $9.50 | $9.45 | $8.01 (cheapest) | $9.00 |
| Claude Sonnet 5 | $3.60 | $3.80 | $3.78 | $3.20 (cheapest) | $3.60 |
| Claude Haiku 4.5 | $1.80 | $1.90 | $1.89 | $1.60 (cheapest) | $1.80 |
| OpenAI | |||||
| GPT-6.1 Sol | $3.60 | $3.80 | $3.78 | $2.68 (cheapest) | $3.60 |
| GPT-6 Luna | $0.18 | $0.19 | $0.19 | $0.13 (cheapest) | $0.18 |
| GPT-5.5 | $10.00 | $10.55 | $10.50 | $7.45 (cheapest) | $10.00 |
| GPT-5.6 Terra | $4.00 | $4.22 | $4.20 | $2.98 (cheapest) | $4.00 |
| GPT-5.6 Luna | $0.40 | $0.42 | $0.42 | $0.30 (cheapest) | $0.40 |
| Gemini 3.1 Pro Preview | $4.00 | $4.22 | $4.20 | $2.78 (cheapest) | $4.00 |
| Gemini 3.5 Flash | $3.00 | $3.16 | $3.15 | $2.08 (cheapest) | $3.00 |
| Open-weight | |||||
| Kimi K3 | $5.40 | $2.24 | – | $1.76 (cheapest) | – |
| DeepSeek V4.1 Flash | $0.48 | $0.13 (cheapest) | – | $0.29 | – |
Anthropic
- Claude Fable 5.1$16.00List $18.00SayGM
- Claude Opus 5.5$6.40List $7.20SayGM
- Claude Sonnet 5.5$3.20List $3.60SayGM
- Claude Opus 5$8.01List $9.00SayGM
- Claude Sonnet 5$3.20List $3.60SayGM
- Claude Haiku 4.5$1.60List $1.80SayGM
OpenAI
- GPT-6.1 Sol$2.68List $3.60SayGM
- GPT-6 Luna$0.13List $0.18SayGM
- GPT-5.5$7.45List $10.00SayGM
- GPT-5.6 Terra$2.98List $4.00SayGM
- GPT-5.6 Luna$0.30List $0.40SayGM
- Gemini 3.1 Pro Preview$2.78List $4.00SayGM
- Gemini 3.5 Flash$2.08List $3.00SayGM
Open-weight
- Kimi K3$1.76List $5.40SayGM
- DeepSeek V4.1 Flash$0.13List $0.48OpenRouter
Input and output prices separately · Price a monthly workload
02 · Explore
Where to go from here.
Prices
Providers
03 · Method
How the numbers are made.
- Basis
- Real-time, uncached tokens in US dollars per million, from each provider's published rate card.
- Fees
- OpenRouter's 5.5% card top-up fee and Requesty's 5% markup are added to the token price, because that is the rate an invoice reflects.
- Open-weight models
- Compared against the model maker's own price. OpenRouter figures use its cheapest listed provider. Requesty and Vercel are not tracked for these models.
- Dates
- Gateway terms as of September 2026; per-token prices as of 9 October 2026. Some providers reprice often, so treat every figure as a snapshot.
Sources: the labs' pricing pages, OpenRouter, Requesty and Vercel documentation, and the model catalogue on saygm.com (opens in a new tab).
04 · Questions
About the comparison.
Why does the same model cost different amounts from different providers?
The labs set a list price for their own API. Gateways resell that capacity and take their margin in different places: some add a fee when credit is bought, some add a percentage to usage, some charge list with no markup, and at least one bills below list by sourcing capacity from independent providers. The model and its output are the same.
What is an LLM gateway?
A service that puts many models behind one API key and one bill. Gateways add routing, fallback between providers and usage tracking. OpenRouter, Requesty, Vercel AI Gateway and SayGM are examples.
What is a blended price?
One number per model that weights input and output prices by a typical mix, here 80 percent input and 20 percent output tokens. It makes models and providers comparable at a glance. For workloads with long outputs, compare input and output prices separately on the price table.
Are gateway fees included?
Yes. OpenRouter's 5.5 percent card top-up fee and Requesty's 5 percent markup are added to their token prices, because that is the effective rate an invoice reflects.
Is buying direct from the lab ever cheaper?
For work that can wait, yes. Anthropic, OpenAI and Google sell batch processing at half of list price, below any real-time route. The table here compares real-time prices only.
How often are prices updated?
Prices are dated snapshots. Some providers reprice often, so confirm the current rate with the provider before budgeting.