live on the first-party API by Anthropicreleased 2026-06-30text
A million-token context with every rate class published, cache and batch included.
Quick answer prices verified
Best for our call
A million-token context with every rate class published, cache and batch included.
Cheapest
$10.00at Claude API · 1M output tokens · claimed
First-party
$10.00at Claude API · 1M output tokens · claimed
Routes selling
1 of 1 tracked
Verified
2026-08-15
Claude Sonnet 5, from Anthropic, is purchasable on the one route we track. 1M output tokens costs $10.00 at Claude API, the cheapest published rate we can cite.
The figure above buys 1M output tokens. The board below prices 7 units in all, each column named. Figures come from each seller's own published rates and carry the date we last read them.
Which route should you take
Each row is the table above read for one question, at the unit named beside the price. Rows marked our call are editorial judgement instead, and appear only where we have something specific to say.
Each column is one unit, and every price in it buys the same thing. The figures quoted elsewhere on this page are 1M output tokens. That is a baseline for comparing providers, not a quote. What you are actually billed depends on the settings you send, and on some routes on how much you top up at once. Every figure is the provider's own published price; anything we derived rather than read off a price list is marked est.
How each figure above was read off the provider's own price list, and what that price list does not say.
Claude API
Every figure here is a rate Anthropic publishes for a million tokens, one per rate class. The cache and batch rates are separate prices for the same token streams, not a discount we applied: Anthropic prints the batch figures in a table of their own. Nothing on this page combines them.
Reference workload
This is not a price Anthropic publishes, and not a measurement of anything. It is a fixed reference scenario, 100K input + 20K output tokens, costed at the rates below and added up, so that models can be compared at one scenario instead of at whichever rate each one looks best on. The mix is ours and is not a claim about what a typical request looks like: change it and every figure here moves. The published rates are in the table under it, unchanged.
Reference scenario: 100K input + 20K output tokens, at the standard rate class
Where a row says not verified, we have not checked that figure yet; it does not mean the model lacks the feature. Every row is read off published documentation for the model, never from our own testing.
These sell 100K input + 20K output tokens too, so their figures on the index can be read against this one. Models that do not are shown with their own unit instead.