Pricing
Every price, published in full.
Per-token billing, metered from request metadata. No seats, no minimums, no sales call to see a number.
Free use is rate limitedpricing_version 2026-08-19
Things that could surprise you on a bill.
Written down here rather than discovered later.
01
Reasoning tokens bill as outputIf the model emits reasoning tokens, they are counted as output tokens at the output price. They appear in the usage breakdown in the console.
02
A cancelled stream still bills what it generatedTokens produced before you disconnect are metered. Cancelling a long generation early reduces the bill but does not zero it.
03
Cache hits are not guaranteedCache pricing applies when a prefix actually hits. A cold or evicted prefix bills at the normal input price, so estimates that assume a high hit rate can run high.
04
Failed requests are not billed, rejected ones are free5xx errors and gateway rejections cost nothing. A request that returns tokens and then errors bills for the tokens it returned.
05
Prices carry a versionEvery price ships under a pricing_version. Changes are published as a new version and announced; nothing changes silently mid-month.
Questions
You can call without an account and without a card, up to 60 requests per minute, 120 per hour and 1,000 per day. Once the account has credit those caps no longer apply — paid throughput runs on your account's own limits, visible in the console.