Inferway
ModelsPricingPlaygroundDocsStatusTransparency
Sign inSign up
ModelsPricingPlaygroundDocsStatusTransparency
Sign inSign up
Pricing

Lower prices, published in full.

What gets billed
InputEverything you send - system prompt, history, retrieved context, tool definitions.
OutputEverything generated, including reasoning tokens. Billed at the standard output token rate. Thinking is off by default; enable it per request with reasoning_effort=low|medium|high (none keeps it off) or chat_template_kwargs.enable_thinking=true.
Cache hitRepeated prefixes served from cache, billed well below the input rate.
per tokenmetered from metadatano seats, no minimumsversioned prices

Video generation

Billed by the second of video ordered, at the published rate. No subscription and no per-request minimum.

MiniMax H3 768P

Asynchronous InteractionModel card
inferway/minimax-h3-768p

MiniMax H3 768P video generation with audio, served with int8 pruned weights.

OutputLoading prices

Ordered in whole seconds; the delivered clip is rounded up to the model's frame grid at no extra charge.

Inferway Media Retention Policy v1: generation outputs are retained temporarily for retrieval.

Video generationINT8

Coming soon

Qwen3.8 Flash

inferway/qwen3.8-flash
Coming soon

Qwen's efficiency-focused Flash model.

Model input
TextImage
Model output
Text
PriceComing soon

DeepSeek V4.1 Flash

inferway/deepseek-v4.1-flash
Coming soon

DeepSeek's efficiency-focused V4.1 Flash release.

Model input
TextImage
Model output
Text
PriceComing soon

GLM-5.3 Flash

inferway/glm-5.3-flash
Coming soon

Z.ai's Flash model for coding and agentic workloads.

Model input
TextImage
Model output
Text
PriceComing soon

Things that could surprise you on a bill.

01
Reasoning tokens bill as outputIf the model emits reasoning tokens, they are counted as output tokens at the output price. They appear in the usage breakdown in the console. Thinking is off by default; enable it per request with reasoning_effort=low|medium|high (none keeps it off) or chat_template_kwargs.enable_thinking=true.
02
A cancelled stream still bills what it generatedTokens produced before you disconnect are metered. Cancelling a long generation early reduces the bill but does not zero it.
03
Cache hits are not guaranteedCache pricing applies when a prefix actually hits. A cold or evicted prefix bills at the normal input price, so estimates that assume a high hit rate can run high. The cache matches whole blocks of tokens, so a short prompt may not hit at all and then bills entirely at the input rate.
04
Failed requests are not billed, rejected ones are free5xx errors and gateway rejections cost nothing. A request that returns tokens and then errors bills for the tokens it returned.
05
Prices carry a versionEvery price ships under a pricing_version. Changes are published as a new version and announced; nothing changes silently mid-month.
Invite friends and earn bonus credit when they spend.
Referral termsGo to referrals
Accepted feedback earns bonus credit.
Feedback terms

Questions

Registered free accounts: 8 requests/second burst, 90 requests/minute, 500 requests/hour and 1,000 requests/day, with 4 concurrent requests. No per-request input or output token cap, and no per-minute or daily token quota. Hourly and daily request counts are rolling windows, not a reset at a fixed time. API keys on the same account share these request counts.
Put the stable part of your prompt first — system prompt, few-shot examples, long retrieval context. When a later request repeats that exact prefix it hits, and those tokens bill at that model's cache rate instead of its input rate. Keep the varying part at the end. The cache matches whole blocks of tokens, so a short prompt may not hit at all and then bills entirely at the input rate.
They can, but under a version. Every price ships with a pricing_version, and changes are published as a new version and announced. Nothing changes silently mid-month.

Every number on this page is public. Go run a request.

Get an API keyRead the quickstart
Inferway

Self-hosted inference API with public metrics. Prompts and completions never touch disk.

Developers
API DocsConsolePlaygroundSystem StatusTransparency Hub
Product
Launch offerModels & PricingModel CatalogContact
© 2026 GWMM LLC (operating as Inferway)
Privacy PolicyTerms of ServiceCookie PolicyData Processing AgreementAcceptable Use PolicyLegal Hub
© 2026 GWMM LLC (operating as Inferway)hello@inferway.aiinferway.ai