Inferway
ModelsQwen3.8 27B

Qwen3.8 27B

Live
inferway/qwen3-8-27b

One model, served at NVFP4 precision on hardware we operate, with a 262K context window and an OpenAI-compatible endpoint.

NVFP4262K ctx32,768 max outputStreamingZero retention

Price

from catalog
Input$0.40per 1M tokens
Output$2.80per 1M tokens
Cache hit$0.0615% of input
Per token, metered from request metadata. No seats, no minimums.pricing_version 2026-08-19
Full pricing and billing rules

Context and output

Context window262,144Tokens per request
Max output32,768Tokens per response

Everything you send counts toward the context window — system prompt, conversation history, retrieved context and tool definitions. Both figures are read from the catalog, so they cannot drift from what the gateway enforces.

Rate limits and quotas

What the gateway enforces, and when each limit stops applying.

Free and keyless use60 / min · 120 / hour · 1,000 / dayApplies to unpaid accounts and to requests with no key at all.
Funded accountsAccount limitsOnce the account has credit, the free caps no longer apply — your account's own limits govern, visible in the console.
Non-streaming timeout~100 s first byteNon-streaming requests pass through a CDN. Long outputs must use stream=true.
Guest playgroundNo key requiredThe box below calls the same gateway under the free-use caps. No sign-up, no card.
Concurrency64 / key, defaultEnforced per API key as a concurrency lease; the default is 64. Account and per-key limits can be adjusted in the console. Free and keyless requests are not governed by this row.
All limits read from the gateway configuration

Benchmarks

Being measured

No number is published before it is measured. These cells fill in when the runs are complete and their conditions can be stated.

First-token latencyWill state GPU model, concurrency, timestamp
Output throughputTokens per second, single stream
90-day uptimePer service, planned windows marked separately
Load reportsWhat a published number will carry

Try it

No key needed · free-use caps apply
Ask anything. The response streams token by token from the same gateway your code will call.
Copy setup for your AI assistant
Installs the SDK, sets base_url and the model id, reads the key from an environment variable, writes a streaming smoke test.
Quickstart

Data and hardware

Data handlingNothing is written to disk

Inference runs in GPU memory and is erased on completion. Only request metadata — token counts, timestamps, latency, routing — is retained, for billing and reliability.

Transparency hub
HardwareHardware we operate

This model is served at NVFP4 on hardware we operate, with no silent downgrade under load. Execution happens in one published region, the United States, and the region each request ran in is recorded on that request.

Transparency hub

Everything about this model is on this page. Go run a request.