Qwen3.8 27B
LiveOne model, served at NVFP4 precision on hardware we operate, with a 262K context window and an OpenAI-compatible endpoint.
Price
from catalogContext and output
Everything you send counts toward the context window — system prompt, conversation history, retrieved context and tool definitions. Both figures are read from the catalog, so they cannot drift from what the gateway enforces.
Rate limits and quotas
What the gateway enforces, and when each limit stops applying.
Benchmarks
Being measuredNo number is published before it is measured. These cells fill in when the runs are complete and their conditions can be stated.
Try it
No key needed · free-use caps applyData and hardware
Inference runs in GPU memory and is erased on completion. Only request metadata — token counts, timestamps, latency, routing — is retained, for billing and reliability.
Transparency hubThis model is served at NVFP4 on hardware we operate, with no silent downgrade under load. Execution happens in one published region, the United States, and the region each request ran in is recorded on that request.
Transparency hub