Qwen3.8 Flash
inferway/qwen3.8-flashQwen's efficiency-focused Flash model.
- Model input
- TextImage
- Model output
- Text
Zero data retention: your data is not our data.
or read the quickstartTLS-encrypted, straight to a dedicated endpoint.
No shared queue in the path.
Paged KV, never written to disk.
Erased on completion; only metadata remains.
Per token, metered from metadata. No seats, no minimums.
Per second of video delivered. No seats, no minimums.
Coming soon
inferway/qwen3.8-flashQwen's efficiency-focused Flash model.
inferway/deepseek-v4.1-flashDeepSeek's efficiency-focused V4.1 Flash release.
inferway/glm-5.3-flashZ.ai's Flash model for coding and agentic workloads.
Free trial — requests are not counted in your account.
Checking live model availability…
Throughput and latency are being re-measured on the current hardware and model. Until that finishes, this section stays empty — we publish measurements with the precision served, concurrency, prompt shape and timestamp attached, or we publish nothing.
curl https://api.inferway.ai/v1/chat/completions \
-H "Authorization: Bearer $INFERWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "inferway/mimo-v2.6-flash",
"stream": true,
"messages": [{"role": "user", "content": "Say hello in one short sentence."}]
}'Longer answers: Docs · Privacy · Transparency