Model catalog
One model, documented in full.
Everything we serve is listed here with its price, limits and hardware. Nothing sits behind a sales call.
inferway/qwen3-8-27b
General-purpose chat and tool use, served at NVFP4 on hardware we operate. OpenAI-compatible — point base_url at us and keep your SDK.
NVFP4262K ctx32,768 max outputStreaming
Input$0.40
Output$2.80
Cache hit$0.06
USD per 1M tokens262K ctx · 32,768 max output
Free use is rate limitedpricing_version 2026-08-19