MiMo-V2.6 Flash
LiveServed at FP8 precision on hardware we operate, with a 262K context window and an OpenAI-compatible endpoint.
Price
from catalogContext and output
Everything you send counts toward the context window — system prompt, conversation history, retrieved context and tool definitions. Both figures are read from the catalog, so they cannot drift from what the gateway enforces.
Rate limits and quotas
Benchmarks
Being measuredTry it
Data and hardware
Inference runs entirely in memory and is erased on completion. Only request metadata — token counts, timestamps, latency, routing — is retained, for up to 90 days, for billing and reliability.
Transparency hubThis model is served at FP8 on hardware we operate, with no silent downgrade under load. Execution happens in one published region, the United States, and the region each request ran in is recorded on that request.
Transparency hub