Qwen Console

multi-model gateway

checking
Live decodetok/s · streaming
TTFTtime to first token
Outputtokens this reply
Backend loadrunning / queued

Ask anything

Pick a model, toggle Think or Fast, and go.

0 / 32k ctx
Throughputtok/s · all models
Activerunning requests
Pendingqueued requests
Total callssince start
Prefix cachehit rate
Mean callinference time

GPUs utilization & VRAM

Throughput tok/s · last 2 min

MTP acceptance draft accept · last 2 min

Backends

ModelGPURunWaittok/sKVCacheMTP

Model endpoints what's serving right now

ModelStatusEngineGPUEndpoint (OpenAI /v1)CtxCapabilities

Access points how to reach these models

Call History

1
TimeModelSrcStatusConnRecvPromptOutPrefillTTFTtok/sKV MB

Call