Recommend sizing
Start with just a model name — we fill the rest, then let you tune every assumption.
Tested: Nemotron, DeepSeek V4, Gemma 4, Kimi, ... — type to autocomplete
Gated model? Add your HF token in Settings →
Based on your configuration — ISL 2048, OSL 128, TTFT target 1s, concurrency 1, TPOT 30 ms.