v0.2.6 · 6f5fc22
2026-09-21 03:42:35 UTC

Performance estimate

See time to first token, throughput, and memory estimates in seconds — refine as needed.

Tested: Nemotron, DeepSeek V4, Gemma 4, Kimi, ... — type to autocomplete
Serving mode:
Based on your configuration — ISL 2048, OSL 128, FP16 KV cache, 1 concurrent users.
Want to change assumptions?

Every number above comes from these. Open a section to tune it — closed sections show their current values.