Experiment · 001 / Live

Inference latency
explorer

Change the workload and watch prefill, decode, and queueing compete for the latency budget.

Conceptual model, not a hardware benchmark. Values expose relationships, not production predictions.

Estimated latency377ms
Token throughput2716tok/s
Prefill67 ms
Decode287 ms
Queue23 ms

Decode dominates: output length is the strongest lever in this configuration.