Experiment · 001 / Live
Inference latency
explorer
Change the workload and watch prefill, decode, and queueing compete for the latency budget.
Conceptual model, not a hardware benchmark. Values expose relationships, not production predictions.
Estimated latency377ms
Token throughput2716tok/s
Decode dominates: output length is the strongest lever in this configuration.
This model extends the active episode on inference as a systems problem.
Read the episode →