A benchmark harness can fail to deliver its requested load while reporting results as if it did, which makes server performance look worse or better than it is.
2
Inference Perf uses a multiprocess load generator, client-side telemetry, and declarative workload configuration to make large-scale tests observable and repeatable.
3
Benchmark validity depends on matching real workloads, including request distributions, temperature, token lengths, multi-turn behavior, and dataset handling.
Summary
Ashok Chandrasekar and Jason Kramberger explain why published LLM performance numbers can be difficult to reproduce. A Python benchmark harness may be limited by the global interpreter lock, so a request for 200 QPS can produce 38 QPS or stop near 170 QPS without making the failure clear. A client under load can also add tens of seconds to measured latency. Other differences come from temperature settings and from benchmark tools sampling and truncating the same dataset in different ways. Their answer is Inference Perf, a CNCF project from the Kubernetes serving working group. It schedules requests across worker processes, reports planned and actual send times, and places client telemetry beside server metrics. Its declarative configuration supports request-rate models, concurrency, conversation replay, and workload distributions. A shared workload catalog defines scenarios such as agentic generation, tree of thought, and batch summarization. Prism displays results for llm-d deployments.
Production-scale tests need more than a simple request-rate script
Chandrasekar separates developer-focused model-server scripts, competitive analysis tools, web benchmarks, and production-scale LLM benchmarks. Production tests may cover online serving, batch workloads, many servers, prefill and decode disaggregation, and autoscaling. They need to sustain high load, represent customer workloads, report accurate metrics, find the point where the server saturates, and measure SLOs such as P90 time to first token.
A harness can report a requested load that it never delivered
The speakers tested benchmark harnesses by asking them to generate 200 QPS. On a small machine, one harness produced only 38 QPS. On a more powerful machine, some single-process harnesses stopped near 170 QPS. Python's global interpreter lock can leave a CPU-bound process limited to one CPU even when the machine has more capacity. If the harness does not expose its actual rate, the reported server results can be mistaken for a 200-QPS test.
Client overload can make a healthy server look slow
A benchmark client may spend so much time collecting streaming token requests that it inflates the latency it measures. In one test, the delay reached 58 seconds. The speakers also ran a 1,000-QPS test against a simulated server, where a scalable harness produced almost no latency. This separated the server from the client and showed that the benchmark tool itself could be the bottleneck.
Benchmark settings can create misleading throughput gains
A shared result claimed 20 percent better throughput, but the harness had set model temperature to zero. That made outputs more deterministic and allowed higher throughput than a workload with a temperature around 0.7. The speakers also used the same ShareGPT dataset with two harnesses and got different input token counts because the tools sampled and truncated the data differently. Dataset handling and generation settings therefore change the workload being measured.
Inference Perf spreads load across processes and reports what actually happened
Kramberger describes Inference Perf, a CNCF project from the Kubernetes serving working group. Its main process schedules requests from a declarative plan, such as a Poisson process, a constant rate, or fixed concurrency. Worker processes send the requests and report both their planned and actual execution times. At 5,000 QPS, the architecture kept up and reported that it had kept up, giving users visibility into the harness as well as the system under test.
Declarative configuration makes complex workloads replayable
Inference Perf configuration can describe more than a random prompt and an endpoint. It supports conversation replay with input and output length distributions and other workload settings. This gives teams a way to describe a test in configuration and repeat it across runs. The configuration is intended to match the workload being tested rather than reduce every test to one synthetic request shape.
A shared workload catalog gives different tools common definitions
The workload catalog contains definitions for multi-turn generation, tree of thought, agentic generation, and batch summarization. Each workload has a natural-language description plus detailed configuration and metrics. The definitions are also expressed in generic terms so that other benchmark tools can adopt the same scenarios. This addresses the problem of tools using different interpretations of the same public dataset or workload.
Valid benchmark results require evidence from both sides of the test
Kramberger's closing principles require client concurrency and observability into whether the client met its configuration. Client and server metrics together show whether the scenario ran as intended. Randomness and determinism must match the real workload, and the dataset must resemble the data the system is meant to handle. Prism presents results for llm-d, including a comparison between combined optimizations and a plain Kubernetes service across eight TPU replicas.