Load testing for LLM applications. We measure what streaming endpoints actually do when more than one person is using them.
A normal load test reports response time. For a streaming endpoint that averages together two different failures: the wait before the first token, which is mostly queueing, and the gaps between tokens while the answer streams, which is generation speed. They have different causes and different fixes, and averaging them hides both.
So we measure them separately, and we publish what we find.
🧠LLM Performance and Load — thirteen free and open models on Hugging Face Inference Providers, measured at 1, 25 and 100 concurrent users. Time to first token, inter-token latency, goodput, throughput, answer quality, error rate and cost, at every level, so each row reads across.
The raw per-request CSVs, the prompt set, the scoring code and the SLO thresholds are all published in that Space. Disagree with the thresholds and the goodput column changes — that's the point of publishing them.
Goodput over averages. A request counts only if it met every threshold at once: first token under 2s, finished under 15s, inter-token latency p95 under 80ms, correctly formatted, no error. A model can post a fast median and still fail most requests.
One provider pinned per row. The same model id can be served by ten providers at different quantisation, hardware and price. An unpinned number isn't a comparable number.
No client-side retries. A 429 is recorded as an error, not quietly retried away.
We publish our own variance. Running the same configuration twice, two hours apart, moved our numbers by more than rounding. That's in the Method tab, because a benchmark that hides its noise is asking to be taken more seriously than it deserves.
InferGauge is at infergauge.com. It's source-available under the Business Source License 1.1: read it, run it internally and in production, gate your CI with it. You can't resell it as a hosted service. It converts to Apache-2.0 in 2030.
The measurements above are provider-side, from a datacenter. They describe how the provider behaves, not how your application will. For numbers about your own system, run it against your own endpoint.