Everything below runs on the same standing infrastructure: a grid of 238 open-model endpoints across 52 providers, swept twice a day, scored by the same evaluation harness we use for our published benchmark work.
The core product. You send 30 to 50 anonymized examples of your real tasks. We run them through our evaluation harness across the candidate models and the actual endpoints you would deploy on, with multiple trials and published margins of error.
You get a ranked pick and runner-up, the evidence for both, transparent cost arithmetic at your volumes, latency and concurrency behavior, and a go/no-go matrix. Models that tie within the margin are reported as tied. Days, not weeks.
delivery/bakeoff/testcases/staging/demo-kamal-2026-08-28, scored
2026-08-30), namespace-isolated so
it never enters any public leaderboard. It exercises the pipeline end
to end; it is not a powered evaluation of these models, and it is not
a customer result.The same model name does not mean the same model. Quantization, serving stack, and configuration differ across providers, and the differences are invisible from a pricing page. qprobe measures endpoints from the outside and flags when what is served does not behave like what is claimed.
Inside a bake-off it answers a second question: which quantization your workload actually tolerates, so you stop paying for fidelity you do not need, and stop losing quality you did.
The ecosystem does not hold still after you choose. gprobe watches it for you: price changes on your endpoints, deprecations, and new releases relevant to your task mix, delivered as alerts plus a monthly digest.
Configured to your stack, so you hear about what affects you and nothing else.
Our benchmark harness, Mfold, and our measurement method are documented on the method page. The receipts are on the live data page.