Everything below runs on the same standing infrastructure: a grid of 205 open-model endpoints across 50 providers, swept twice a day, scored by the same evaluation harness we use for our published benchmark work.
Model selection on your workload.
The core product. You send 30 to 50 anonymized examples of your real tasks. We run them through our evaluation harness across the candidate models and the actual endpoints you would deploy on, with multiple trials and published margins of error.
You get a ranked pick and runner-up, the evidence for both, transparent cost arithmetic at your volumes, latency and concurrency behavior, and a go/no-go matrix. Models that tie within the margin are reported as tied. Days, not weeks.
Know what you are actually being served.
The same model name does not mean the same model. Quantization, serving stack, and configuration differ across providers, and the differences are invisible from a pricing page. qprobe measures endpoints from the outside and flags when what is served does not behave like what is claimed.
Inside a bake-off it answers a second question: which quantization your workload actually tolerates, so you stop paying for fidelity you do not need, and stop losing quality you did.
The no-rude-shocks subscription.
The ecosystem does not hold still after you choose. gprobe watches it for you: price changes on your endpoints, deprecations, and new releases relevant to your task mix, delivered as alerts plus a monthly digest.
Configured to your stack, so you hear about what affects you and nothing else.
Our benchmark harness, Mfold, and our measurement method are documented on the method page. The receipts are on the live data page.