probearc

Three instruments, one measurement system.

Everything below runs on the same standing infrastructure: a grid of 238 open-model endpoints across 52 providers, swept twice a day, scored by the same evaluation harness we use for our published benchmark work.

mprobe
Model selection on your workload.

The core product. You send 30 to 50 anonymized examples of your real tasks. We run them through our evaluation harness across the candidate models and the actual endpoints you would deploy on, with multiple trials and published margins of error.

You get a ranked pick and runner-up, the evidence for both, transparent cost arithmetic at your volumes, latency and concurrency behavior, and a go/no-go matrix. Models that tie within the margin are reported as tied. Days, not weeks.

What a bake-off verdict looks like

emaildrafthinglishreplyinvoiceextractionreconciliation1deepseek-v4-flashopenrouter · T1 ready1.001.001.001.001qwen36-35b-a3bllamacpp · T1 ready1.001.001.001.003haiku-4.5anthropic · T2 usable with review0.840.950.670.64gpt-5.4-miniopenai · NRnot reached — 15/15 requests rate-limited (HTTP 429) at the time of the run; not scored
Rank, model, and fit per task family. Bars run 0 to 1 on the buyer's own task mix (email drafting to spec, hinglish customer reply, invoice field extraction, tabular reconciliation); teal is a go on that family, iris means usable with review, and a muted-grey bar is a no-go. A bar left empty is not a low score: it means that family was never scored for that row. Ranked by tier first and then average fit — the two top models tie and are shown tied, never split by decimals the evidence cannot support. One candidate was not reached during this run and is reported that way, unscored, rather than quietly dropped: gpt-5.4-mini (15/15 requests rate-limited (HTTP 429) at the time of the run). That is a fact about the run, not about the model.
Where the data comes from: the 12-task synthetic demo set our friends-and-family lab runs (delivery/bakeoff/testcases/staging/demo-kamal-2026-08-28, scored 2026-08-30), namespace-isolated so it never enters any public leaderboard. It exercises the pipeline end to end; it is not a powered evaluation of these models, and it is not a customer result.
v1-era demo data
qprobe
Know what you are actually being served.

The same model name does not mean the same model. Quantization, serving stack, and configuration differ across providers, and the differences are invisible from a pricing page. qprobe measures endpoints from the outside and flags when what is served does not behave like what is claimed.

Inside a bake-off it answers a second question: which quantization your workload actually tolerates, so you stop paying for fidelity you do not need, and stop losing quality you did.

What a qprobe verdict looks like

higher fidelitylower fidelitybf16 referencea reference we serve ourselvesfp8-classQ8-classQ6-classQ5-classQ4-classthis endpoint behaves like this rungQ3-classverdictinconsistent with thebf16 referenceclosest to Q4-classno numbers — illustrative
A schematic, not a measurement. A qprobe verdict places an endpoint against references we serve ourselves and says what it is closest to. This drawing shows the shape of that answer with no numbers on it at all, because our own rule is calibration before claims: we publish no quantization verdict about any endpoint until the calibration matrix reports. Nothing on this figure refers to a real endpoint.
illustrative — calibration in progress
gprobe
The no-rude-shocks subscription.

The ecosystem does not hold still after you choose. gprobe watches it for you: price changes on your endpoints, deprecations, and new releases relevant to your task mix, delivered as alerts plus a monthly digest.

Configured to your stack, so you hear about what affects you and nothing else.

A move gprobe caught

gemma4-31b at Together · pin together · input price, dollars per million tokens$0.280$0.390+39%caught by the sweep of 2026-08-28
A real receipt, not an example. gemma4-31b at Together (pin together) went from $0.280 to $0.390 per million input tokens, a 39% rise, caught by our sweep of 2026-08-28 — one of the moves listed in full on the live data page.
Where the data comes from: the twice-daily endpoint sweep, adjudicated day over day; every move that cleared the margin in the current window is listed on the live data page.
real sweep data

Our benchmark harness, Mfold, and our measurement method are documented on the method page. The receipts are on the live data page.