probearc

We measure, so you decide with evidence.

ProbeArc is the independent measurement lab for open models. Which model fits your workload. What it really costs. Whether to go hosted or run your own hardware. And what changed overnight.

Every instrument, at once

The same standing grid, read six different ways. Nothing below is illustrative — every mark is a real request we ran.

Price moves caught this window

Change in input price, percent, by endpoint.
deepseek-v4-pro @ StreamLake+146%glm-5.2 @ Novita+117%deepseek-v4-flash @ DigitalOcean+106%glm-5.2 @ StreamLake+106%deepseek-v4-pro @ Baidu+104%deepseek-v3.2 @ DigitalOcean+100%deepseek-v4-pro @ DigitalOcean+100%glm-5.2 @ DigitalOcean+100%
91 input-price moves cleared the margin.

Endpoints adjudicated, by day

08-2508-2608-2708-2808-2908-3008-31
unchanged changed new gone
1,578 endpoint-days adjudicated, 256 changed.
endpoints, last sweep
219flat vs 08-30
models, last sweep
17flat vs 08-30
providers, last sweep
51flat vs 08-30
changed, last sweep
37+7 vs 08-30
price moves, window
9112d
requests, window
130,7022×/day
ttft, fastest
116msGroq
ttft, median
1,044msn=205
last adjudicated day 2026-08-31, deltas against 2026-08-30 · window 2026-08-20 to 2026-08-31 (7 of 12 days shown) · TTFT gauges = latest timing record per endpoint as of 2026-08-31 · generated 2026-09-01T20:29:33Z

One endpoint, 7 days

$0.081 /M in, deepseek-v4-flash at Baidu
$0.081$0.08108-2508-31

Latency moves

modelprovider from msto ms changeday
inklingDeepInfra2,53759,975+2264%2026-08-29
muse-glimmer-30bDeepInfra1,63228,535+1648%2026-08-28
gpt-oss-120bNovita1,63627,620+1588%2026-08-31
inklingTogether2,52239,362+1461%2026-08-29
nemotron3-ultraVenice3,13624,381+678%2026-08-27
muse-glimmer-30bDeepInfra1,63812,030+634%2026-08-26
inkling at DeepInfra: 1,838 to 93,067 ms.

Time to first token

Latest timing record per endpoint as of 2026-08-31 (205 with TTFT among 219 endpoint files seen that day; records may come from more than one sweep, of 238 tracked across the window), log-scaled because the spread runs 116ms to 36,846ms in one sample.
200ms500ms1s2s5s10s30s
Fastest: gpt-oss-120b at Groq, 116ms. Slowest: glm-5.3-flash at Wafer, 36,846ms — an outlier worth its own look, not a typical reading. Median 1,044ms.

The receipts

The four biggest price moves and two biggest latency moves the sweep caught in the window 2026-08-20 to 2026-08-31 on endpoints that had been holding steady beforehand (91 price and 30 latency moves in total; the full ledgers are on the live data page), dated, wide enough to read at a glance instead of scrolling a column.

2026-08-28

deepseek-v4-flash at DigitalOcean+106%

$0.068 to $0.140 per million input tokens.
pin digitalocean · caught by the sweep of 2026-08-28
2026-08-31

deepseek-v4-pro at Baidu+104%

$0.426 to $0.869 per million input tokens.
pin baidu/fp8 · caught by the sweep of 2026-08-31
2026-08-28

deepseek-v3.2 at DigitalOcean+100%

$0.250 to $0.500 per million input tokens.
pin digitalocean · caught by the sweep of 2026-08-28
2026-08-28

deepseek-v4-pro at DigitalOcean+100%

$0.870 to $1.740 per million input tokens.
pin digitalocean · caught by the sweep of 2026-08-28
2026-08-29

inkling at DeepInfra+2264%

Median 2,537ms to 59,975ms.
pin deepinfra/fp8 · caught by the sweep of 2026-08-29
2026-08-29

inkling at Together+1461%

Median 2,522ms to 39,362ms.
pin together · caught by the sweep of 2026-08-29

Three instruments, one standing system

mprobe · model selection

A model bake-off on your own workload. Ranked pick, evidence, cost math, go/no-go.

qprobe · serving verification

Verification of what an endpoint actually serves: model fidelity and quantization, measured from the outside.

gprobe · the radar

Continuous sweeps across the endpoint grid, with alerts when something that affects you changes.