ProbeArc is the independent measurement lab for open models. Which model fits your workload. What it really costs. Whether to go hosted or run your own hardware. And what changed overnight.
The same standing grid, read six different ways. Nothing below is illustrative — every mark is a real request we ran.
| model | provider | from ms | to ms | change | day |
|---|---|---|---|---|---|
| inkling | DeepInfra | 2,537 | 59,975 | +2264% | 2026-08-29 |
| muse-glimmer-30b | DeepInfra | 1,632 | 28,535 | +1648% | 2026-08-28 |
| gpt-oss-120b | Novita | 1,636 | 27,620 | +1588% | 2026-08-31 |
| inkling | Together | 2,522 | 39,362 | +1461% | 2026-08-29 |
| nemotron3-ultra | Venice | 3,136 | 24,381 | +678% | 2026-08-27 |
| muse-glimmer-30b | DeepInfra | 1,638 | 12,030 | +634% | 2026-08-26 |
The four biggest price moves and two biggest latency moves the sweep caught in the window 2026-08-20 to 2026-08-31 on endpoints that had been holding steady beforehand (91 price and 30 latency moves in total; the full ledgers are on the live data page), dated, wide enough to read at a glance instead of scrolling a column.
A model bake-off on your own workload. Ranked pick, evidence, cost math, go/no-go.
Verification of what an endpoint actually serves: model fidelity and quantization, measured from the outside.
Continuous sweeps across the endpoint grid, with alerts when something that affects you changes.