We measure, so you decide with evidence.

ProbeArc is the independent measurement lab for open models. Which model fits your workload. What it really costs. Whether to go hosted or run your own hardware. And what changed overnight.

205endpoints
16models
50providers
30,007requests this window
2×/daysweep cadence
$0.336/M$0.693/M
glm-5.2 @ StreamLake: input price +106% overnight, caught 2026-08-24

The open-model world moves faster than anyone can track by hand.

New models land every week. Prices change without notice, sometimes doubling in a day. The same model name can be served at different quantization and different speed depending on the provider, and the differences rarely appear on a pricing page. Teams building on open models end up choosing on vibes, benchmarks that do not match their workload, or whatever worked last quarter.

We exist so you do not have to track any of this alone. ProbeArc is a standing measurement system pointed at the open-model ecosystem, and every product we sell is an answer generated from it.

Why ProbeArc

Independent

We take no money from model providers or hosts. Our conflict-of-interest policy is public. If a conflict ever arises, we say so upfront.

How we stay independent

Your usage, not averages

A leaderboard tells you who wins on average. We test on your actual tasks, at your volumes, against the endpoints you would really use.

How a bake-off works

Measured, continuously

Answers come from a system that sweeps hundreds of endpoints every day, not from opinion. Ask the same question next month and the data has not gone stale.

How we measure

Three instruments

mprobe · model selection

A model bake-off on your own workload. Send us a sample of real tasks, get a ranked pick with evidence, cost math, and a go/no-go.

qprobe · serving verification

Verification of what an endpoint actually serves: model fidelity and quantization, measured from the outside.

gprobe · the radar

Continuous sweeps across the endpoint grid, with alerts when something that affects you changes.

Start free. Pay when it gets serious.

Shortlist, free, from the standing grid. Bake-off, when the decision is worth doing properly. Enterprise, when staying right matters every week.

Choosing a model is easy. Staying right is the hard part.

Tell us what you are building. We will show you what the data says.