ProbeArc is the independent measurement lab for open models. Which model fits your workload. What it really costs. Whether to go hosted or run your own hardware. And what changed overnight.
New models land every week. Prices change without notice, sometimes doubling in a day. The same model name can be served at different quantization and different speed depending on the provider, and the differences rarely appear on a pricing page. Teams building on open models end up choosing on vibes, benchmarks that do not match their workload, or whatever worked last quarter.
We exist so you do not have to track any of this alone. ProbeArc is a standing measurement system pointed at the open-model ecosystem, and every product we sell is an answer generated from it.
We take no money from model providers or hosts. Our conflict-of-interest policy is public. If a conflict ever arises, we say so upfront.
How we stay independentA leaderboard tells you who wins on average. We test on your actual tasks, at your volumes, against the endpoints you would really use.
How a bake-off worksAnswers come from a system that sweeps hundreds of endpoints every day, not from opinion. Ask the same question next month and the data has not gone stale.
How we measureA model bake-off on your own workload. Send us a sample of real tasks, get a ranked pick with evidence, cost math, and a go/no-go.
Verification of what an endpoint actually serves: model fidelity and quantization, measured from the outside.
Continuous sweeps across the endpoint grid, with alerts when something that affects you changes.
Shortlist, free, from the standing grid. Bake-off, when the decision is worth doing properly. Enterprise, when staying right matters every week.
Tell us what you are building. We will show you what the data says.