probearc

How we measure.

Every claim on this site and in our reports comes from a standing measurement system, not from opinion. This page says what that system actually does.

Daily sweeps

We sweep every endpoint we track on a fixed schedule, at least twice a day. Each sweep is stored, not just the latest one, so the record is a history rather than a snapshot, and a price, latency or error-rate move can be dated instead of guessed at.

Our own accounts

We run every probe through accounts we open and pay for ourselves, the same way any customer would. We do not use free or partner credits from a provider we measure without saying so on the pages that use that data.

Temperature-zero probes

Probes run at temperature zero. The same question gets a comparable answer sweep to sweep, so a change in the answer is a change in what is actually being served, not a roll of the dice.

Published margins

We only call something a move when it clears a real margin, not noise. A latency change is reported only when the two measurement windows stop overlapping. A sample too small to support a claim is left out rather than rounded into one.

The Mfold harness

Model quality claims come from Mfold, our benchmark harness: banded tasks, a judging panel, multiple trials, and a published margin of error. Models that tie within the margin are reported as tied, never ranked by decimals the margin cannot support.

Corrections

When a number we published turns out wrong, we fix it in place and say so. We do not quietly reissue it.

How we stay independent from the providers and models we measure is its own page: independence.