DocsCore ConceptsWhat You Pay For

What You Pay For

Monitors, runs and triggers — and why inference is included.

One unit: what PeerLM runs

Your plan buys monitors, runs and triggers. Model inference is included — PeerLM pays the providers, and we do not pass a token bill through to you. Hosted-generation ceilings bound what PeerLM funds at list price.

Runs

A run is one comparison cycle: a sample of your own production prompts, compared against each candidate, judged by a rotating panel, and aggregated into a verdict. Run capacity is pooled across every Monitor in your organization.

A Run counts immediately before its first provider request. Queue and pre-provider failures release the reservation; after inference begins, the Run counts even if later work fails.

Observed and replayed control

Production Monitors default to observed control: your captured production response is compared against up to five challengers, subject to the plan's hosted-generation ceiling, without calling the incumbent provider. PeerLM requires at least 90% usable incumbent-output coverage and blocks the Run when the evidence is incomplete.

Choose replayed control to replay both sides as they behave today. Prompt CI always uses replayed control on the same model. PeerLM never changes control modes automatically.

What a run costs us

Before you start a run, PeerLM shows what the inference will cost at provider list price — always as a range, with the date those prices were last synced. The range is not hedging: the figure is derived from cached list prices and typical response lengths, and a product built to catch overbilling has no business printing false precision about its own numbers.

Why can't I pick the judges? PeerLM selects them. A model from the vendor under test can't judge it, and a panel the customer picks is a panel we can't stand behind as evidence. Choosing your own judges is available on Enterprise.