Overview
Continuous comparison: ingest traffic, run comparisons, living switch/route verdict.
A Monitor is a standing decision about one production workload. It owns the exact production model, candidate policy, traffic sources, decision contract, judge pool, schedule, and living verdict.
Set up a Monitor
- Choose where traffic will come from. OpenTelemetry is recommended because it preserves roles, tool structure, and trace identity.
- Select the exact catalog model serving production. PeerLM preserves source model strings but never guesses the executable incumbent from them.
- Choose a task policy, which sets how much quality loss the Monitor can tolerate before recommending a change.
- Accept rolling candidate coverage or choose candidates manually. Advanced settings can change the Monitor name, candidate categories, and monthly production volume.
Manual first: saving setup does not start a Run, consume pooled capacity, or enable a schedule. A saved connector also does not mean traffic has arrived; the Monitor shows received records separately.
From traffic to verdict
- Ingest — normalize production records and record their structural fidelity.
- Sample — select a fixed, eligible sample, including required critical categories.
- Replay — generate every candidate and use either captured or regenerated control evidence.
- Judge — seat the full rotating panel for pairwise decisions.
- Aggregate — apply the frozen decision contract and publish switch, route, or hold.
Run receipt and Monitor verdict
The Run freezes its inputs and contains the comparison receipt: prompts, outputs, judge votes, failure patterns, cost range, latency evidence, and methodology. The Monitor rolls completed Runs into the current living verdict. Changing Monitor settings after queueing never changes what an existing Run executes.