Monitors & comparison Runs
Compare production traffic against candidates for quality and cost.
Monitors let you plug production traffic into PeerLM and accumulate comparison Runs against candidates for quality and cost. The Monitor shows a living switch / route / hold verdict; each Run holds the receipts. Authored one-off Suite Runs remain under Suites / Runs for shareable growth reports.
Starting from a Monitor
Navigate to Monitors. Create a Monitor (or Configure with AI), set control and candidates, add a data source, then start a Run.
1. Import or connect traffic
Upload CSV/JSONL/JSON or connect the SDK / an observability connector. Include response and model when possible for control-parity judging.
2. Select candidates & judges
Up to 5 candidates. Judges can be auto, by tier, or manual. Default criteria cover adherence, hallucination, completion, and format.
3. Launch a Run
Review the credit estimate and start. The pipeline samples, replays, judges, and aggregates—then refreshes the Monitor verdict.
Reading results
- Monitor verdict — quality retained vs control, savings, switch / route / hold
- Trends — deltas across Runs (not absolute scores)
- Routing — per-task-slice recommendations + exports
- Run receipts — loss examples and methodology for one batch