Glossary
Definitions of key terms used throughout PeerLM.
| Term | Meaning |
|---|---|
| Monitor | A standing production decision with sources, policies, Runs, and a living verdict |
| Run | One frozen sample/replay/judge/aggregate cycle; the evidence receipt |
| Trigger | A manual, scheduled, catalog, price, drift, or deploy event that starts a Run |
| Control | The incumbent evidence: captured production output in observed mode or regenerated output in replayed mode |
| Candidate | An alternative model compared one-to-one with the control |
| Judge pool | The models frozen on a Monitor from which each item's full panel is seated |
| Decision contract | The thresholds and critical categories frozen before a Run |
| Evidence strength | A 0–100 description of sample sufficiency, execution cleanliness, and determinate agreement |
| Switch | A candidate held quality under the contract and is worth replacing the incumbent with |
| Route | Evidence supports a candidate only for specific stable workload categories |
| Hold | Evidence does not support a change because of quality, economics, or insufficient evidence |
| Observed control | Uses the exact captured production response without calling the incumbent |
| Replayed control | Regenerates the incumbent under the frozen Run context |
| Sponsored first look | A free, directional comparison on pasted prompt/response pairs; not a standing Monitor |
| Run pack | A $99/month self-serve expansion adding one Monitor slot and 10 pooled Runs |
| Prompt CI | A deploy Run comparing old and new system prompts on the same model |
Compatibility terminology: REST and MCP may still expose Suite evaluation endpoints for existing integrations. Suites are an internal engine and are not a customer dashboard surface.