DocsMonitorsDecision contracts & readiness

Decision contracts & readiness

Understand roles, evidence strength, traffic eligibility, and spend caps.

Control, candidates, and judges

The control is the exact catalog model serving production. Candidates are alternatives replayed on the same inputs. Judges compare each pair; they are not deployment targets. PeerLM excludes models from the vendor under test and seats the full panel from a pool frozen on the Monitor.

Replayability

A record is eligible only when PeerLM can reproduce the request without inventing missing turns, tool state, or media. Images, audio, files, and unresolved tool calls are excluded from the text-only replay. A flattened transcript can also be ineligible because roles and tool boundaries cannot be recovered from prose.

The readiness panel separates structural coverage, trace identity, and tool-schema coverage instead of presenting one vague health score. File uploads can be valid replay material without trace IDs, but their verdicts cannot be written back to an observability record.

Observed-control eligibility

Observed mode requires zero retention to be off, the current classifier version, enough replayable records to fill the requested sample, an exact incumbent model match, one known system-prompt boundary, and at least 90% usable production-output coverage. If any check fails, the Run is blocked and replayed control is offered explicitly; PeerLM never changes the mode for you.

When reingest is required

Older corpora classified before readiness version 2 must be reingested. Appending records or recounting old rows cannot recreate structural data that was discarded. Completed Run snapshots remain unchanged.

Decision contract

The contract freezes the variable under test, allowed material and critical-category loss, savings threshold, production volume, and active critical category IDs before the Run. Uncertain evidence is inconclusive. Proposed categories can be reported but do not become critical identities until a person activates them.

Evidence strength and human review

Evidence strength describes sample sufficiency, execution cleanliness, and determinate panel agreement. Blind human spot-checks are append-only gold evidence. The report can say how often the panel agreed with your reviewers, but adjudications do not recalibrate judges, alter completed decisions, or authorize a switch.

Sample and cost gates

A standard Run uses 150 prompts and never fewer than 30. Critical categories reserve their required evidence before the remaining sample is distributed. If the corpus, plan allowance, or hosted-generation ceiling cannot fund the requested configuration, PeerLM rejects it rather than silently shrinking the sample.