DocsGetting StartedUnderstanding Your Results

Understanding Your Results

Read a Run report and the Monitor's living verdict.

There are two things to read: the Run report, which is the receipt for one comparison, and the Monitor's living verdict, which is what all its Runs add up to. The verdict lives on the Monitor; the evidence lives on the Run.

The verdict

Switch means a candidate held quality within the Monitor's decision contract and costs less. Route means it held on some workload categories but not others, and the per-task routing table says which. Hold means the evidence does not support a change — either the candidate lost quality, or there was not enough evidence to tell. A hold for lack of evidence is labelled inconclusive rather than dressed up as a loss.

Decision

The top of a Run report states the recommendation, the evidence strength behind it, and the projected saving as a 0.5×–2× range stamped with the date the prices were read. When savings cannot be resolved — an unpriced model, missing token counts — the figure is absent, never zero.

Quality & cost summary and candidate comparison

Per candidate: wins, ties and losses against the control, the share of judged items where quality held, observed cost per request, and latency when both sides were measured the same way. Latency is only compared when the measurement scope matches; a missing number is a missing number, not a regression.

Do the judges agree?

How often the panel reached a determinate verdict and how often the seated judges agreed with each other. When your team has done blind spot-checks, this also shows how often the panel agreed with them — a readout, not an input; adjudications never change a verdict.

Item-level evidence

Every judged item is inspectable: the prompt, both outputs, the panel seated for it, each judge's vote and its one-sentence reason. Where the candidate loses and Failure patterns group those losses by failure tag, so a hold tells you what kind of prompt the cheaper model gets wrong.

Methodology

The bottom of every report states what was measured: sample size, coverage of replayable traffic, control mode, judge pool, price date. A shared report carries the same section, so anyone you send it to can check the claim against its basis.

Actions

  • Mark reviewed / Mark adopted — record that a person looked at the recommendation, and that you acted on it. Only reviewed evidence counts toward the dashboard's opportunity total, and adopting happens on the report, never from Slack.
  • Share — create a public link (see Sharing & Exporting)
  • Start a Run — the next comparison on the same Monitor, uses one pooled Run.