What is PeerLM?
Overview of the platform and core workflow.
PeerLM answers one question about your production LLM: should you switch, route, or hold — and can you prove it on your own traffic. Instead of hand-testing prompts in chat interfaces, it samples what your product actually sends and compares candidates against the response your current model really gave.
Core Workflow
- Set up a Monitor — Name the model serving production today, choose a source and task policy, and preview the candidate policy. Saving remains manual-first.
- Connect traffic — Point an OpenTelemetry exporter at PeerLM, connect a tracing tool, or upload an export. A saved connection is not counted as received traffic.
- Runs accumulate — Each Run replays candidates on real sampled prompts and seats a rotating panel of judges.
- Read the verdict — The Monitor holds a living switch / route / hold call, with the receipt behind it. It arrives in Slack, your webhook, or back on the original trace.
Prompt CI runs the same comparison on every deploy, checking a prompt change against real traffic before it ships.
Who is it for?
- Product teams choosing which LLM to integrate into their product.
- AI engineers testing prompt changes across models before deploying.
- Procurement teams comparing vendors with objective, shareable reports.
What you get
- 200+ models across providers (OpenAI, Anthropic, xAI, and more), priced from dated catalog prices
- Monitors that turn production traffic into a living switch / route / hold verdict
- A five-judge panel per item, rotated from a pool frozen on the Monitor — PeerLM picks the judges
- Triggers that start a new Run when a model ships, a price changes, or traffic drifts
- Prompt CI: the same comparison on every deploy, before a prompt change ships
- Verdicts delivered to Slack, a signed webhook, or back onto the original trace
- Shareable public reports, with the methodology attached