Evaluation Assistant (AI)
Build evaluations conversationally with the AI assistant.
The Evaluation Assistant is an AI-powered chat interface that helps you design and launch evaluations conversationally. Instead of manually configuring every setting, describe what you want to test and the assistant recommends models, generates prompts, and sets up the configuration for you.
Getting Started
Navigate to Suites and click Create with AI. This opens a chat interface with a configuration sidebar.
Start by describing your use case — for example:
- "I want to compare coding models for Python refactoring tasks"
- "Help me evaluate customer support chatbots on a $50 budget"
- "Which models are best for structured data extraction from medical notes?"
What the Assistant Can Do
The assistant has access to five specialized tools:
- Search models — browse your workspace's available models by provider, tier, budget, or capability
- Generate prompts — create system prompts and test prompts tailored to your use case
- Save to library — persist generated prompts to your workspace library for reuse
- Configure the run — update the sidebar with recommended models, criteria, and settings
- Estimate credits — calculate the cost before you launch
Progress Tracker
The sidebar shows a 4-phase progress indicator:
- Define goal — describe what you're evaluating
- Prompts selected — system and test prompts are configured
- Models selected — generator and evaluator models are chosen
- Ready to launch — all required fields are set
Launching
Once your configuration is complete, you have two options:
- Start evaluation — run immediately as a one-off evaluation without saving a suite
- Save suite & run — save the configuration as a reusable suite, then trigger the first run
Tip: Your conversation is saved in your browser session. If you navigate away and come back, you can pick up where you left off.