Run Your Own LLM Evaluation
Test any models with your prompts. Start free, no credit card required.
Coding Performance with 10 Evaluators — Run
Comprehensive evaluation of 2 language models across 1 system prompt with rigorous benchmarking and scoring criteria.
7.89
o3
5.00
Spread: 5.78 pts
157ms
Response time
80
8 total responses
Executive Insights
Key takeaways from this evaluation
Top Performer
o3
7.89
5.78 pts ahead of #2
Model Rankings
Ranked by overall performance score
o3
openai/o3
Derived from how judges ranked this response against the others — not an absolute quality rating.
Responses
4
Avg Latency
157ms
Cost
$0.0264
deepseek-r1
deepseek/deepseek-r1
Derived from how judges ranked this response against the others — not an absolute quality rating.
Responses
4
Avg Latency
—
Cost
$0.0277
Evaluator Consensus
How 10 evaluator models ranked the candidates via blind comparison
unanimous Agreement
All 10 evaluators agree on the top model
o3
Avg Rank
1.0
Range
#1
#1 Votes
10/10
Latency
157ms
deepseek-r1
Avg Rank
2.0
Range
#2
#1 Votes
0/10
Latency
—
gpt-5.4-mini
claude-sonnet-4.6
minimax-m2.7
gemini-3.1-flash-lite-preview
kimi-k2.5
deepseek-v3.2
grok-4.1-fast
qwen3.5-27b
mistral-small-2603
nova-2-lite-v1
Score Comparison
Visual comparison of all model scores
Performance by System Prompt
How each model performs across different evaluation contexts
Top Performer
o3
7.89
Performance by Test Prompt
Model results broken down by individual test prompts
| Test Prompt | Avg Score |
|---|---|
Javascript Function 2 responses | 5.00 |
Write an Interval Merge Function 2 responses | 5.00 |
Debug Python 2 responses | 5.00 |
Refactor Javascript 2 responses | 5.00 |
About This Evaluation
Methodology, criteria weights, and evaluation confidence
8
Total Responses
80
Total Evaluations