Run Your Own LLM Evaluation
Test any models with your prompts. Start free, no credit card required.
Coding Performance with 10 Evaluators — Run
Comprehensive evaluation of 2 language models across 1 system prompt with rigorous benchmarking and scoring criteria.
6.84
o3
5.00
Spread: 3.68 pts
—
Response time
80
8 total responses
Executive Insights
Key takeaways from this evaluation
Top Performer
o3
6.84
3.68 pts ahead of #2
Model Rankings
Ranked by overall performance score
o3
openai/o3
Derived from how judges ranked this response against the others — not an absolute quality rating.
Responses
4
Avg Latency
—
Cost
$0.0264
grok-4
x-ai/grok-4
Derived from how judges ranked this response against the others — not an absolute quality rating.
Responses
4
Avg Latency
—
Cost
$0.0925
Evaluator Consensus
How 10 evaluator models ranked the candidates via blind comparison
majority Agreement
7 of 10 evaluators agree on the top model
o3
Avg Rank
1.3
Range
#1–2
#1 Votes
7/10
Latency
—
grok-4
Avg Rank
1.7
Range
#1–2
#1 Votes
3/10
Latency
—
gpt-5.4-mini
gemini-3.1-flash-lite-preview
claude-sonnet-4.6
minimax-m2.7
deepseek-v3.2
grok-4.1-fast
mistral-small-2603
qwen3.5-27b
kimi-k2.5
nova-2-lite-v1
Score Comparison
Visual comparison of all model scores
Performance by System Prompt
How each model performs across different evaluation contexts
Top Performer
o3
6.84
Performance by Test Prompt
Model results broken down by individual test prompts
| Test Prompt | Avg Score |
|---|---|
Javascript Function 2 responses | 5.00 |
Write an Interval Merge Function 2 responses | 5.00 |
Debug Python 2 responses | 5.00 |
Refactor Javascript 2 responses | 5.00 |
About This Evaluation
Methodology, criteria weights, and evaluation confidence
8
Total Responses
80
Total Evaluations