Run Your Own LLM Evaluation
Test any models with your prompts. Start free, no credit card required.
Coding Performance with 10 Evaluators — Run
Comprehensive evaluation of 2 language models across 1 system prompt with rigorous benchmarking and scoring criteria.
5.76
qwen3.8-max-0902
5.00
Spread: 1.52 pts
2144ms
Response time
80
8 total responses
Executive Insights
Key takeaways from this evaluation
Top Performer
qwen3.8-max-0902
5.76
1.52 pts ahead of #2
Model Rankings
Ranked by overall performance score
qwen3.8-max-0902
qwen/qwen3.8-max-0902
Derived from how judges ranked this response against the others — not an absolute quality rating.
Responses
4
Avg Latency
3754ms
Cost
$0.0211
gpt-6-astra
openai/gpt-6-astra
Derived from how judges ranked this response against the others — not an absolute quality rating.
Responses
4
Avg Latency
534ms
Cost
$0.0475
Evaluator Consensus
How 9 evaluator models ranked the candidates via blind comparison
majority Agreement
7 of 9 evaluators agree on the top model
qwen3.8-max-0902
Avg Rank
1.2
Range
#1–2
#1 Votes
7/9
Latency
3754ms
gpt-6-astra
Avg Rank
1.8
Range
#1–2
#1 Votes
2/9
Latency
534ms
gpt-5.4-mini
gemini-3.1-flash-lite-preview
claude-sonnet-4.6
minimax-m2.7
kimi-k2.5
deepseek-v3.2
mistral-small-2603
qwen3.5-27b
nova-2-lite-v1
Score Comparison
Visual comparison of all model scores
Performance by System Prompt
How each model performs across different evaluation contexts
Top Performer
qwen3.8-max-0902
5.76
Performance by Test Prompt
Model results broken down by individual test prompts
| Test Prompt | Avg Score |
|---|---|
Javascript Function 2 responses | 5.00 |
Write an Interval Merge Function 2 responses | 5.00 |
Debug Python 2 responses | 5.00 |
Refactor Javascript 2 responses | 5.00 |
About This Evaluation
Methodology, criteria weights, and evaluation confidence
8
Total Responses
80
Total Evaluations