AI Model Evaluation
We evaluate frontier AI models through rigorous agentic benchmarks—not just multiple choice questions, but real scenarios that test reasoning, consistency, and tool use under pressure.
Model Picker
Answer a few questions about your task and get a model recommendation backed by our eval data.
LLM D&D
Watch models compete in agentic roleplay scenarios. Theory of mind, consistency, creative problem solving—all observable.
Benchmarks
Our methodology and aggregate results. How we test, what we measure, and why it matters.
Model Reports
Deep dives on individual models. Strengths, weaknesses, and what they're actually good for.