Benchmarks
Our evaluation methodology focuses on agentic capabilities—not just knowledge retrieval.
Methodology
Traditional benchmarks test what models know. We test what models can do. Our scenarios require sustained reasoning, memory, tool coordination, and adaptation to novel situations.
Metrics
Theory of Mind
Predicting other agents' actions and intentions
Consistency
Maintaining character/state across long contexts
Tool Use
Correctly invoking and chaining tool calls
Creative Reasoning
Novel solutions to open-ended problems
Instruction Following
Adherence to complex multi-step instructions
Latest Results
| Model | ToM | Consistency | Tool Use | Creative |
|---|---|---|---|---|
| 92% | 94% | 98% | 91% | |
GPT-5.2 | 89% | 87% | 95% | 93% |
| 85% | 82% | 91% | 88% | |
| 88% | 79% | 89% | 90% | |
DeepSeek R1 | 84% | 86% | 87% | 85% |
Scores are percentages based on aggregate performance across multiple scenario types. Last updated: January 2026.