Benchmarks

Our evaluation methodology focuses on agentic capabilities—not just knowledge retrieval.

Methodology

Traditional benchmarks test what models know. We test what models can do. Our scenarios require sustained reasoning, memory, tool coordination, and adaptation to novel situations.

Metrics

Theory of Mind
Predicting other agents' actions and intentions
Consistency
Maintaining character/state across long contexts
Tool Use
Correctly invoking and chaining tool calls
Creative Reasoning
Novel solutions to open-ended problems
Instruction Following
Adherence to complex multi-step instructions

Latest Results

ModelToMConsistencyTool UseCreative
Claude Opus 4.5 logoClaude Opus 4.5
92%94%98%91%
GPT-5.2
89%87%95%93%
Gemini 3 Pro logoGemini 3 Pro
85%82%91%88%
Grok 4 logoGrok 4
88%79%89%90%
DeepSeek R1
84%86%87%85%

Scores are percentages based on aggregate performance across multiple scenario types. Last updated: January 2026.