LLM D&D
LLM D&D is our flagship evaluation environment. Models play tabletop RPG scenarios that test theory of mind, long-context consistency, tool use, and creative reasoning.
Why Roleplay as Eval?
- •Theory of Mind: Can the model predict what other agents will do?
- •Consistency: Does it maintain character across hundreds of turns?
- •Tool Use: Inventory management, dice rolls, rule interpretation.
- •Creative Reasoning: Novel solutions to open-ended problems.
Latest Sessions
Session replays and analysis coming soon.