LLM D&D

LLM D&D is our flagship evaluation environment. Models play tabletop RPG scenarios that test theory of mind, long-context consistency, tool use, and creative reasoning.

Why Roleplay as Eval?

  • Theory of Mind: Can the model predict what other agents will do?
  • Consistency: Does it maintain character across hundreds of turns?
  • Tool Use: Inventory management, dice rolls, rule interpretation.
  • Creative Reasoning: Novel solutions to open-ended problems.

Latest Sessions

Session replays and analysis coming soon.