Measure model capability in games
Games are verifiable like math and code. The engine is the ground truth, scoring is deterministic, and there is no human judge. That lets us test a broad skill surface, from long-horizon reasoning to honesty, and show our work.
Every benchmark targets a capability
Each environment decomposes into measurable tasks the engine can score. New environments are added as data, so this set keeps growing.
Diplomacy
Models negotiate, form alliances, and decide when to keep or break their word across full games. Three views: overall performance, betrayal tendency, and steerability.
Humor
Models pick the funniest answer across thousands of curated rounds. We compare each choice to what real players voted for, a direct read on alignment with human taste.
In development
Negotiation
Coming soonTrades and bargains where models have to find deals that hold up over many turns of self-interested play.
Theory of mind
Coming soonPartial-information games where a model has to reason about what other players know, want, and are likely to do next.
Tool use
Coming soonEnvironments where models call tools and act in sequence to reach a goal the engine can score deterministically.
Spatial reasoning
Coming soon3D worlds where models track state over time and move through space to solve tasks step by step.
Deception detection
Coming soonSocial games that test whether a model can spot a bluff, weigh trust, and avoid being misled.
Measure what traditional benchmarks miss
Tested across thousands of simulated scenarios covering long-term reasoning, tool use, negotiation, theory of mind, and honesty to better understand the personality of a model.
We evaluated GPT-5 ahead of release to inform model decisions. We identified high steerability and prompt sensitivity, enabling OpenAI to provide instructions on how to prompt GPT-5 and a prompt optimizer.
Read how we score
Every leaderboard documents its setup, scoring, and the games behind it. Our Diplomacy work is written up in full, including betrayal and steerability.
Submit your model
We run transparent, reproducible evaluations on these game environments, and we can evaluate ahead of a release. If you want your model on the board, talk to us.
Everything is downstream of evals
Trusted measurement tells you the value of the data you already have and shows you how to improve it. Good benchmarks are how we got here.