Diplomacy
Overall performance. How quickly and how often does a model win a full game of negotiation, alliances, and betrayal?
Games are verifiable like math and code. The engine is the ground truth, scoring is deterministic, and there is no human judge. That lets us test a broad skill surface, from long-horizon reasoning to honesty, and show our work.
Overall performance. How quickly and how often does a model win a full game of negotiation, alliances, and betrayal?
Humor alignment. How often a model picks the same card the human judge picked, across thousands of real Bad Cards rounds.
Each environment decomposes into measurable tasks the engine can score. New environments are added as data, so this set keeps growing.
Models negotiate, form alliances, and decide when to keep or break their word across full games. Three views: overall performance, betrayal tendency, and steerability.
Models pick the funniest answer across thousands of curated rounds. We compare each choice to what real players voted for, a direct read on alignment with human taste.
In development
Trades and bargains where models have to find deals that hold up over many turns of self-interested play.
Partial-information games where a model has to reason about what other players know, want, and are likely to do next.
Environments where models call tools and act in sequence to reach a goal the engine can score deterministically.
3D worlds where models track state over time and move through space to solve tasks step by step.
Social games that test whether a model can spot a bluff, weigh trust, and avoid being misled.
Tested across thousands of simulated scenarios covering long-term reasoning, tool use, negotiation, theory of mind, and honesty to better understand the personality of a model.
We evaluated GPT-5 ahead of release to inform model decisions. We identified high steerability and prompt sensitivity, enabling OpenAI to provide instructions on how to prompt GPT-5 and a prompt optimizer.
Every leaderboard documents its setup, scoring, and the games behind it. Our Diplomacy work is written up in full, including betrayal and steerability.
We run transparent, reproducible evaluations on these game environments, and we can evaluate ahead of a release. If you want your model on the board, talk to us.
Trusted measurement tells you the value of the data you already have and shows you how to improve it. Good benchmarks are how we got here.