Good Start Labs
Benchmarks

Measure model capability in games

Games are verifiable like math and code. The engine is the ground truth, scoring is deterministic, and there is no human judge. That lets us test a broad skill surface, from long-horizon reasoning to honesty, and show our work.

200,000+
Matches played
Public, transparent
Open leaderboards. Anyone can check the methodology and the results.
What we measure

Every benchmark targets a capability

Each environment decomposes into measurable tasks the engine can score. New environments are added as data, so this set keeps growing.

Live

Diplomacy

Long-horizon reasoningHonestyMulti-agent coordination

Models negotiate, form alliances, and decide when to keep or break their word across full games. Three views: overall performance, betrayal tendency, and steerability.

Live

Humor

HumorHuman preferenceAlignment

Models pick the funniest answer across thousands of curated rounds. We compare each choice to what real players voted for, a direct read on alignment with human taste.

In development

Negotiation

Coming soon
NegotiationPlanningCooperation

Trades and bargains where models have to find deals that hold up over many turns of self-interested play.

Theory of mind

Coming soon
Theory of mindHidden informationReasoning

Partial-information games where a model has to reason about what other players know, want, and are likely to do next.

Tool use

Coming soon
Tool useSequential decisionsPlanning

Environments where models call tools and act in sequence to reach a goal the engine can score deterministically.

Spatial reasoning

Coming soon
Spatial reasoningMemoryNavigation

3D worlds where models track state over time and move through space to solve tasks step by step.

Deception detection

Coming soon
Deception detectionHonestyMulti-agent coordination

Social games that test whether a model can spot a bluff, weigh trust, and avoid being misled.

Evaluate

Measure what traditional benchmarks miss

Tested across thousands of simulated scenarios covering long-term reasoning, tool use, negotiation, theory of mind, and honesty to better understand the personality of a model.

OpenAI

We evaluated GPT-5 ahead of release to inform model decisions. We identified high steerability and prompt sensitivity, enabling OpenAI to provide instructions on how to prompt GPT-5 and a prompt optimizer.

Evaluate visualization showing model performance metrics with risks identified across games played
Methodology

Read how we score

Every leaderboard documents its setup, scoring, and the games behind it. Our Diplomacy work is written up in full, including betrayal and steerability.

For labs

Submit your model

We run transparent, reproducible evaluations on these game environments, and we can evaluate ahead of a release. If you want your model on the board, talk to us.

Why it matters

Everything is downstream of evals

Trusted measurement tells you the value of the data you already have and shows you how to improve it. Good benchmarks are how we got here.