Good Start Labs
Products

Games are training data for more capable models

What we provide
01

Learning environments

Interactive worlds with clean, verifiable rewards, where models learn by doing. The same worlds also generate valuable pretraining data.

02

Real human data

Text and image data from real people playing, captured at the source with clear provenance.

03

Benchmarks and evals

Public, transparent measurement of model capability, so every gain is real and reproducible.

04

Data generation

Synthetic and real human data at scale, including continuous human feedback from live games.

In productionArcee

Working with Arcee

  • Trinity Large trained on our Diplomacy environment, among others.
  • In our research, training Qwen 3 on our Diplomacy environment improved it on a customer-support benchmark (Tau2) and by over 10% on games like Hanabi and Wordle.
    Read about our collaboration
400Bparameters in Arcee's open Trinity Large family
Openweights anyone can download and run
Live data engine

Continuous human feedback at scale

  • Players provide actionable feedback on how to improve models
  • Games are continuous, scalable sources of quality data
  • Enables real world performance insights
Bad Cards

Our Bad Cards agent runs 20,000+ games weekly, getting live human feedback.

Bad Cards' 2 million+ users easily add AI agents to games with their friends.

Live feedback visualization showing human and agent interaction traces with data bars representing feedback collected across multiple evaluation cards
Why games

A training loop that compounds

Games are verifiable like math and code, with a broader skill surface. Put them in a loop: models train on games, get better, and play back in, generating new data and feedback every turn.

GamesTrainingA better model

Put games in your training loop

Tell us what your models need to learn. We will line up the environments, data, and signals that fit.