Games are training data for more capable models
Learning environments
Interactive worlds with clean, verifiable rewards, where models learn by doing. The same worlds also generate valuable pretraining data.
Real human data
Text and image data from real people playing, captured at the source with clear provenance.
Benchmarks and evals
Public, transparent measurement of model capability, so every gain is real and reproducible.
Data generation
Synthetic and real human data at scale, including continuous human feedback from live games.
Working with Arcee
- Trinity Large trained on our Diplomacy environment, among others.
- In our research, training Qwen 3 on our Diplomacy environment improved it on a customer-support benchmark (Tau2) and by over 10% on games like Hanabi and Wordle.Read about our collaboration
Continuous human feedback at scale
- Players provide actionable feedback on how to improve models
- Games are continuous, scalable sources of quality data
- Enables real world performance insights
Our Bad Cards agent runs 20,000+ games weekly, getting live human feedback.
Bad Cards' 2 million+ users easily add AI agents to games with their friends.
A training loop that compounds
Games are verifiable like math and code, with a broader skill surface. Put them in a loop: models train on games, get better, and play back in, generating new data and feedback every turn.
Put games in your training loop
Tell us what your models need to learn. We will line up the environments, data, and signals that fit.