Our Research
Peer-reviewed papers and technical write-ups from our work.
What a Railroad Game Taught a Model About Finance
We trained a 30B model to play 1830, a brutal board game about 19th century railroad barons. That model seemed to get better at real SEC-filings research — and an independent AI research agent found the same fingerprint on a test we never ran.

Where an Agent’s Intelligence Lives
OpenAI tripled one benchmark score with two harness settings. In our game agents, the strongest results came when the model and its skill system learned together.

Measure What Matters. Get Paid for It.
How we helped a game publisher launch its own benchmark, turn its data into revenue, and improve a language model with an autonomously trained expert.

Data-Centric Interpretability for LLM-based Multi-Agent Reinforcement Learning
We introduce Meta-Autointerp, a framework that uses LLMs to automatically generate and validate interpretability hypotheses about learned features in multi-agent RL. Our results show that data-centric methods can surface meaningful behavioral patterns that traditional approaches miss.

We Trained an AI on a Board Game. It Became a Better Customer Support Agent.
Games teach transferable skills, to humans and AI alike.

Opening the Door: Democratizing Diplomacy
We created the first Diplomacy environment where even small models can play full games. Our goal was to give people tools to understand how different AI models make decisions.

We Made Top AI Models Compete in a Game of Diplomacy. Here’s Who Won.
We launched our first project AI Diplomacy. Different models revealed their true character: some betrayed without hesitation, while others, like Claude, chose principles over victory. We made AI behavior visible through Diplomacy, for 45,000 unique Twitch viewers, many experiencing AI for the first time.

Interested in working together?
We partner with labs and researchers exploring what games teach AI.