Measure what matters.Get paid for it.
How we helped a game publisher launch its own benchmark, turn its data into revenue, and improve a language model with an autonomously trained expert.
Adapted from an essay at Every · July 2026
The same models that make novel discoveries in math and science lose 90 percent of their Gin Rummy games against casual players.
Ten hands of Gin Rummy, each against a casual human player. Each icon shows who won the hand: robot for the model, person for the human. Frontier results from Arkadium's Game Lab; expert model at full strength.
Over the past year at Good Start Labs, we've built benchmarks, trained agents, and helped game publishers operationalize and monetize their data for the lab market. Arkadium is one publisher with hundreds of games played by tens of millions of players. We supported its recent launch of Game Lab, a public leaderboard scoring how well frontier models play simple games, in partnership with Meta and DeepMind. Arkadium set a clear goal: give its players a good game against AI. Together we built the benchmarks and evaluated them against real users. We discovered that frontier models absolutely could not deliver.
That's because these models are shaped by what they've seen: lots of math, and almost no Gin Rummy. (Opus can't beat the daily crossword, either.) Popular AI models may not have experience with whatever you work on all day.
But defining the goal revealed options for Arkadium beyond a large language model, because different forms of AI work well for different goals. Instead of an LLM, we trained an expert model with 4.6 million parameters in an 18-megabyte file that runs on a regular CPU. At full strength, it beats human players about 90 percent of the time, and the economics are as lopsided in its favor:
Model scale and serving cost, side by side
The circles show parameter scale. Costs update with traffic; bars use a log scale.
One well-defined job did not require the largest possible model. It revealed a smaller model that performed better and cost dramatically less.
An AI improving an AI
Our expert model was the product of a custom autonomous Autoresearchloop. With Claude Opus as the researcher: it monitored live training runs, benchmarked them against a diverse field of opponents, critiqued its own hypotheses, and closed every cycle by proposing exactly two research-informed experiments: one iterative, one creative. We deployed both, every time, building on top of the winner.
The frontier model runs the lab
It designs experiments, trains candidate experts, benchmarks them, and keeps the best.
Performance curve is illustrative. Opus designed the experiments; many expert candidates trained inside the environment; a human approved each cycle before deployment.
You build the learning environment
The frontier model never taught the expert how to play. What we were after was a world where learning could happen on its own: Gin Rummy, fully simulated, with rules, opponents, and a score you can trust. With that in place, the frontier model could send a small expert through it, watch the results, and change the experiment, over and over. Each pass came out stronger. When one expert was clearly the best, it went from student to coach: the frontier model took its own turn in the same environment, guided by the expert's ratings, and came out stronger at the game—and some of what it learns may travel to domains like math and finance. (A small hedge: RL fine-tuning can move math benchmarks even with spurious rewards, signals that carry no real information, so we hold those gains loosely.)
Two trips through one environment
First to find the best expert. Then to make the frontier model stronger.
The frontier model designs the experiment, sends a small expert through the game, watches the results, and adjusts. Every pass levels the expert up.
Then the frontier model takes its own turn, with the best expert as coach, and comes out stronger.
One well-designed environment can level up any model.
One goal pays twice
Meanwhile, frontier labs are under pressure to improve at economically valuable work like finance, life sciences, and general reasoning. They're paying for data that helps them get there. For Arkadium, that "good game" goal paid twice: a better experience for its players and a new revenue line in the form of selling anonymous data to labs for millions. Frontier models performed poorly at games; Arkadium's well-structured gameplay data from "good" games with "good" human players was what the labs needed to improve. As Arkadium continues to sell that anonymized data to frontier labs, its models will learn from it and apply that intelligence. (Maybe soon, Opus will be able to beat that crossword.)
Define one outcome well enough and it can pay twice: once in the product, then again in the data produced while achieving it.
Reddit, Shutterstock, and News Corp have already turned their data into hundreds of millions of dollars in recurring revenue. Most companies outside the labs' competitive focus could benefit from selling them data. One major caveat: companies whose data is their product, like Figma or Cursor, must protect that IP. If the lab you'd sell to is entering your business, retaining a data edge is essential. The full Figma, Cursor, and pharma stories are in the essay at Every.
Nobody knows yet how defensible the data market is long term, only that the pot promises to be large. But choosing the right goal for AI, measuring progress, and valuing what you learn will only grow more important as AI becomes table stakes for most companies. Defining and measuring "good" is emerging as the next stage of AI adoption. It never ends, and it is becoming a requirement to compete.
We help companies build benchmarks, train expert models, and operationalize their data for the lab market.