Keep going
Handing an agent a goal, guardrails, walking away, and reviewing results is becoming how work gets done. So which agent can you actually walk away from?
Alex Duffy · Good Start Labs · October 1, 2026
Last week a frontier physics calculation fell to an agentNINE LOOPSPhysicist Matt von Hippel dared AI companies to push a famous particle physics calculation from eight loops to nine. Two Anthropic physicists gave it to Claude Fable 5.1 and mostly told it to keep going, in a run that would have cost an end user one or two thousand dollars.Opens anthropic.com ↗ whose main instruction was “keep going.”
That’s auto researchAUTORESEARCHKarpathy’s autoresearch keeps only changes that beat its best result, about 100 experiments a night on one GPU, and in a two-day run it caught an oversight in his own code. Spin-offs aim the loop at test coverage, query speed, anything a command can score.Opens github.com ↗.
An agent gets a goal and a score. It tries something, checks, keeps what works, and tries again. Point it at AI research and you get recursive self-improvementRECURSIVE SELF-IMPROVEMENTModels doing the work that makes the next model better. Same loop, higher stakes.Opens anthropic.com ↗, the holy grail that AI labs are looking for.
Keep what works
Four real runs of the same loop. Up is always better: grey marks are tries, the ink line is the best so far, and blue is the takeaway. Hover or tap to read any mark.
NEVER STOP. Do NOT ask “should I keep going?”the agent's instructions
One night of Andrej Karpathy's autoresearch, from the run log its agent posted in the repo's discussions.
So the practical question is which models are any good at it? Which agent is still finding improvements at hour 40, and which one is going in circles? Where do their approaches differ? How do we measure this? Games are the perfect environment.
Since 2010, programmers have been hand-writing bots for StarCraft: Brood WarBROOD WAR BOTSAn open-source toolkit lets code play the 1998 game, and since 2011 the longest-running tournament has published every entrant’s source code, which is why a public library of human bots exists. They still play each other around the clock on an open ladder.Opens davechurchill.ca ↗ and pitting them against each other. That’s sixteen years of human expertise, written as code, with win rates attached.
StarSkirmishSTARSKIRMISHBuilt by Kai McPheeters. Good Start Labs is partnering with him on the reinforcement learning environment.Opens starskirmish.com ↗ turns this competition into an auto-research laboratory. An agent writes its own bot in C++. It plays practice games against human-written bots THE LADDERSStardust, PurpleWaveABananaBrain, LocutusBtscmoop2, Steamhammer, McRaveCHladhammer, PylonPuller, SkynetDDemo botsStardust, by Bruce Nielsen, is the top-rated human-written bot: seven titles since 2020, and 365–1 on the StarSkirmish Bench., reads what happened, and rewrites. The game is the grader. No AI-written bot has reached the top yet.
Saturday’s StarSkirmish BenchSTARSKIRMISH BENCHEach model gets one hour to write a StarCraft bot, then every bot plays a tournament: 100 is Stardust, 0 the weakest demo bot. GPT-6 Astra and Claude Opus 5.5 are functionally tied at the top; with GPT-6 Sol, they’re a clear step above every other model tested.Opens starskirmish.com ↗ measures how strong a bot each model can write in an hour. They create hypotheses, do research, experiment, learn from their mistakes, and iterate.
Today we’re kicking off a 48-hour livestream that extends the benchmark: GPT-6 Astra in Codex against Opus 5.5 in Claude Code. Which one is the better scientist?
The loop, live
Each agent runs this loop on its own bot. When a marker jumps a row, it just beat bots people spent years writing. Hover or tap a step for a real moment from the one-hour bench, or a rung to see who wrote it.
- = Top tier
- Cleared a tier
- First to clear
- Graded try
- Live position
Live from the StarSkirmish feed. The moments come from the one-hour StarSkirmish Bench.
We’re watching for a few things. How high each agent climbs, how fast, can they beat the record. How each one works: what it tries, when it drops an idea, what it takes from a loss. And how different their bots end up.
It’s a small experiment on purpose: WHAT COMES NEXTIf it goes well, new models drop in as they ship and the leader gets tried in other harnesses.. Model and harness come bundled, so a win says something about both.
Games make this easy to follow. You don’t need to read C++ to watch a bot lose its base. Every rung on the ladder is WHO WROTE THE LADDERStardust48,800Bruce Nielsen · since 2020Steamhammer†44,500Jay Scott · since 2016PurpleWave43,300Dan Gant · since 2017BananaBrain34,700Johan de Jong · since 2018Locutus†29,900Bruce Nielsen · 2018–2020McRave20,200Christian McCrave · since 2017Skynet16,700Andrew Smith · 2011–2013PylonPuller2,300Hao Pan · 2022Approximate lines of code, our count of each bot’s published source. † Includes code it was built on., so when an agent passes a bot, you know exactly whose work it just outplayed.
That comparison got more interesting last month. PlutoPLUTOtscmoo’s network learned only by playing itself, won the 2026 CoG competition by beating PurpleWave 4–1 in the final, and has since beaten pros StRyKeR 19–1 and Nesh 5–0 in exhibitions. tscmoo also wrote tscmoop2, a tier-B opponent on this ladder.Opens github.com ↗, a network trained purely by self-play, beat PurpleWave and Stardust, two of the strongest hand-written bots, and then beat professional players in exhibition matches. Our agents write the policy. Pluto is the policy.
We don’t know yet how good these agents are at open-ended research. We need to, and a game where anyone can watch the score is a good place to start finding out.
OUR LEARNING ENVIRONMENTStarSkirmish is built on our learning environment, which has many tasks besides this one, focused on algorithm optimization, strategy and long-horizon reasoning. Labs want what this one has: a grader that’s hard to fake and a wide gap between frontier models and the rest.. Every attempt gets scored by the game, so once agents start climbing, the same setup can reward the climbing and train it. If you want to build on that, or have a harness you want in the race, get in touch.
Watch live at twitch.tv/llmskirmish. Thanks to Kai for building StarSkirmish and being a creative, diligent partner.