The job: build the machinery that turns real games into training grounds for AI.
You wrap live games into deterministic, scalable RL environments. You build the pipelines that turn gameplay into training data. Your environments are live games with tens of millions of real players, not sandboxes. The research runs on your systems.
What You'll Do
- Design, build, and scale RL environments around real games from our publisher partners.
- Build the system that mass-creates LLM harnesses across a publisher's entire game library. Wiring up one game should take hours, not weeks.
- Build high-speed game engines for expert-agent training. Rollout throughput sets the ceiling on how good our experts get.
- Ship public leaderboards and benchmarks where models and humans compete in real games.
- Build pipelines that capture human data from live games: ranked preferences, chat, play traces. Turn it into training-ready datasets.
- Ship agent infrastructure that runs in production games with tens of millions of real players.
- Make experiments fast. The researcher's iteration speed is your metric.
What We Look For
- Has shipped production ML systems end to end, largely alone.
- Deep Python and systems skills. RL, simulation, or environment-building exposure.
- Has built LLM agent harnesses with tool-use APIs or agent SDKs.
- Evidence of speed. Show us what you shipped and how long it took.
- Benchmark profile (our last MLE hire): built a Dota 2 RL environment with Madrona and PufferLib, productionized models for 150M+ users, shipped agent platforms from scratch.
Comp and Location
- $165K-$250K base plus significant equity.
- Remote-first with regular in-person gatherings. Must be in the US or Canada.
- Apply: apply.goodstartlabs.com