The job: prove that games make models better.
You build expert agents that play games at a superhuman level, then distill that intelligence into general models. You isolate what game data and game-based RL do to model capability, and you make the transfer repeatable.
What You'll Do
- Train expert game agents and distill their intelligence into general language models.
- Design dense reward signals and shape what gets learned. Our first transfer result came from sparse rewards and expensive rollouts. You help improve that process.
- Build training systems that target a benchmark: generate tasks that mirror its format and use the game engine as the verifier.
- Run controlled SFT (supervised fine-tuning) and RL experiments that measure how game environments change model performance.
- Design rewards and evaluations for multi-agent games: negotiation, long-horizon planning, cooperation, deception.
- Build public evaluations and benchmarks that show what games measure and math or code benchmarks miss.
- Publish papers, technical reports, and blog posts. Your results will help carry our research brand.
- Feed findings back into environment design with the engineering team.
What We Look For
- The profile we want most: you built a superhuman game agent, in the spirit of AlphaStar, AlphaGo, or OpenAI Five. Any company, any game, show us the agent and what it beat.
- Down for the mission. You love games and you view them as serious training grounds for intelligence.
- Hands-on RL post-training experience: PPO, GRPO, or RLVR.
- Evaluation design experience on language models.
- A public research record: papers, models, or benchmarks that other people used, cited, or built on.
- Designs small, fast experiments and pulls real conclusions from messy results.
Comp and Location
- $220K-$300K base plus significant equity.
- Remote-first with regular in-person gatherings. Must be in the US or Canada.
- Apply: https://apply.goodstartlabs.com