Announcement
Introducing Steve
The first AI to teach itself to beat Minecraft, and a step towards RSI.
Steve is an agent system that gives a language model an environment for developing its own reliable software, testing its understanding against the world, and carrying what works into later attempts, until it achieves goal-oriented mastery.
The goal is a system that can build and run a capable, working 'body' for any environment.
It starts with a stock pretrained model, a fresh world, a high-level API (the AI's 'keyboard and mouse'), and our proprietary tools for externalizing cognition into muscle memory and useful patterns-of-thought.
It develops everything else itself, and System One model Jev takes over at 92% lower cost once those capabilities are built.
Why it matters
Steve is a model substrate we believe has profound implications for how much useful work we can get out of language models in robotics, agentic coding, and long-horizon automation.
For games, this opens a possibility we care deeply about: characters that can develop new abilities through experience, with artists shaping where and how that development belongs in the world.
The next great games will look more like worlds that are alive, where characters, creatures, and ecosystems respond to what the player does and to each other. Soon, in-game towns, civilizations, and whole ecosystems will increasingly look and feel like our own, or like a great story.
The economics matter as much as the capabilities: a system that needs frontier-model reasoning for every action faces a very different path to a playable experience than one whose learned procedures can run cheaply. Open-weight models closed the gap with the closed frontier from ~13 points to ~6 in a year. Inference now costs 6x to 60x less per token than the closed labs charge for the same job.
A model’s useful reasoning does not have to disappear at the end of a response. It can become working software: calculations, controllers, diagnostic tools, tests, and procedures for solving whole classes of problems.
Once a stronger model has developed a capable system, a substantially less capable model can take over using the capabilities already built. That opens a different way of deploying intelligence: invest stronger reasoning upfront to develop and validate useful capabilities, then use much cheaper computation to operate them.
We should not have to keep paying the full cost of figuring something out after an agent has already learned how to do it.
Research
Four papers in the coming weeks.
- [1.]
Autonomous capability development.
- [2.]
Patterns of thought: agent building their own learning and investigation tools.
- [3.]
Muscle memory: agents designing the accumulation & transfer of software capabilities across longer undertakings.
- [4.]
Stronger-to-weaker model handoffs.
We’ll present the methods, comparisons, costs, experimental boundaries, and limitations alongside the results.
Over the next month we’ll also announce a mass-scale, publicly playable intelligent-NPC playtest, and partnerships with teams behind some of the last decade’s most popular games, both AAA and indie.