What if we’ve misunderstood what AI actually learned from play?
Mia Sundstrom recently wrote an engaging Substack article arguing that the story of modern AI is importantly a story about play.
She recounts how DeepMind achieved many of its breakthroughs by setting up systems to learn through games. Rather than programming solutions directly, the researchers created environments in which AI could explore, receive feedback and gradually improve through repeated experience.
DeepMind then transferred the method into fields such as protein folding. It’s a fascinating reminder that intelligence often develops through exploration rather than instruction.
One detail, though, puzzled me. In describing reinforcement learning, she suddenly introduces “failure” as though it were an essential part of the explanation:
It experimented.
It failed.
It iterated.
Once she hits upon that idea, it dominates her explanation, even though it isn’t actually what the AI is doing.
A reinforcement learner has no concept of failure. It receives information. Losing a game is one kind of feedback. Winning is another. Every move changes the system’s expectations. Every consequence helps shape future behaviour. The process is organised around difference, not failure.
The distinction matters because “failure” keeps appearing as one of the stock explanations of learning: fail fast, celebrate failure, learn from failure. The implication is that failure itself drives improvement.
But does it? To become good at anything, you eventually have to get it right at least once. You may fail many times on the way there, or perhaps not at all. That varies enormously. But once something works, the important question is: What happened there that I can do again?
Failure tells us, at best, that something needs changing, perhaps eliminating one wrong move. Success tells us what to repeat. That’s an asymmetry that’s easy to overlook.
Learning may involve both kinds of information, but only one contains the recipe for success. Learning needs more than a record of unsuccessful attempts. It has to recognise the conditions, actions and relationships that produced something worth repeating, then refine and transfer them.
We see learning without meaningful failure all the time. A child acquires a new word simply by hearing it and using it successfully. A dancer copies a partner’s movement.
When I play the guitar, I hear a phrase and can reproduce it first time, (on a good day). A cook adds a squeeze of lemon to a sauce and immediately knows, “That’s it.”
It’s simply wrong to say ‘we can’t learn without making mistakes‘, even though we often make mistakes while learning.
For skilled performance, when I play tennis I can incorporate a coach’s suggestion to toss the ball a fraction further forward to produce a faster serve. An improviser notices when a scene suddenly comes alive and internalises what made that possible.
A reinforcement learner updates its expectations because the reward signal has changed. Like a player, it’s exploring the conditions, actions and relationships that produce something worth doing again.
So the learning sequence is better described as:
Explore.
Notice the consequences.
Drop what doesn’t help.
Recognise what works.
Repeat and refine it.
None of this reduces the importance of play. If anything, it strengthens the case.
Play creates a space for exploration, where consequences arrive quickly and it’s easier to notice useful patterns, often under minimal pressure.
That seems closer both to the science of reinforcement learning and to the lived experience of human play. It also fits well with Solutions Focus, where attention is directed towards identifying, repeating and amplifying what works, rather than treating failure as the primary teacher.
If you’d like to discuss how we can help strengthen your facilitation practice, team collaboration or workshop design, book an exploratory conversation here:
https://calendly.com/paulzjackson/30min
Best wishes
Paul & the team

