OpenAI researchers released an open-source algorithm earlier this week that lets artificial intelligence agents learn from their own failures by reframing them as unintended successes. The method, called Hindsight Experience Replay (HER), addresses a long-standing problem in reinforcement learning: how to give an AI a useful learning signal when it never reaches its intended goal.
In a blog post, the researchers described the core idea behind HER as something humans do intuitively. "Even though we have not succeeded at a specific goal, we have at least achieved a different one," they wrote. "So why not just pretend that we wanted to achieve this goal to begin with, instead of the one that we set out to achieve originally?"
Under conventional reinforcement learning, an AI agent typically receives a reward only when it completes a specified task. A second approach, known as reward shaping, gives partial credit based on how close the agent gets to the goal. According to IEEE Spectrum, the first method can stall learning because the agent either succeeds or receives nothing, while the second can be difficult to implement effectively.
How HER reframes failure
HER treats every failed attempt as a virtual goal. If an AI agent tries to move an object to one location but ends up moving it elsewhere, HER retroactively treats the actual outcome as the goal the agent was trying to achieve. This substitution provides a learning signal even when the original objective was missed.
"By doing this substitution, the reinforcement learning algorithm can obtain a learning signal since it has achieved some goal; even if it wasn't the one that you meant to achieve originally," the OpenAI researchers wrote. "If you repeat this process, you will eventually learn how to achieve arbitrary goals, including the goals that you really want to achieve."
OpenAI demonstrated HER using its Fetch simulation, in which a robotic arm learns to manipulate objects. The algorithm does not make learning easy in all settings. Matthias Plappert, an OpenAI researcher, told IEEE Spectrum that "learning with HER on real robots is still hard since it still requires a significant amount of samples."
The research builds on OpenAI Baselines, the organization's existing reinforcement learning methods. By allowing agents to extract value from unsuccessful attempts, HER aims to speed up learning and improve its quality compared with sparse-reward approaches that offer no feedback until success.
The technique does not eliminate the challenges of training AI, particularly in physical systems where data collection remains costly. But in simulation, HER showed it can encourage agents to learn from mistakes in a way that resembles human trial-and-error, without the frustration that humans experience along the way.
Editorial Notes