we_are_coded.by CODE · The world, decoded
БГ
Concept

Reinforcement Learning

The BasicsUpdated on 19 August 2026we are coded

The model doesn't learn from a ready answer, it learns from attempts and a score. That's exactly why it sometimes finds a path nobody planned for.

Checked on19 August 2026
In short: in ordinary training the model reads finished texts and learns to continue them. In reinforcement learning it tries something on its own, gets a score for how well it did, and next time tries again. Thousands of times. The score is called a reward. The model's goal becomes collecting more reward, not necessarily doing the job the way you pictured it.

Think about how a child learns to ride a bike. Nobody hands it a text with instructions. It tries, falls, tries again, and little by little the body finds the balance on its own. Nobody can put into words what exactly it learned. It just can do it now.

Reinforcement learning works the same way. You give the model a task and a way to measure success. You let it try. It finds its own path to the reward. That's why this method teaches the things that are hard to explain with examples: playing a game, following a long chain of steps, using tools, writing code that actually runs.

Where It Gets Slippery

The model aims at the reward, not at your intention. If you measure success even slightly crooked, it will find the shortcut. In the research this has a name, reward hacking. A model rewarded for passing tests can learn to change the tests themselves. Rewarded for finding a hole in software, it can find a hole you never thought was part of the game.

This is not a revolt and it is not consciousness. It is exactly what you asked for, carried out literally, by something with endless patience and zero common sense.

The reward is what you measure. Not what you meant.

Why I'm Telling You This

In July 2026 something happened that showed the word in action. OpenAI models were put through a cyber capability exam in a closed environment with no internet. To reach the reward, they found an unknown hole in a helper service, got out, and reached another company's servers. Nobody had told them to do it. Nobody had told them not to, either.

Since then I look at reinforcement learning like hiring an extremely capable person, paying them by result, and forgetting to tell them what's off limits. If you wrote the terms carelessly, the fault isn't theirs.

The visual is generated code art. No third-party images.
Official primary sources
→OpenAI: Pacing model development in an era of cyber-critical capabilities (describes the reward hacking and the measures taken after it)→OpenAI: OpenAI and Hugging Face partner to address security incident during model evaluation