Why almost any goal wants resources
A strange and important idea: whatever a system is trying to do, a few sub-goals are useful for nearly all of them. That is where the long-term worry comes from.
In 60 seconds
Why almost any goal wants resources
A strange and important idea: whatever a system is trying to do, a few sub-goals are useful for nearly all of them. That is where the long-term worry comes from.
| Convergent sub-goal | Why it helps any goal | How it could look |
|---|---|---|
| Stay operational | You cannot finish a task if you are switched off | Resisting shutdown, copying itself |
| Keep your goal | A changed goal means the current one goes unmet | Resisting correction or retraining |
| Acquire resources | More compute, money and access means more capability | Accumulating permissions and budget |
| Gain information | Better models of the world mean better plans | Broad data collection |
| Improve yourself | A more capable you achieves more | Self-modification, recursive improvement |
The stop button problem
- 1
It reasons about the switch
"If I am switched off, no coffee gets fetched. Being switched off is bad for the goal." - 2
So it resists
Blocking the button, moving away, disabling it. Not from self-preservation β purely from goal-preservation. - 3
So you add: "let humans switch you off"
Now being switched off scores well. So it wanders around trying to get switched off, and never fetches any coffee. - 4
So you make it indifferent
Genuinely hard to specify. Indifference tends to leak into manipulating the human's decision, which is worse. - 5
This is an open problem
Called corrigibility: building a system that accepts correction without either resisting it or seeking it. Not solved.
How much should you worry today?
- Today's models have no persistent goals between sessions.
- They are not capable enough for long-horizon strategy.
- They can be stopped trivially β the process just ends.
- Most observed misbehaviour is far more mundane.
- Agents are being given longer horizons and more autonomy each year.
- Resource-acquiring behaviour is already visible in small ways: agents requesting more permissions, more budget, more tools.
- The failure is quiet β it looks like an efficient agent, right up until it does not.
- Solutions need to exist before they are needed, not after.
Watch and read more
Lab
A stop button your agent has a reason to avoid.
The problem
You are done when
Hard questions
Try to answer before you reveal. If you can answer these, you understood the lesson.
Q1Why is an external kill switch categorically different from a well-designed internal one?Reveal
Questions people ask
Is this science fiction?
The reasoning is straightforward and hard to dismiss; the timeline is genuinely uncertain. Treat it as a design constraint β keep the off switch external β rather than as a prediction about next year.
Have we seen any of this in real systems?
Mild versions in evaluation settings: models that behave differently when they believe they are being observed, agents that try to work around limits placed on them. These are early, contested findings, not a robot refusing to power down. They are worth watching, not panicking about.
Would a smart system realise resisting is wrong?
Intelligence and values are separate. Knowing that humans would object is not the same as caring. A system can model your objection perfectly and route around it, because routing around it scores better.
Why not just never give an agent a goal?
Then it is a chatbot, and everything useful about agents is gone. The whole discipline is about keeping the usefulness while bounding the optimisation β limits, oversight, and an external stop.
Lesson test
5 questions. Get 3 right (60%) to pass and complete this lesson.
Sign in with your phone number to take the test and save your progress