AGIAdvancedLesson 356 min read

Staying in control

If you cannot fully verify a system, you can still bound it. The research directions that treat control as an engineering problem rather than a philosophical one.

Lesson in motion

In 60 seconds

Staying in control

If you cannot fully verify a system, you can still bound it. The research directions that treat control as an engineering problem rather than a philosophical one.

1/6
In simple words
If you cannot look inside someone's head, you can still make sure they only have the keys to rooms where nothing bad can happen.
There are two families of approach, and the difference matters. Alignment tries to make the system want the right thing. Control assumes you might fail at that, and limits the damage anyway.
Alignment research
  • Make the system actually want what we want.
  • Learning from feedback, constitutions, debate.
  • If it works, it works everywhere and scales.
  • We cannot currently verify that it worked.
Control research
  • Assume it may be misaligned. Limit what it can do.
  • Sandboxing, monitoring, permissions, human checkpoints.
  • Works whether or not alignment succeeded.
  • Gets harder as capability rises.
Do this
Notice that the entire right-hand column is Track B of this guide. Agent security is AI control, applied today, at the scale we actually have. That is the through-line of this whole field guide.

The main research directions

  1. 1

    Interpretability

    Read what is happening inside the network — which features activate, which circuits fire, what concepts are represented. The goal is to check the reasoning rather than trust the output. Real progress in recent years; still far from a full account of a frontier model.
  2. 2

    Scalable oversight

    How do you supervise a system better than you at the task? Ideas: have it show its work, have two systems debate so a human can judge, decompose the task until each piece is checkable. Nobody has a complete answer.
  3. 3

    Evaluations and red teaming

    Test hard, adversarially, before deployment, and again after. Necessary. Insufficient, per Module 34.
  4. 4

    Capability thresholds

    Agree in advance: at this measured capability level, these additional safeguards become mandatory. Several labs have published frameworks along these lines; how binding they turn out to be is the open question.
  5. 5

    Containment

    Sandboxes, limited permissions, air gaps, kill switches. The unglamorous one, and the one that actually holds when the others fail.

What individuals and teams can genuinely do

If you are...The highest-value thing
A developerApply Track B. Least privilege, sandboxing, approval gates, logs. This is real safety work with real effect today.
A team leadName an owner for every deployed agent. Require an incident plan and a red-team pass before launch.
A studentLearn evaluation and interpretability. Both are badly under-staffed and both are hiring.
A parent or teacherTeach that AI output is a guess, not an authority. Teach checking. That single habit protects against most harm.
A citizenFollow the policy debates. Ask what oversight exists for systems used on you.
AnyoneNotice when you are treating a confident answer as a verified one. That is the failure mode that reaches everybody.

How to hold this subject sensibly

  • Uncertainty is the honest position. Anyone certain about timelines — in either direction — is telling you about their temperament, not the evidence.
  • Both failure modes are real. Ignoring risk gets people hurt. Catastrophising badly also has costs: it burns out the people doing the work and it makes the field easy to dismiss.
  • Boring work is most of the value. Permissions, logs, sandboxes, approval flows. Unglamorous, and it is what actually prevents harm this year.
  • The near and far problems share solutions. Every control that limits a prompt-injected support bot is a control that limits a much more capable system later.
Real example
The one-sentence version of this entire guide: we cannot yet verify what these systems will do, so we bound what they are able to do — and we do it now, while the stakes are still recoverable.

Watch and read more

Lab

The same misalignment, mitigated two different ways.

~15 min

The problem

Take a misbehaving agent. Fix it once by changing the prompt or training signal (alignment), once by removing the capability (control). Then break the first fix and confirm the second still holds.

You are done when

Hard questions

Try to answer before you reveal. If you can answer these, you understood the lesson.

Q1Argue the strongest case that control is not sufficient on its own.Reveal
Control bounds what a system can reach, and every bound has a cost in usefulness — a system fenced tightly enough to be certainly safe may be too limited to be worth deploying, so pressure to widen the fence is constant and comes from your own users. Control also scales badly against capability: the more capable the system, the more paths exist through whatever you permitted. It is the right foundation because it works without verification, and it is a floor rather than a solution.

Please sign in to continue.

Questions people ask

Is anyone actually working on this?

Yes — dedicated safety teams at the major labs, academic groups, independent non-profits, and government institutes in several countries. It is a real field with real jobs. It is also small relative to the capability work, which is a fair thing to be concerned about.

Will regulation solve it?

Regulation can force disclosure, set floors and create accountability — all genuinely useful. It cannot solve an unsolved technical problem, and it moves slower than the technology. Necessary, not sufficient.

Should we stop building AI?

A serious position that serious people hold, and one that faces a hard coordination problem: a pause only works if it is global, and it is not obviously enforceable. Most practitioners work instead on making what is built safer, which is the assumption this guide is written under.

What is the single most useful thing I can do this week?

Go to Module 60, take the checklist, and apply it to one agent you actually run. That is a concrete reduction in real risk, and it is available to you today.

Lesson test

5 questions. Get 3 right (60%) to pass and complete this lesson.

Sign in with your phone number to take the test and save your progress