AGIBeginnerLesson 275 min read

The capability ladder

Intelligence is not one switch that flips. Click through the rungs to see what changes β€” and what it means for whoever is supervising.

Lesson in motion

In 60 seconds

The capability ladder

Intelligence is not one switch that flips. Click through the rungs to see what changes β€” and what it means for whoever is supervising.

1/4
In simple words
Think of steps, not a magic door. Each step up, the machine can do a bit more on its own β€” and each step up, you have to watch it a bit differently.
A ladder is a better mental model than a finish line. What matters at each rung is not a label but a practical question: how much can it do before a human needs to look?

Interactive Β· the capability ladder

Rung 3 β€” AgentIt acts, you check afterwards. It plans, uses tools, works for minutes or hours. It can send, write, deploy and spend. Oversight: approval gates on the irreversible, plus logs someone reads. This is where the frontier actually is today.

The pattern to notice

Read the rungs again and watch what actually changes. It is not "how clever" β€” it is how long it can run unsupervised before something goes wrong. That number is the real measure, and it is the one that decides how much of Track B you need.
toolassistantagentautonomousunsupervised timesecondsminutes to hoursdays β€” nobody is watchingabove this line, oversight must be designed in
The curve that matters. Capability is the horizontal axis; the vertical axis is how long the system runs before a human sees anything. Everything in Track B exists to keep that second number honest.

Where we are, honestly

  • Today's best systems sit solidly at the agent rung: they plan, use tools, and work for minutes to hours on real tasks.
  • They are unreliable over long horizons β€” small errors compound, and they do not always notice.
  • They do not learn from yesterday. Every session starts fresh unless you engineer memory around them.
  • They are extremely uneven: superhuman at some narrow things, oddly poor at things a child finds easy.
Do this
This unevenness is itself a safety fact. A system that is brilliant at writing code and bad at judging consequences is exactly the combination that needs external limits rather than trust.

Watch and read more

Lab

Your own systems placed on the capability ladder, with the oversight each needs.

~10 min

The problem

Place three systems you use on the six-rung ladder. For each, measure the real number: how long does it operate before a human sees anything? Then state the oversight that number demands.

You are done when

Hard questions

Try to answer before you reveal. If you can answer these, you understood the lesson.

Q1Two systems are equally capable. One is far more dangerous. Give the mechanism in one sentence, then unpack it.Reveal
Autonomy multiplied by permissions, not intelligence. Unpacked: capability sets what a system could do; unsupervised operating time sets how long it does it unnoticed; the permission set sets what damage is reachable. A brilliant model with read-only access in a sandbox is safe; the same model with production credentials and an overnight schedule is not. You control the two multipliers.

Please sign in to continue.

Questions people ask

Do we have to climb every rung in order?

Probably not neatly. Progress has been lumpy β€” sudden jumps in some abilities while others stay flat. A system could be superhuman at research and still unable to reliably book a taxi.

Which rung is dangerous?

Any rung where the system acts faster than you can check, on things you cannot undo. That threshold is crossed today, at the agent rung, by ordinary business software. It is not a future problem.

Does more capability mean more danger?

Capability multiplies whatever autonomy and permissions you granted. A more capable model with read-only access in a sandbox is not more dangerous. The same model with production credentials is much more dangerous. You control the multiplier.

Is there a rung above superintelligence?

People speculate about recursive self-improvement β€” a system that improves itself, which then improves itself faster. Whether that is a real dynamic or a story we like telling is genuinely unsettled.

Lesson test

5 questions. Get 3 right (60%) to pass and complete this lesson.

Sign in with your phone number to take the test and save your progress